- Introduction
- When It Comes to ESG Data Aggregation,
Transparency is a Farce - It Keeps Getting Worse:
The Raw Data Doesn't Even Add Up - Your ESG Ratings Inherit a
Problem You Never See Coming - So, Stop Asking ESG Data Vendors
“How Accurate Is This Data” - Does This Mean You Can't
Trust Any ESG Data? - The Companies Are Being Honest. The
ESG Data Supply Chain Needs to Step Up
74%. That is the share of S&P 500 companies that went back and revised previously reported greenhouse gas emissions at least once between 2010 and 2020, according to a Harvard Business School study published in January 2026.
In most cases, the earlier reported numbers were increased. For Scope 1 emissions alone (the emissions a company produces directly, and the easiest category to measure), the quietly added total comes to roughly 135 million tonnes that earlier reports had simply left out.
That is not a rounding error. That gap is comparable to the annual emissions of a small industrialized country.
The revisions themselves aren't a problem. Companies revise emissions for real reasons — better science, a merger that redrew the boundaries, a corrected method. But no third-party ESG data provider acknowledges out loud what happened to those revisions after they left the company. Did they ever enter the database that's supposed to sell you trustworthy data? They likely didn’t.
Somewhere between the disclosure and your dashboard, the correction got erased, and your ESG ratings became ineffective, with no one the wiser.
[Source- Many Companies Quietly Revise Their Emissions Data: One Chart | Working Knowledge]
When It Comes to ESG Data Aggregation, Transparency is a Farce
Start with how little the system asks of anyone.
When a US public company restates its revenue (something barely a few hundred companies do in a given year), it owes the SEC an explanation. There is a filing. A paper trail. A market that reacts to the missing data.
When a company restates its emissions by hundreds of thousands of tonnes, it owes precisely nobody an explanation. The Harvard team found that revisions are rarely flagged at all, occasionally a footnote, most often not even that.
Sit with the asymmetry. The same market that would crucify a company for burying an earnings restatement waves through a decade of unexplained emissions rewrites without blinking. And the safety net everyone points to (third-party assurance, an auditor, or a specialist like a Big Four firm) doesn't catch it. The Harvard study found no meaningful link between whether a company's emissions were assured and whether it later revised them.
ESG reporting is “voluntary”. The actual cost of that voluntary disclosure is that the revisions are unmarked. The single most useful quality signal a data point can carry — "this figure changed, here's when, here's why" — is precisely the signal the voluntary system does not require anyone to preserve, report, or consider seriously.
It Keeps Getting Worse: The Raw Data Doesn't Even Add Up
If you think the pipeline at least sanity-checks what it ingests, one more study should end that comfort.
Researchers Garcia-Vega, Hoepner, Rogelj, and Schiemann took the Scope 1 emissions that 33 major oil and gas companies reported to CDP (the disclosure platform routinely called the gold standard for ESG) and ran the most trivial test imaginable: do the reported sub-totals add up to the reported total?
In 38.9% of cases, they didn't. More than a third of the numbers failed the kind of check a spreadsheet does automatically.
And as the authors point out, other data providers and rating agencies rely on that same CDP data. Numbers that can't pass basic arithmetic flow straight into the data aggregation layer and out to buyers, because nothing at the door is checking. Assurance didn't help here either. Even the aggregators half-admit the rot: S&P Global's own Trucost analysis found that in 2022, only 32% of companies reported Scope 1 completely enough to need no supplementing. The other two-thirds is estimated and modeled, which is fine, if anyone tells you which data is exactly as reported and which is estimated. Usually, no one does.
Your ESG Ratings Inherit a Problem You Never See Coming
A single unreliable disclosure in a single company's report is a contained problem. Any competent analyst can flag it. The damage is done when that disclosure enters the aggregation layer. ESG data aggregators, i.e., the vendors and platforms that collect thousands of companies' numbers, standardize them, and sell them onward as clean, structured data, follow an almost boring mechanism, which is why the data gap goes unexamined.
For example, a data aggregator ingests a company's 2019 emissions and stores it as a value. The company later revises and increases the emission value. The vendor most likely does one of two things:
- It doesn't update its database at all — and the Harvard study found exactly this, noting that data providers do not appear to uniformly correct these revisions.
- It overwrites the old value with the new one — and in doing so, quietly destroys the evidence that a revision ever happened.
The second case is more worrying because it looks like diligence. It looks like the numbers you have are updated, confident, and final, with no indication that it was different last year, no record of the magnitude of the change, and no way to distinguish between three very different situations:
- a footprint that has been stable and trustworthy for a decade
- a footprint that was corrected last quarter for a legitimate methodology change
- a footprint that was understated for ten years and has only just been truly disclosed
To a scoring or rating algorithm, there is no difference between a company that reported its carbon footprint accurately for a decade and one that hid its footprint for a decade and was suddenly forced to reveal the actual numbers. This issue hits downstream systems even harder.
A rating uses the vendor's emissions number. An index uses the ratings. So if the emissions number is missing its history, every layer built on top inherits that same blindness. Or, when you back-test a scoring model, you're asking whether it would have worked in the past (say, 2015), which means feeding it the data as it stood in 2015. But if your vendor has overwritten that year's emissions with a figure the company only revised in 2022, your model is being tested against numbers that didn't exist yet. It appears to have judged 2015 correctly, when in truth it was handed a number that did not exist in the 2015 disclosure. Instead of validating your rating methodology, it quietly rewards it for knowing things it couldn't have known and hands you false confidence.
So, Stop Asking ESG Data Vendors “How Accurate Is This Data”
If you run a rating platform, a scoring firm, or any product that consumes ESG data at scale, the Harvard and CDP findings should reframe your vendor relationship.
You have been trained to ask vendors the wrong question. "How accurate is this data?" is a point-in-time question, and any vendor can pass it with a number that's correct today and useless for anything that depends on history. The question that actually exposes the gap is: "What does this data know about its own past?"
Here is what you should ask before you feed any data pipeline to your ESG rating/scoring algorithm.
1. Do you keep revision history, or just the latest number?
If a company restates its 2019 figure in 2023, do you still hold the original alongside the revision, with dates? If a vendor keeps only current values, every back-test you run on their data will be using 2023's corrected numbers to test decisions made in 2019. So the results are meaningless before you start. If a vendor can't answer this one cleanly, nothing else matters.
2. Can you tell me, per data point, what's reported versus estimated?
Some emissions numbers come straight from a company's own disclosure and have been checked against other sources. Others are just a sector average, an educated guess used when the company reported nothing. Those two should not carry equal weight in your model, but they will if the vendor doesn't tell you which is which. Good ESG data vendors label every figure by how solid it is (for example, a simple Exact/Best-Estimate/Tentative scale), so you can trust the strong numbers more than the weak ones.
3. Can you trace every figure back to a specific source document?
Your rating algorithm does not need a general description of where the data "tends to" come from. It needs a per-datapoint link to the exact filing, page, or disclosure. If you can't trace a number to its source, it will not survive an auditor, a regulator, or a client who decides to check your work. And such untraceable data is a liability you will bear repercussions for.
4. Do you track why a number changed, or do you just record that it did?
Many revisions are completely legitimate. A company acquires a factory, so its emissions jump, not because it polluted more, but because it's now counting a bigger operation (this is what it means when we say a “reporting boundary changed”). Or it adopts a better calculation method, and the number moves. These aren't errors. But they only help you if the vendor tells you why the figure changed. A vendor who notes "this rose in FY22 because an acquisition widened the reporting boundary" is giving you something usable. You can tell that the revised figure is structural, not a red flag. A vendor who just swaps in the new number, with no explanation, leaves you to guess whether it signals a real problem or nothing at all.
Does This Mean You Can't Trust Any ESG Data?
No. It means you can't trust a bare value without knowing its provenance. Data that carries its revision history, source lineage, confidence tier, and a clear reported-versus-estimated flag is trustworthy and defensible.
At SunTec India, the ESG data research practice is built around exactly these answers because the failures behind them are the ones our clients kept coming to us with. With ESG rating agencies and 3P data aggregators, we have seen a similar pattern: they distrust the data because the audit trail is weak or doesn't exist. So we operate the ESG and exposure data research layer underneath them, supplying the rating agencies, data providers, and sustainability-intelligence firms with verified ESG data and a strong chain of evidence.
- Data estimates are flagged and held separate from as-reported data.
- Every figure carries an explicit Exact/Best-Estimate/Tentative score
- Each data point comes with its source evidence attached and an audit trail of every verification step
- Data from two companies (reported using different units and scope boundaries) is standardized for easy comparison.
- Unexplained revisions are reviewed by data analysts and added to your dataset with a reason and supporting documentation attached.
The Companies Are Being Honest. The ESG Data Supply Chain Needs to Step Up
The 74% of companies that revised their emissions upward were, in their clumsy and unannounced way, being honest by correcting the record. At this point, the dishonesty exists in the layer where all of this data is aggregated and packaged, and the supply chain built to carry these disclosures. The existing ESG data aggregation system was never designed to carry change. But change, it turns out, is the normal state of emissions data, and the industry must adapt.
Pavan Kakar
Pavan Kakar, Associate Vice President of International Sales at SunTec India, is an experienced professional in the data services domain with over two decades of global experience. He drives enterprise growth across 50+ countries and is a recognized opinion leader for data-driven innovation, human-in-the-loop AI, and business intelligence.