Why general-purpose LLMs fall apart on sustainability data

by

Emma Ylivainio

by

Emma Ylivainio

CPO

Follow me on:

Last updated:

CONTENT ai modified

Paste a sustainability report into any frontier model today and you'll get a fluent, well-organized summary in seconds. It'll look right. It'll read like someone who understood the document wrote it. And for a meaningful share of the claims in it, you won't actually know if it's true. The failure modes that make sustainability data hard aren't the ones general-purpose language models were built to handle.

This isn't a knock on the models. It's a mismatch between what they optimize for and what this specific data demands.

The data itself is adversarial to a single-pass read

ESG disclosure isn't clean, structured, machine-friendly content. It's narrative prose, footnoted tables, infographics, numbers that only make sense next to a methodology note two pages away. When documents get chunked and embedded as plain text, a PDF flattens into a token stream, and that's exactly where the structural context goes missing: which reporting year a number belongs to, what scope boundary it assumes, whether it's a restated figure or the original. Reading the page as an image instead, and treating a table cell as its own retrievable unit rather than smearing it into a paragraph, is the difference between citing the real number and citing a plausible-looking one.

There's no ground truth to check against, by default

Nobody has secret access to better data. The sustainability report everyone reads is the same one. What differs is what happens next. A general model reads it once and answers with the same confidence whether it's right or wrong, with no check on whether the number it just quoted is the number actually on the page or just sounds like it. Working from the same document as everyone else isn't the hard part. Reading it properly is.

Confident language and vague commitments look identical to a language model

Sustainability reporting is full of aspirational phrasing that's deliberately soft: "working toward," "exploring options for," "committed to a pathway." A model trained to produce fluent, confident text doesn't have a strong signal for the difference between a measurable, dated commitment and a sentence engineered to sound like one. Left alone, it'll launder marketing language into what reads like a hard fact. Catching that distinction is closer to a classification problem than a generation one. You have to decide what a passage is before you decide what it says.

Ask it twice, get two different answers

A general model doesn't look up facts from a fixed record. Every time you ask it something, it searches for whatever passage seems close enough to your question, and "close enough" can land somewhere different next time, even for the exact same question. Ask about the same company twice, get two different answers pulled from two different places, no explanation for why they differ. Fine if you're browsing. Not fine if the number ends up in a supplier scorecard or a board deck, where you need to ask again next quarter and know exactly why the figure held or changed.

What actually holds up

None of this gets fixed by a bigger model or a longer context window. It's an architecture problem, not a scale problem, and the architecture has to do more than retrieve text. It has to construct data that doesn't exist yet.

Take Acts and Goals: what a company has actually done, and what it's committed to doing. Neither is a clean data point sitting in a report waiting to be pulled out. A report just says, somewhere in a paragraph, "we reduced emissions." Turning that into a measurable record (a defined metric, a figure, a unit, a year, a source page) is synthesis, not extraction. It's a new layer of validated data that didn't exist until the pipeline built it: the difference between a static PDF and a database row you can filter, score, and rank with confidence.

Getting from raw disclosure to that structured record runs through roughly 34 processing steps: classification, cross-checking, scoring, etc. Every step is a gate. Is this actually a measurable commitment or just aspirational phrasing? Does this figure match the exact page it's cited from? Has this same claim already been logged from another document? A claim only becomes a permanent record once it clears every gate. From there it's written into a structured, citable database and immediately discoverable across the platform, instead of sitting unread in a PDF.

That's the real output: verified, structured sustainability data at a depth that's close to impossible to reproduce by hand. A single analyst, or even a small team with limited resources, can't manually classify, cross-check, and structure disclosure at this volume across hundreds of companies and keep it current. That's precisely what the pipeline is for.

A general-purpose model can tell you what a document says. Turning that into a verified, structured, comparable data asset at scale, kept current, is a different engineering problem, and it's the one worth solving.

That's what we've built. Not for people who just want a document summarized, but for sustainability teams expected to have a defensible answer ready, and for everyone around them: sales, procurement, leadership, everyone who turns to that team the moment they need to know what a customer, supplier, or competitor is actually doing.

text ai modified

Insights

Read more articles

Questions & answers

Frequently
Asked Questions

What are the technologies behind PlanetAI?

PlanetAI combines AI, data engineering, and automation to turn scattered sustainability information into clear, actionable insights. AI Engine: Reads, interprets, and compares sustainability reports, news, and websites using advanced language models and our own custom-built classifiers Multi-Agent System: Specialized AI agents (for ESG topics, CSRD, and quality assurance) work together to deliver precise, structured insights. Secure Cloud Infrastructure: Enterprise-grade security, GDPR-compliant, and built to scale. Interactive Dashboards: Explore ESG actions, goals, comparisons, and trends in real time. In short, PlanetAI is a full sustainability intelligence engine, combining generative AI with validated ESG data on a secure, scalable platform.

How is PlanetAI different from other AI tools?
What kind of data are the profiles based on?
How reliable is the data on the platform?