Insights

What fixing the kilogram taught us

For a hundred years, one kilogram meant one lump of metal in a vault outside Paris. How physicists replaced it is the cleanest way we have found to explain why we built Synthia.

The short answer

An evaluation set that is its own reference will drift, and you will not be able to tell. So Synthia fixes the truth first, as a graph of facts, and generates the company around it. Every artefact points back to the fact it is claiming to express, which means every question about that company comes with an answer key that was written before the question was.

For over a century, if you wanted to know what one kilogram actually was, the answer was a platinum-iridium cylinder sitting in a vault near Paris. It had a name, Le Grand K, and a nervous team of metrologists who would occasionally take it out and compare it against the official copies scattered around the world. Every scale on earth was calibrated, eventually, against that one object.

The problem was drift. The cylinder picked up contamination from the air. The copies drifted too, at different rates. And nobody could say whether the original had gained mass or the copies had lost it, because the only way to check the original was to compare it against the copies, and the only way to check the copies was to compare them against the original. Le Grand K was, by definition, one kilogram. If it gained weight, the kilogram gained weight.

In 2019 physicists retired it. They redefined the kilogram in terms of the Planck constant, a number you cannot contaminate or drop on the floor, and any well-equipped lab can now realise a kilogram on demand with an instrument called a Kibble balance. The reference stopped being a thing and started being a property of nature.

Real enterprise data is the lump in the vault

Agents meant to work inside an organisation read threads, update tickets, draft incident reports and summarise yesterday’s stand-up. If you are building those, you have to evaluate them against the kind of messy, half-finished, context-dependent material they will actually meet. The options available today are all bad, and they are bad in related ways.

You can get hold of a real company’s data under an agreement, which is scarce, cannot be shared, cannot be varied, and is a single snapshot of a single organisation at a single moment. You can scrape whatever is public, which is noisy, the wrong shape, and carries no ground truth at all: you do not know what the correct answer was. Or you can write scenarios by hand, which works right up to the moment you want a second one, or a variant, or five thousand.

All three share the deeper problem. There is no independent reference for what is true, because the dataset is the reference. If it is quietly biased, or incomplete, or missing the exact workflow your agent is going to meet, nothing tells you. You optimise against the lump, the lump drifts, and the numbers go up and to the right regardless.

Stage one is the Planck constant

The first thing Synthia does is build a fact graph: a structured record of everything that is true about a fictional project. What the team decided at kickoff, which modules depend on which, what got renegotiated in week two because somebody realised the original plan would not work. These are stored as facts with relationships between them, not as text.

What matters about the fact graph is that it sits underneath any artefact generated from it. You can inspect it, query it, and diff two versions of it. It does not drift, because it is not a measurement of anything. It is the definition the rest is derived from.

Stage two is the Kibble balance

Having the graph is not enough. It has to be realised as the cluttered artefact stream a real company produces, and that is the second stage. It wraps the fact graph in an enterprise: it assigns people with calendars, workloads and communication styles, builds workstreams with owners, schedules meetings, and spins up the threads, the tickets and the documents. Every message, every ticket comment and every revision points back to the specific fact it is claiming to express.

Which means we always know what an artefact is supposed to be saying, even when the fictional person writing the fictional message is sloppy, or miscommunicates, or forgets to mention the thing they agreed to an hour earlier. The truth lives upstream of the text.

What that actually buys you

The obvious thing is answer keys. Ask an agent whether the team approved the new pricing model and we can grade the reply, not because we remember generating it, but because stage one recorded the decision and stage two tracked which artefacts mentioned it. Once answer keys exist, the rest arrives for free.

  • Answer keys

    Grading becomes a comparison against a planted fact rather than a second opinion from another model.

  • Counterfactuals

    The same project across three time zones, or with the decision quietly reversed a month later. Re-runs, not new datasets.

  • The hard cases

    Not whether it does fine on average, but whether it finds the decision buried in a message nobody pinned.

That last one is the point of the whole exercise. The interesting question is never whether an agent performs acceptably on average. It is whether it gets the consequential decision right when the only record of it is a three-day-old direct message and the ticket says nothing more than “per our discussion Tuesday”. Because evidence pointers track which facts ended up in which artefacts, that exact scenario can be constructed on purpose and scored.

The panel below is one run against a synthetic company. As with everything on this site, the figures are illustrative unless a source is named.

Jean
→ MODULE DAG→ SYNTHIA→ THE TWIN
Module DAG
Synthia
The twin
How the twin is generatedcomplete
The digital twin is ready

Every claim traces back to the module that owns it.

29
modules
5,637
claims
11,838
Q&A
Answer accuracy
96%
Path correctness
91%
Questions per run
1,480
Refusals when ungrounded
100%

Flight simulators, for anyone allergic to metrology

If the kilogram does not land, here is the shorter version. Airlines do not train pilots by crashing real aircraft. They build simulators grounded in aerodynamics. The physics is the ground truth, the simulator is the pipeline, and the scenarios pilots fly are the artefacts. Unlimited realistic experience, nobody hurt, and every scenario arrives with an answer key because the instructor built it.

Synthia is that, for the enterprise agents everyone is trying to ship this year.

Why this is not just fake messages

Asking a model to write a fake channel gives you text that looks like enterprise data. It does not give you a truth layer. Ask that same model, or the agent you are trying to evaluate, what the team decided in sprint three, and you get another plausible-looking answer with no independent way to check whether the two even agree, let alone whether either is right.

Synthia’s fact graph is the reference. The artefacts are what comes out when the graph is run through the Kibble balance. The whole point of evaluation is the gap between those two things, and if you skip the graph, there is no gap to measure.

Le Grand K did not get heavier or lighter on the day it was retired, and no bathroom scale noticed. What changed was where the definition lived. It stopped being a lump in a vault and became something any lab could derive from first principles.

That is what we are trying to do with evaluation.

Because a public benchmark has never seen your company. It reports how a model handles general knowledge, which is not the question a risk committee is asking. The question is whether the system is right about your contracts, your reporting lines and your systems, and no published test set covers that. Synthia exists to make one.

The value is not the text, it is the answer key. A synthetic company is generated from a fact graph, so every fact in it was placed deliberately and the correct answer is known before the question is asked. That turns evaluation from a judgement call into a comparison. It also breaks the circular problem of needing confidential records to test a system nobody has agreed to trust yet.

A model asked to invent a channel produces text that looks like enterprise data and carries no truth layer underneath it. Ask that same model what the team decided in sprint three and you get a second plausible answer with no independent way to check whether the two even agree. In Synthia the fact graph is the reference and the artefacts are derived from it, so the gap between the two is measurable.

No. They are from a run against a synthetic company and they are illustrative. Nothing on this site is a customer result unless a source is named. In a briefing we will run the measurement against a synthetic company modelled on your sector and show you the steps it failed.

Watch a run fail on purpose.

Thirty minutes with Alex or Daniel. We will show you the steps Jean got wrong, not only the ones it got right.

Book a demo