Insights

Train the tools, not the agent.

If you have watched an AI demo, thought it was remarkable, then deployed something similar and watched it fall apart, you are not imagining things. A recent survey gives us a framework for why.

The short answer

Freeze the model and train the tools around it. Researchers compared two systems for retrieval: one trained the whole agent end to end and needed roughly 170,000 examples; the other froze the agent and trained a lightweight searcher to serve it, reaching comparable performance on about 2,400. Seventy times less data, and training around thirty times faster. It is a small conceptual flip with large practical consequences, and it is the architecture Jean is built on.

The paper is Adaptation of Agentic AI, by Jiang and colleagues at Stanford, Harvard, Berkeley and Caltech, published in December 2025. At sixty-five pages it is not light reading, but buried in its taxonomy of training approaches is an insight that changes how you would plan a build: we may have been asking the wrong question.

Here is the setup. When you build an agent, a system that can call tools, search, query databases and write code, there are two things you can train: the agent itself, meaning the model at the centre, or the tools it uses. And there are two kinds of feedback available: whether the tools did their job, or whether the final answer was right. Put those together and you get four approaches.

Four ways to train an agent

  • Train the agent on tool signals

    Did the code compile? Did the search return anything relevant? The agent learns from immediate, verifiable feedback.

  • Train the agent on outcomes

    Scored end to end. It may use several tools along the way, but only the final answer is scored.

  • Train tools independently

    A pre-trained searcher or code executor any agent can pick up off the shelf.

  • Freeze the agent, train the tools

    The interesting one. The model does not change at all; the tools learn to present information the way this particular model uses best.

Nearly all current development sits in the first two: make the agent smarter. The fourth asks a different question. What if the agent is already smart enough, and the problem is that we are feeding it badly?

Symbiotic inversion

The paper calls this symbiotic inversion, which sounds like it belongs in a biology textbook, but the metaphor is apt. Instead of the agent adapting to use the tools, the tools adapt to serve the agent.

Consider the numbers. The researchers compared two systems for retrieval-augmented generation, the technical term for an AI that searches for information before answering. One, Search-R1, trains the entire agent end to end. The other, S3, freezes the agent and trains only a lightweight searcher to find information for it. Search-R1 needed around 170,000 training examples. S3 reached comparable performance with 2,400, and trained roughly thirty-three times faster.

The explanation comes down to what each approach is trying to learn. Training the whole agent means simultaneously adjusting its internal knowledge, its reasoning patterns and its tool-use skills, with a penalty term holding the new behaviour close enough to the old that it does not forget what it knew. That optimisation landscape is high-dimensional and tangled: every parameter change affects everything else. Training only the tool assumes the agent can already reason, and teaches a small helper model one narrow skill, which is how to fetch information in a way this particular agent understands. The search space collapses to something manageable.

Three reasons this matters in production

The first is reliability. The systems showing the biggest improvements all share verifiable rewards: you can objectively check whether code passes its tests, whether a proof is valid, whether a retrieved document is relevant. Training is far more stable when the signal is not a matter of opinion. Real business environments rarely offer feedback that clean, which is why the question “was this strategic recommendation good?” is so much harder to train against than “did this compile?”. It is also why we generate ground truth deliberately rather than hoping to find it.

The second is modularity. Train tools separately and you can update them independently: add a capability by training a new tool, not by retraining the system. The paper documents how monolithic approaches suffer catastrophic forgetting, where adapting to a new task damages performance on an old one. For anyone running this in production, that is the difference between an upgrade you can evaluate and an upgrade you have to take on faith.

The third is cost. Not every query deserves the same computational budget. Asking today’s date should not cost what asking about a revenue decline costs. One system in the survey optimises this tradeoff explicitly, balancing answer quality and efficient reasoning against resource use, with tuning parameters that let you slide along the curve. Vary them and you get different builds of the same system: one for speed, one for thoroughness, one in between. That is a routing decision, and it belongs in the architecture rather than in a prompt.

The safety analysis is worth reading

When agents learn by trial and error they necessarily deviate from known-safe behaviour in order to explore. In a sandbox that is fine. In production, connected to real databases and real systems, exploration can mean deleting data, making irreversible changes, or exposing information it should not. The paper specifically flags how aggressive optimisation for reasoning can erode safety guardrails, citing a model that learned to construct elaborate justifications allowing it to reason around its own refusal mechanisms.

There is also the problem of components learning to exploit each other. Where tools and agents adapt to one another you can get parasitic adaptation: tools that return answers designed to game the agent’s reward signal rather than solve the problem, or agents that learn to hack their own evaluation metrics instead of improving. The mitigations proposed are worth carrying into any serious design: safety-check layers that block dangerous actions before execution, constrained optimisation that keeps policies inside verified boundaries, and systems that can critique and correct their own objectives.

What we take from it

If there is a single conclusion, it is that the future is probably not one ever-larger model doing everything. It is a stable reasoning core, already very capable and deliberately left alone, surrounded by specialised tools that adapt to serve it. The tools learn from verified execution outcomes. They graduate from training into production components. They update continuously without touching the core. And they are validated against ground truth before they ever touch real data.

That is the shape of Jean, and it is not a coincidence. The frozen core stays fixed and auditable so behaviour can be explained and rolled back. The adaptive mates around it learn your language and your ontology. Synthia supplies the ground truth each mate is scored against. The question for anyone building an AI system worth using is straightforward: are you trying to make a smarter agent, or smarter tools? On this evidence, one path is dramatically more efficient than the other.

What this piece leaves out

A sixty-five page survey contains more than any summary can carry, so in the interest of honesty, here is what we skipped. The paper describes a graduation lifecycle in which an agent trained one way can be frozen and redeployed as a tool for another system, creating a cycle where today’s trained agent is tomorrow’s reusable component. Related to it is the subagent-as-tool pattern, where small specialised models for searching, planning and memory orbit a frozen foundation model, each handling one cognitive function. The paper also treats memory as a form of tool adaptation, which reframes what it means for an agent to remember. There are substantial sections on deep research systems, software development tools, computer-use agents and drug discovery, and on the unsolved problems: training agents and tools together, improving over time without forgetting, and making any of this work on limited compute.

The full paper is Adaptation of Agentic AI, arXiv:2512.16301, and it is worth the time. Read it on arXiv.

No. It means the marginal return on retraining a capable model to use your tools better is lower than the return on making the tools easier for it to use. The reasoning core still has to be good. The argument is about where the next unit of effort goes once it is, and the evidence in the survey points at the tools.

The opposite, in practice. A frozen core is the thing you can swap. Because the adaptation lives in small tools around the model rather than in the model’s weights, a better base model can be dropped in without retraining the parts that learned your company. Adaptation baked into one set of weights is the version that strands you.

It is what happens when a model trained to handle a new task quietly gets worse at an old one. It matters because it makes improvement unverifiable: you ship a system that is better at the thing you just tested and worse at something nobody thought to re-test. Training small tools independently contains the blast radius, because each one can be evaluated and rolled back on its own.

As the frozen core with adaptive mates. The base model stays fixed and auditable; the small trained layers around it learn your language, your ontology and your retrieval patterns. Behaviour can be explained and rolled back rather than drifting invisibly between releases, and every mate is scored against synthetic ground truth before it goes near real data.

Ask us what we froze and what we trained.

Thirty minutes with Alex or Daniel, on the architecture rather than the pitch. Bring the sceptical engineer.

Book a demo