Before your business can deploy an agent, you need to answer a question that sounds simple and turns out to be devastating: how do you know when the output is good?
Not approximately good. Not good enough for a human who can fill in the gaps. Good by a standard that can be evaluated consistently, by something with no judgment, context, or professional intuition.
Most businesses cannot answer this question. Not because you haven’t thought about it, but because you’ve never had to. Quality in most businesses has always been a human judgment. You know good work when you see it. You have trained newer practitioners to develop the same instinct. The standard has lived in people, passed down through apprenticeship, feedback, and example.
That works when every output passes through a human who carries your standard. It stops working when your agent produces the output and no one checks it until it reaches your client.
Why evaluation is harder than it looks
The instinct is to define quality as the absence of errors. No factual mistakes. No logical gaps. No missed steps. This is necessary but not sufficient. Business outputs are not just technically correct — they are appropriately calibrated to your context.
A memo that is technically accurate but addresses the wrong question is a bad memo. A financial model that is arithmetically correct but built on assumptions you never validated is a dangerous model. A due diligence report that covers every item on your standard checklist but misses the non-standard risk that was visible in your data is a failure of judgment, not a failure of process.
These are the things you catch with experience. They are also the things agents reliably miss, because they are not in the specification.
This is what makes evaluation hard. The standard is not just what is in the output. It is the relationship between the output and the context in which you will use it — and you have almost never written that relationship down.
The proxy problem
When you try to build evaluation, you usually reach for proxies. Completeness: did your agent address all the required sections? Consistency: does the output match your template? Accuracy on verifiable facts: did you cite the right numbers?
These proxies are better than nothing. They catch obvious failures. But they also create a false sense of coverage — a feeling that because your agent scored well on the measurable things, it probably did well on the unmeasurable ones.
It usually didn’t.
The problem is that the failures that matter most — the missed implication, the unwarranted confidence, the context that should have changed the conclusion — are precisely the failures that proxies don’t capture. They pass your checklist and fail your review, if there is one.
When there isn’t one, they reach your client.
Building an evaluable standard
The path forward is not to find better proxies. It is to do the harder work of making the actual standard explicit.
This means starting with the outputs you are proud of — the work that, when you send it, you know it is right — and working backwards to ask what makes it right. Not “it covers the required sections” but “it frames the issue the way your client will think about it.” Not “it is consistent with your template” but “it identifies the one thing your client needs to act on and leads with that.”
Then it means asking whether those things can be evaluated without a human, and if not, what a human evaluator would need to see to make the call quickly and reliably.
For most businesses, this process reveals that the evaluable standard is smaller than you thought, and the judgment-dependent residual is larger. That is uncomfortable but useful. It tells you exactly where agents can operate autonomously and where they need human review — and it makes the review itself more efficient, because the reviewer knows what to look for.
The standard as competitive infrastructure
There is a second reason to do this work beyond enabling agents. Businesses that build an explicit, evaluable quality standard have built something that did not exist before: a transferable definition of their quality.
This is not a small thing. It means you can calibrate new practitioners faster. It means quality can be monitored across your business, not just in your leadership team’s heads. It means your business has an asset — a codified standard — that compounds over time as it is refined and improved.
The businesses that will lead in the agent era are not the ones that deploy agents fastest. They are the ones that know, with precision, what good looks like — and have built the infrastructure to enforce it at scale.
That standard is built before the agents arrive. It is one of the things that makes agent deployment possible at all.