
Table of contents
A practical framework for evaluating agentic AI systems for reliability, behavior, and production readiness.
An agent can pass an evaluation and still fail the business process it was built to run.
A text-generation model gives you an answer to inspect. An agent can issue a refund, update a ticket, or write to the system of record. Before you let it act, you need evidence not only that it can reach the right outcome, but that it can do so repeatedly, within the right boundaries, and at an acceptable cost.
That evidence is harder to build when the agent is new. There is no production traffic to sample, no labels, and often no baseline. The outcome also depends on more than the model: the harness, tools, permissions, prompts, memory, retries, API versions, and environment state can all change the trajectory.
This is the problem addressed here. We start with lessons from classical ML evaluation and follow them into agentic systems: how to build an initial dataset before production usage exists, make runs comparable, verify behavior as well as outcomes, and trust the metrics used to gate releases.
Our earlier article, The top-scoring model is not always the best production choice, looked at a comparison problem: given several candidate systems, which one performs best under the same conditions?
This article addresses the earlier question: how do you build enough evidence to evaluate an agent when the system — and its production data — do not yet exist?
The framework has four layers: the dataset, the execution setup, behavior verification, and the metric. A strong offline score does not guarantee that the system will hold up in production. An agent can clear a benchmark while its orchestration is still broken.

TL;DR
- Start with a small, hand-authored evaluation set. Extend it with carefully reviewed model-generated cases and simulated users, then replace both with real traces as soon as the agent gets production exposure.
- Version the harness and environment. Use the same tools, schemas, permissions, prompts, and seeded state when comparing runs.
- Run every case multiple times and gate releases on pass^k, not pass@k.
- Verify the trajectory, not only the final answer: tool choice, order, failure handling, state changes, cost, and latency all matter.
- Use a small set of bounded, anchored, monotone metrics, and validate every LLM judge against held-out human labels.
- Once the dataset reflects real failures, use it as a regression suite so refactors can ship without repeating manual verification.
1. Build the dataset before production exists
The most trustworthy dataset is one that’s a representative sample of your production usage. That’s not available when you’re developing an agent: something that doesn’t exist yet can’t have run in production, so there’s no usage yet to sample. Building the tools an agent relies on, MCP servers, skills, any integration point, is pure engineering: you write code before any interaction data exists.
At this point you can already run engineering checks on the tools and the harness themselves: contract tests on each tool’s inputs and outputs, structural checks on what the harness exposes. Zero interaction data needed, so do this regardless. What you can’t evaluate yet is behavior: how your agent actually uses those tools once it’s making its own decisions.
That changes what a case has to contain. A classical row is an input and the label you expect back. An agentic case starts from a state rather than a single input: the conversation up to that point, and the environment the run assumes, seeded to a known starting position. What it expects back is a description of acceptable behavior rather than one right answer, because several different trajectories can each be fine. Which tool calls have to appear, which must not, what the run has to stay inside on time and cost. §3 covers how to write those.
Build a starting dataset yourself. A small hand-authored set covering the use cases and happy paths you expect, aiming for uniform coverage across the capabilities you claim it has, extended with two techniques that scale it further once writing every case by hand stops being feasible:
- Synthetic data generation. Have a model expand your hand-authored cases, or generate new ones from a description of the use case. The economics work because reviewing a generated case is quick where authoring one is slow: you’re approving, not writing. Approving is the part you can’t skip. Eval data nobody looked at was produced by a model, which means it can be wrong in exactly the ways the agent you’re testing is wrong — and then the dataset inherits the failure it was supposed to catch. Confirm it’s plausible enough to be a real conversation, and that it exercises the capability you meant it to and not a neighboring one. Then check that the set as a whole covers everything you expect the agent to handle.
- Simulated users. A simulator drives multi-turn conversations against the agent, so you get emergent behavior instead of a scripted case. Reach for it when even generation can’t get you enough real-shaped cases. Conversations are generated on the fly, so you can’t review each one before it’s used the way you would a synthetic case. Oversee the simulator instead: periodically check that it’s holding reasonable conversations, and go in with a plan to replace it with real production data once that exists. This is the less reliable of the two techniques. It puts enough noise into the score that this data can’t gate a release.
None of this, hand-authored or generated, is your production dataset. It’s simple by construction, and it’s what lets you start testing before real usage exists. The moment the agent gets any real exposure, even limited or shadow use, that’s what actually starts closing the gap:
From there, the dataset should start absorbing that real traffic. Expect a saturation curve as you go: early on you’re adding plenty of cases the agent still fails, each one a real gap worth having. As development continues, you push the pass rate on that same set toward 100%. Once it’s there, the set doesn’t retire; it becomes your regression suite: the gate a new version has to clear before it ships.
One caveat on timing: don’t scale the dataset before the solution is stable. UAT tends to force redesigns rather than prompt tweaks, and a large dataset built before that is rework – pipelines rebuilt, benchmarks meaningless. Orchestration before optimization puts the threshold at ~75% reliable behavior with real users. Method first, then data.
2. Make runs comparable: harness, environment, and nondeterminism
In classical ML, the same input normally produces the same output. For an agent, the harness, the environment, and the run itself can all introduce variation.
An agent breaks that in three separate places. The same prompt can land on a different outcome depending on the harness it’s handed (which tools exist, live or mocked) and the environment it perceives the moment it starts working (a database row, a file on disk). And even with harness and environment held fixed, the run itself isn’t guaranteed to repeat: that’s nondeterminism.

Harness
Which tools are registered, whether their calls hit a live system or a mock, and what permission scope they run under. The analogy from classical ML is feature engineering: you keep the raw data, but the model consumes engineered features – change that feature set and every earlier score stops being comparable. The harness plays the same role for an agent, and your eval data doesn’t need to store it either.
So don’t freeze it. Replay recorded conversations against your current tools and system prompt, which is what lets you iterate on the harness at all. A recorded conversation can’t call a tool that no longer exists. Rename a tool or change its schema, and every case that used it is now invalid, so a harness change means going through those cases by hand before you can re-run them.
Environment
The state the agent perceives when it starts acting isn’t the same as the tools it’s handed. Two runs with identical tools can still diverge if the world looks different: extra files an ls turns up in the working directory, a stale cache, a database left over from the last test run, a directory listing that happens to come back in a different order.
Build trust the way you would for an integration test. Pin the environment down so it mimics production: seed it and snapshot it before the run starts, then diff it after. Snapshot everything a tool might read without being asked, including working directory contents and ambient config. Give every run its own clean, isolated sandbox, so nothing one case writes can reach the next.
Nondeterminism
The same input, harness, and environment can still produce a different trajectory, even with temperature set to 0: tool retries and tool ordering can simply differ from one run to the next, and that variance lands before any metric sees the result. One run is one sample from the agent.
Treat it the way you’d treat any noisy measurement: run each case k times. That leaves you with k scores where you used to have one, so the case needs an aggregation strategy before it can report a single number, and there are two worth knowing.
Two measures are particularly important:
- pass@k asks whether at least one of the k runs succeeded. It tells you what the agent can achieve on a good day — an upper bound, not a reliable release criterion.
- pass^k requires all k runs to succeed. This is the stricter measure to use when you need evidence that the behavior is dependable.
3. Verify behavior, not just the final answer
In classical ML, verification usually meant checking whether the prediction matched a label or whether a score was close enough to a reference. For generative models, it often means comparing the output text with a ground truth.
An agent can also have ground truth, including ground truth for the path it took. But creating it requires subject-matter expertise, and it rarely takes the form of one canonical answer. Several different trajectories may be equally correct. The practical alternative is to describe what acceptable behavior looks like and what the agent must achieve.
Agentic eval has to score the agent’s entire behavior instead of a single output, and there’s no equivalent reference to check it against: a resolution that might take five turns, three tool calls, and a fallback to get there. So now you have to care about far more than the final answer. A 100% completion rate says nothing on its own about how the agent got there. The agent can call the wrong tool and self-correct, retry blindly until something sticks, or wander through a detour that happens to land on the right final state anyway. Trusting the agent means checking the path it took:
- Tool choice. Did it pick the right tool for the request, called with arguments that actually matched the request, not just arguments that returned a plausible-looking response?
- Order and structure. Where sequence matters, did the calls happen in the required order, and does the trace contain the spans a correct run should contain?
- Failure handling. When a tool times out, returns a malformed response, or hands back bad data, does the agent surface it and fall back, or improvise its way around the gap?
- Cost and latency. Turns, tokens, latency, dollars per run. An agent can be correct and still too expensive or too slow to put in front of a customer. Report it with the score.
Three runs of the same prompt, three different paths:
So guard the eval case with several checks instead of one score, and write each one against something that should happen: this tool was called, with arguments that came from the conversation rather than from thin air, before that one, and the whole thing finished inside your time and cost threshold. Each check is narrow on purpose, so a run that lands on the right answer the wrong way fails at least one of them.
What you don’t do is pin the whole path. Depending on the use case, the agent may need room to explore, and two runs that reach the same place by different routes can both be fine. Assert the steps that must be there and the boundaries it must stay inside; leave the rest of the trajectory free.
4. Make the metric earn your trust
A score earns trust when it behaves predictably: full marks mean the job was done perfectly, zero means it was not done, and better work receives a higher score. In agentic systems, the challenge is that this number often comes from an LLM judge rather than a simple function.
The same properties apply to classical and agentic metrics: they should be bounded, and monotone. These are not new requirements. Boundedness and monotonicity are established principles in metric design and measurement theory (Amigó & Mizzaro, 2020). They matter just as much when the score is produced by an LLM.
- Bounded. The score lives in a fixed range, so the same number means the same thing every time you read it, and two runs are comparable. An unbounded count – “17 tool calls wasted” – doesn’t tell you whether that’s near-perfect or a disaster.
- Monotone. The score is order-preserving with respect to quality, in a fixed direction – for an error rate the value never rises when the run gets better. Practically: given two runs, the one with the better score has to be the better run. Where that relation breaks, the number is untrustworthy.
Prove each one before you trust a number: which cases actually produce 100%, which produce 0%, and confirmation that the ranking in between holds. That part hasn’t changed. At deepsense.ai we went one step further, because an average doesn’t give you the whole picture. An agent that returns a solid 70% on every run and an agent that swings between 0%, 30%, 70% and 100% can average out to the same place — but in production you need the first one. Reliability is the property you’re buying, so we built it into the metric rather than checking for it afterwards. It’s called the Business Utility score – it penalizes high variance across runs instead of averaging it away, and the write-up is worth a read if you’re designing your own.
What changed is how you get there. §3 covered the first half: the object being scored moved from a response to a task. The second half is the scoring mechanism itself. Because an agent can reach a good outcome by many different paths, there’s rarely one canonical output left to exact-match against, so rubrics, state checks, and LLM judges take over from exact match. That’s convenient and dangerous in the same move: the judge is now the thing that has to carry boundedness, anchoring, and monotonicity, and an LLM judge doesn’t earn them just by sounding confident.
The same care applies to how many metrics you keep. Three scores you have validated and know how to act on are worth more than fifteen nobody has checked against the properties above. If a number moving wouldn’t change what you do next, it is costing attention and returning nothing. That is separate from the narrow checks feeding each score, which should stay numerous. It’s the headline numbers you report and gate releases on that should be few.
The LLM judge needs its own evaluation
LLM judges usually aren’t reliable enough to be monotone. Ask for a score from 1 to 10 and the ordering won’t hold, mostly because nobody can say what separates a 6 from a 7. So stick to binary checks: was the refund issued, did the answer cite the document it claims to. Those are easier to apply and verify, and they’re bounded and anchored already. You still get granularity: count how many checks passed.
How do you know the judge itself is trustworthy? Evaluate it on a held-out set it had no part in calibrating, labeled by people, with known passes and known fails. Measure agreement with Cohen’s κ. If it doesn’t match the human labels, don’t trust its scores. This is also why it’s better to have many small judges than one big one: a judge that checks a single thing is easier to verify.
5. Worked example: how an evaluation set de-risked a refactor
The framework above was applied end to end in an engagement involving a technical-support agent used in a supply-chain operation.
A customer writes in that the page won’t load and the connection is slow. The agent loads its ticket-analysis skill and runs a set of commands to pull that user’s data out of several internal systems. The tasks are narrow enough that the right tool calls are known before the run, so the checks are deterministic ones on the trajectory, like whether the skill was loaded and whether the commands were called.
Every case came out of production. Wherever the agent went wrong, everything up to that point stayed as recorded, and only the failing step was rewritten to what it should have been. That case went into the capability set, the things the agent can’t do yet. Once the fix landed and the case passed, it moved into regression. That set grew to around a thousand cases.

Later in the engagement, the system was restructured, a planned piece of architectural work that changed how the code was organized without changing anything the agent was meant to do. A refactor on that scale usually costs weeks of manual verification before anyone will sign off on it. The regression set replaced that. It ran against the restructured system and came back at the same level as before, so the migration went to production as soon as the code review was done, and the client took on no risk to get it there.
6. Summary: from a trustworthy metric to a release decision
A trustworthy metric should help the team make a clear deployment decision, not simply produce a score. Before treating an evaluation result as evidence for release, check whether the dataset, execution setup, verification process, judges, and metrics are reliable enough to support that decision.
Release checklist Before you ship, you should be able to answer yes to each of these:
- Coverage. The dataset covers every capability you claim the agent has, and the cases that came from production outnumber the ones you generated.
- Same harness. The harness that produced the score is the harness you’re shipping – same tools, same schemas, same permission scope.
- Clean environment. Every case ran in its own seeded, isolated environment, and you diffed the state afterwards.
- Repeats. Each case ran a few times and the gate was pass^k, not one green run.
- Path, not outcome. The checks assert how the agent got there, not just where it landed – tool choice, order, failure handling.
- Judges validated. Every judge in the loop has been scored against human labels on a held-out set it had no part in calibrating.
- Readable metrics. – You know what each number has to reach to be shippable, and what you’d fix to move it
None of this makes an agent safe to deploy. It makes the decision to deploy one an informed decision – you know what you checked, you know what you didn’t, and you know which of those the business is carrying.
That’s the whole difference between evidence and a green dashboard.e..
Table of contents





