A busy AI pilot can still be a bad AI pilot.
The team may have hundreds of prompts, a growing user count, positive comments, and a slide claiming that employees saved 200 hours. None of that proves the agent improved the business. It proves people used it and liked parts of it.
That is useful feedback. It is not a production decision.
An AI agent belongs in production when it improves a defined workflow without creating unacceptable errors, hidden review work, or operational risk. If your pilot scorecard cannot show that, you are measuring activity instead of value.
Model accuracy, token count, and self-reported time savings can help diagnose a pilot. They rarely show whether work finished faster or decisions improved.
A useful scorecard connects the agent to the workflow it is supposed to change.
Start with the decision the scorecard must support
Every pilot should end with one of three decisions: move to controlled production, revise and test again, or stop.
Most pilot reports avoid that choice. They summarize usage, collect a few quotes, and recommend "continued exploration." Okay, cool. What are you approving? What will it cost? What risk are you accepting? Which workflow gets changed?
Write the decision at the top of the scorecard before the pilot begins. Then choose evidence that helps leadership make it.
If the agent handles invoice exceptions, the decision may be whether it can prepare a resolution and route it for approval. If it handles inbound calls, the decision may be whether it can book routine appointments while transferring sensitive or unusual calls to a person. If it updates the CRM, the decision may be whether it can write to a limited set of fields without increasing bad data.
Keep the decision narrow. A pilot that tries to prove "AI works for our company" is almost guaranteed to produce a vague answer.
Metric 1: Did the business outcome improve?
Choose one primary outcome tied to the workflow. It should be something the business already cares about, not a new AI-only metric.
For a support workflow, that could be time to resolution or the share of tickets resolved correctly on first contact. For accounts receivable, it could be days to resolve an invoice dispute. For lead intake, it could be the time from inquiry to a correctly routed next step.
Measure the baseline before the agent touches the workflow. Use the same definition during the pilot. If the team changes the metric after seeing the result, the comparison is no longer clean.
Self-reported hours saved can sit beside the outcome, but it should not replace it. People are bad at estimating how long fragmented work took before automation. Even when the time estimate is accurate, freed capacity does not automatically become revenue or lower cost. Leadership still has to decide what that capacity will do.
Metric 2: Was the work correct?
Speed without quality is a fast way to create rework.
Track the percentage of outputs accepted without correction, the percentage needing minor edits, and the percentage that should never have reached the user or system. Define those categories in advance. "Looks good" is not a quality standard.
The definition should match the consequence. A slightly awkward internal summary is different from an incorrect customer commitment. A missing CRM field is different from changing a financial record. Use a severity scale that reflects the actual business exposure.
Sample the agent's successes, not only the flagged failures. If reviewers inspect only the work the agent escalated, silent mistakes can hide inside the completion rate.
Also count rework. An agent that drafts something in 30 seconds but takes a manager ten minutes to verify may have moved the labor instead of removing it.
Metric 3: How often did the agent finish the right work?
Teams often celebrate autonomy rate, which is the share of tasks completed without human involvement. That can be useful, but higher is not always better.
An agent should escalate when information is missing, policy is unclear, a customer asks for an exception, or an action falls outside its authority. Those escalations are evidence that the control is working.
Separate good escalations from avoidable ones. A good escalation protects the business. An avoidable escalation happens because the instructions, data, or integration failed. The second group is where the team should improve the system.
Track task completion, correct escalation, false completion, and abandoned work. That gives you a much more honest picture than one big automation percentage.
Metric 4: Can the system run when the demo team leaves?
A pilot usually has people watching it closely. Production does not get that luxury forever.
Measure integration failures, unavailable dependencies, retries, duplicate actions, and time to recover. Record whether the system preserved enough context for a person to take over. Check whether logs show what the agent saw, what it decided, which tool it used, and whether the action succeeded.
The NIST Generative AI Profile recommends recording human oversight roles, known issues, underlying model versions, and access modes in the system inventory. It also calls for protocols that let an organization deactivate a generative AI system when necessary. That is practical operating guidance. If nobody knows which version is running, who owns the exceptions, or how to stop the agent, the pilot is not ready for production.
Reliability is not only model quality. The CRM API can time out. Authentication can expire. A vendor can change a field. A queue can process the same event twice. Your scorecard needs to show how the whole system behaves, because the business experiences the whole system.
Metric 5: Did the controls work?
Do not list controls and assume they worked. Test them.
Try records the agent should not access. Give it a request that requires approval. Remove a required field. Present conflicting instructions. Confirm that it stops, escalates, and records the reason.
Track permission violations, attempted out-of-scope actions, missed approval gates, audit gaps, and the time required to investigate an incident. One severe control failure may outweigh a month of clean routine transactions. Your go or no-go rule should say that before testing starts.
This is also why a pilot should begin with narrow permissions. Read access before write access. Draft before send. One system before five. The agent can earn more authority with evidence. It should not receive broad authority because the demo is easier to build that way.
Metric 6: What did a successful outcome actually cost?
License fees and model tokens are only part of the bill.
Add integration work, workflow design, testing, human review, exception handling, monitoring, support, and ongoing changes. Include the internal people who maintain prompts, rules, data mappings, and approval paths. Then divide the full operating cost by successful business outcomes, not by prompts or agent runs.
That gives you a useful comparison against the current workflow. It also exposes a common pilot problem: the automation looks cheap because several employees are quietly keeping it alive.
Use the broader AI ROI framework to account for implementation and ongoing operating costs. For this scorecard, keep the unit simple. Cost per correctly completed case is harder to hide behind than a broad productivity estimate.
Run the pilot in stages
Start in shadow mode when the workflow allows it. Let the agent produce a recommendation without taking the action. Compare its output with what the team did and review the disagreements.
Next, allow a limited action with human approval. Keep the user group, data, and integration scope tight. Expand only after the error, escalation, reliability, and control evidence meets the thresholds you set.
Do not change five things at once between stages. If you switch the model, rewrite the workflow, add two integrations, and broaden the user group, you will not know what caused the result to change.
If you have not selected the workflow yet, use the seven-point workflow readiness test first. Measurement cannot rescue a process that was a poor automation candidate from the start.
The production scorecard
A decision-ready scorecard should fit on one page before the supporting detail. Include:
Primary outcome: Baseline, pilot result, target, and measurement period.
Quality: Accepted work, corrected work, severe errors, and review time.
Completion: Correct completions, good escalations, avoidable escalations, false completions, and abandoned tasks.
Reliability: Integration failures, duplicate actions, recovery time, and audit completeness.
Controls: Approval-gate performance, permission tests, out-of-scope attempts, and stop procedure.
Economics: Full pilot cost, expected production cost, human oversight cost, and cost per successful outcome.
Decision: Go, revise, or stop, with a named owner and the next evidence required.
Set the thresholds before the pilot. A team can reasonably accept different error rates for an internal draft and a customer-facing financial action. The important part is that leadership makes the risk choice deliberately instead of letting enthusiasm make it after the fact.
Stop rewarding motion
A pilot with heavy usage can still fail. A pilot with modest usage can succeed if it proves that one valuable workflow is safer, faster, or less expensive to run.
Count the prompts if you want. Track adoption. Ask employees what they think. Just do not confuse those signals with the decision evidence.
Measure the outcome. Measure the mistakes. Measure the human work hiding around the agent. Then decide whether the system has earned a place in production.