Operations

The real cost of running an AI agent.

Model usage is one line in the budget. A useful cost model follows the whole task through review, retries, and completion.

Model, tool, and human effort contribute to cost per accepted result.MODELTOOLSPEOPLEACCEPTED TASKSCOST PERACCEPTED RESULTCOUNT THE WHOLE TASK

An agent can be inexpensive to call and expensive to use. If its output needs substantial checking, or if a failed run repeats paid work, the model bill tells you very little about the cost of the result. To budget for an agent, start with the unit of work the business actually needs.

Choose the unit you want to improve

Suppose a team uses an agent to prepare a weekly supplier update. A model request is not the unit of value. The useful unit is an update the team has reviewed and accepted for the intended audience.

A single update may involve document retrieval, several model calls, a paid data lookup, and human checking. It may also involve a second attempt after the first draft fails a check. Attach those costs to the same task so they remain visible together.

For a defined period, divide the operating costs attributable to that workload by the number of accepted updates. Include spend on failed attempts in the numerator. Show failed and unfinished tasks alongside the result so a lower completion rate cannot disappear behind a neat average.

Figure 01Unit economics

Measure the cost of an accepted result

Cost peraccepted task
Model+Retrieval+Tools+Infrastructure+Human work
Accepted tasks

Count every attempt.
Failed runs still consume resources.

Count the human work.
Review, corrections & upkeep.

Use the same workload and time period for both sides of the fraction. Include failed attempts in the costs, and report unfinished tasks alongside the result.

Include the work left for people

Separate preparing the update from reviewing it. An agent may remove document collection while leaving the reviewer with more factual checks. Measure both activities before deciding that the system has saved time.

If you convert that time into money, state the labor-rate assumption. Time released for another task is useful capacity, but it does not automatically reduce payroll or become revenue. Report the operational improvement first and make the financial interpretation explicit.

Include ongoing work such as correcting reference material, investigating failures, and maintaining integrations. Keep initial build costs separate from recurring operating costs, then decide over what period you want to recover the initial investment. That makes the budget easier to revisit when the scope changes.

For the wider project budget, see how we scope and price an AI project.

Look for repeated work in the trace

Before changing models, inspect a few complete tasks. Does the agent fetch the same supplier profile more than once? Does a formatting error make it repeat the research? Does every section of the report receive the entire document collection? These are specific things an engineer can investigate.

A useful cost breakdown separates retrieval, model usage, paid tools, infrastructure, and review. It should also show how many attempts each accepted result required. Compare normal tasks with unusually expensive ones rather than assuming the average represents both.

Optimize one source of waste at a time and check the output against the same acceptance criteria. Saving an earlier research result may avoid a repeated lookup. Passing only relevant source passages may reduce unnecessary context. Either change still needs testing: an omitted detail can make a cheaper draft more expensive to correct.

Figure 02Execution trace

Retry the step that needs work

After reviewThe draft needs revision. The research is still valid.
Restart everything
Fetch againRepeated workRedraftCheck quality
Resume from saved work
Reuse sourcesAlready verifiedRedraftCheck quality

The acceptance criteria stay the same.

If only the draft needs revision, reuse valid research from the same task. Refresh it if the inputs or sources change, and apply the same quality checks.

Compare complete tasks when testing a model

A cheaper model is worth trying on a well-defined subtask, such as assigning an update to a category. Measure whether that substitution changes error rates, retries, review effort, or completion time. Use the same representative inputs for the comparison.

A more capable model may justify its cost on a difficult synthesis step if it reduces expensive rework. It may add no useful value to a straightforward extraction. The choice depends on the behavior you observe in that part of the workload.

Keep the comparison reproducible. Record the configuration, the input set, the criteria, and the result. An offline test provides a basis for a limited rollout; verify the effect on live work before treating the estimate as an operating saving.

Set a budget the system can enforce

Decide what the agent should do when a task exceeds its expected effort. It might stop with a partial draft and flag the remaining research for a person. The output should clearly identify what is unfinished.

Set limits on run duration, paid tool usage, and repeated attempts. Review alerts against the volume of work: an increase in total spend may simply reflect more accepted tasks, while a flat bill can hide a decline in completion.

A useful operating report puts cost per accepted task beside quality, completion rate, review time, and turnaround. Read those measures together. That is enough to make a concrete decision about what to improve, what to expand, and what still needs work.

Budget limits work alongside the recovery and operational controls in our production-readiness article.

Have a workflow in mind?

Tell us what you want to improve. We’ll help you work out where to start.

Book a call