Engineering
What an AI agent needs before it goes live.
The engineering decisions that let a team trust an agent with everyday work, including the days something goes wrong.
Imagine an agent that prepares a service visit from an incoming maintenance request. In a demo, it finds the equipment, identifies the fault, and drafts the work order. Before it runs in the business, you need to know what happens when the equipment record is missing, the request arrives twice, or the scheduling system stops responding.
Test the job you are giving it
Start with a written description of what the agent may receive and what it must produce. For the maintenance request, a usable draft might require the equipment identifier, the reported symptoms, the relevant source material, and any unresolved questions. A fluent paragraph alone is not a finished work order.
Build a test set from examples the team is allowed to use. Include ordinary requests, ambiguous equipment names, conflicting instructions, and requests with no matching record. For each, define the acceptable result and when the system should ask for help.
Evaluate fields and actions separately. Correctly identifying the equipment does not make an invented repair instruction acceptable. A total score can hide that distinction; review failures by type and by consequence.
Make permissions part of the software
Reading a maintenance history, drafting a work order, and dispatching a technician are different capabilities. Give the system only the capabilities it needs for the current release. Enforce those limits in the application and the tools it can call.
For a draft-only agent, the scheduling integration should not expose a dispatch action. If dispatch is added later, require the relevant checks and approval before the application executes it. A sentence in the prompt is useful guidance; the permission boundary belongs in code.
Keep incoming content separate from operating instructions. An attachment may contain text that asks the agent to ignore its task or reveal other records. Treat that content as material to inspect, and check any proposed action against the user's access and the application's policy.
Use the responsibilities in our human-review guide to decide which actions need approval.
Figure 01System design
Enforce permissions where actions happen
Handle the failure after the write
One awkward failure happens when a system saves a work order but the connection drops before it returns confirmation. Retrying the same creation request can produce a duplicate. The agent needs a way to distinguish “failed to save” from “saved, but the response was lost.”
Use a stable operation identifier and the destination system's deduplication support where available. If a write's outcome is uncertain, check for the existing result before issuing another write. If that check is impossible, route the uncertainty for review instead of guessing.
Treat a repeated incoming request similarly. Store enough state to connect it to the work already performed. Resuming an interrupted job should not mean repeating every external action it took earlier.
Figure 02Failure recovery
A missing response does not mean a failed save
Give the team a way to operate it
Someone needs to be able to answer a practical question: what happened to this request? Keep a trace of the inputs used, source references, tool calls, checks, and outcomes. Record only the sensitive content needed for diagnosis, restrict access, and decide how long those records should remain.
Before release, walk through an operational failure with the team. The exercise should answer:
- Where does an incomplete job appear, and who will see it?
- Can a person finish the task without restarting the entire process?
- What stops repeated tool calls or an unexpectedly long run?
- How is automatic processing paused while an issue is investigated?
- Which version of the model, instructions, and reference material produced the result?
Include ongoing operation when scoping and pricing the AI project.
Expand from a release you understand
Run the first release with a bounded set of requests and an explicit review step. Compare the agent's work with the existing process, including the time required to inspect and correct it. Track unresolved jobs as carefully as accepted ones.
When the instructions, model, tools, or reference material change, rerun the relevant tests before widening access. Keep a previous working configuration available and a clear procedure for returning to it.
Production readiness means the team can use the system, see where it failed, and recover without improvising. Those capabilities deserve a place in the build plan alongside the agent's core task.
Have a workflow in mind?
Tell us what you want to improve. We’ll help you work out where to start.
Book a call