Strategy
Give your first AI agent a smaller job.
How to choose a useful starting point, define a finished result, and make the first deployment worth expanding.
“We want to use AI in operations” is a reasonable ambition. It is a difficult build brief. Before choosing a model or connecting any tools, choose a piece of work whose result the team can inspect. A small job with a clear finish gives you something concrete to improve.
Describe what finished looks like
Consider a purchasing team that receives supplier quotes by email. “Automate procurement” leaves almost every decision open. A more useful first brief is: turn the quotes for a purchase request into a comparison that includes price, delivery date, exclusions, and a link to each source.
That brief has a beginning, an output, and a person who can check it. It also tells the engineer what the system should do when a supplier leaves something out: flag the missing information. It should not fill the gap with a plausible answer.
Write the acceptance criteria beside the brief. Each price must retain its currency. Delivery dates must distinguish a promise from an estimate. A product substitution must be visible. These details turn “looks good” into a result the team can actually accept.
Figure 01Workflow design
From supplier quotes to a usable decision
Missing information stays visible for review.
Watch the work before estimating the value
Ask someone to complete a recent example using their usual tools. Notice which documents they open, where they look up missing context, and what makes them ask a colleague for help. A written procedure can miss the small judgments that determine whether the output is useful.
For the quote comparison, time spent reading attachments may be only part of the effort. The buyer may spend longer checking whether two suppliers have quoted the same item. If the agent extracts prices perfectly but overlooks product differences, it has made the comparison faster to produce and harder to trust.
Record a baseline that matches the proposed result: preparation time, corrections after review, and whether the comparison was usable. Keep waiting for a supplier response separate from active preparation time. Otherwise, the pilot may be judged against a delay it cannot change.
For a broader evaluation of the proposed investment, use our guide to assessing an AI project before building.
Choose a boundary you can test
For a first version, the system might read attachments and prepare the comparison, while the buyer chooses the supplier. That boundary makes the review specific. The buyer checks an artifact they already understand, and the team can learn which parts of its preparation are reliable.
Compare possible projects using questions that expose the practical work of delivery:
- Can we get representative inputs, including incomplete and awkward examples?
- Can the team agree on what a correct result contains?
- Does this task recur often enough for the improvement to matter?
- Can we identify an error before it creates an expensive downstream problem?
- Will someone use the result in their normal working day?
The review step deserves its own design. Here is how to plan human review for AI agents.
Figure 02Scope
Draw the boundary of the first release
Make the pilot useful in its own right
A pilot should produce something the team would want even if the project stopped there. In this example, that could be a source-linked comparison ready for review. It does not need automatic negotiation, supplier selection, and purchase-order creation to earn its place.
Put the result where the buyer already works. A comparison saved against the purchase request is easier to use than another chat window that requires copying the answer into a separate system. Include a way to mark a field as wrong and explain the correction. That feedback becomes material for the next round of testing.
Decide in advance what would justify extending the scope. Perhaps the comparisons are consistently accepted, the remaining corrections are well understood, and preparing them takes less effort overall. If those conditions do not hold, investigate before adding another capability.
A well-chosen first job gives you a reusable way to test, review, and improve AI inside the business. Start with a result people need, make its quality visible, and let the evidence guide the next build.
Have a workflow in mind?
Tell us what you want to improve. We’ll help you work out where to start.
Book a call