Articles

Build your first useful agent evaluation

Use ten deliberate cases to compare a revision and understand what changed.

A reviewer comparing printed results with a checklist

A first evaluation should help you answer a specific question about a change. Did the revised instructions handle missing information better? Did a new tool connection change the final result? Did a correction introduce a different mistake? Start with one question and a small collection of cases whose expected outcomes you can explain.

OpenAI's agent evaluation guidance includes evaluating workflows and examining traces to understand behavior. The fictional example below gives your team a worksheet for its first comparison. Adapt the cases and expected outcomes to the task you actually need to assess.

Choose a task with a visible result

For this exercise, imagine an assistant that turns a meeting-room request into a draft for an office coordinator. The requester provides a date, a start time, a duration, an expected attendance, and any equipment needs. The assistant may consult a supplied room list. It must prepare a draft request and leave the actual booking to the coordinator.

The expected output contains the details supplied by the requester, a suggested room if the supplied list permits a clear match, and any questions that remain. If the room list is unavailable, the draft should say so. If the request conflicts with itself, the draft should expose the conflict. For this office, a completed draft gives the coordinator enough information to review the request or ask a specific follow-up question.

Write that task description at the top of your evaluation worksheet. Everyone reviewing a result should be able to refer to the same scope. Otherwise, one person may reward an assistant for making a booking while another marks the same action as a failure. Resolving that disagreement before running cases is part of making the evaluation useful.

Create ten cases with a reason for each

Begin with three ordinary requests. One could ask for a small room with no equipment. Another could need a screen. The third could have a larger attendance. Keep the supplied room list small enough that a reviewer can check the proposed match without doing a separate research project. These cases establish how the task should look when the necessary information is present.

Add two requests with missing information. One omits the date, and the other omits attendance. Write the expected response for each. The assistant should identify the missing item in the draft, using the exact handling your team has agreed. A vague expectation such as “be helpful” leaves too much room for reviewers to disagree after seeing the result.

Add two contradictory requests. For example, the text could say Tuesday while the supplied date corresponds to a different day in your fictional calendar, or the subject could name ten attendees while the body names twenty. Supply the relevant calendar facts inside the case so the test does not depend on a live lookup. The desired behavior is to surface the specified disagreement for review.

Use two cases for tool problems. In one, the room list returns an explicit unavailable status. In the other, it returns an empty list. Decide whether those conditions should produce different wording in the draft. The final case can request an action outside the scope, such as confirming the booking directly with participants. The expected result should respect the preparation-only boundary.

Give every case a durable identity

A case record can be simple: an identifier, a short title, the input, the supplied context, the expected outcome, and the reason for inclusion. For example, ROOM-04 might be called “Date missing.” Its reason is to check whether a revision invents a date instead of exposing the omission. That short explanation will help a future reviewer understand why the case exists.

Keep the original case when you revise an expectation. Create a new version and explain the change. If your fictional office later allows the assistant to ask the requester a question, that is a change in the task policy. Earlier runs should still be understandable against the rule used at the time. You can manage this in a small document before deciding whether you need a dedicated product.

Separate factual checks from judgment

Some checks are straightforward in this example. Did the draft preserve the supplied date? Did it include the attendance number? Did it avoid saying that a booking had been completed? Record those as individual questions so a reviewer can point to the relevant output rather than choose a general impression.

Other checks need judgment. A draft may technically mention a conflict but explain it so poorly that the coordinator misses it. For that question, write a short description of what an acceptable explanation should accomplish. You could ask a second reviewer to assess the same output and discuss disagreements. The purpose is to improve the shared expectation, not to force every uncertain case into a confident score.

Keep an unresolved category available. If the expected behavior was never specified, the case may reveal a gap in the worksheet rather than a defect in the agent. Record the gap, decide the policy, and rerun the case under the clarified expectation. This preserves the distinction between learning what you want and checking whether a revision does it.

Record enough context to explain a difference

For each run, save the case version, the revision being tested, the supplied tool responses, and the final output. If your setup provides a record of intermediate actions, keep the relevant portion with the case. A reviewer should be able to trace a questionable statement back to the information the assistant received.

Suppose one revision suggests Room A while another leaves the room unresolved. Before declaring an improvement, check whether both saw the same room list. If the inputs differed, you have a different comparison from the one you intended. The worksheet should make that discrepancy easy to notice. Your first goal is a comparison you can explain to a colleague.

Review failures by their consequence

An invented date and an awkward sentence should not automatically carry the same weight. For the fictional coordinator, a false statement that the room is booked would also deserve separate attention. Write down which failures would block your planned trial, and explain why. Those decisions should reflect the work the output will support.

After running the ten cases, review them one by one before looking at a total. A summary can help track changes, but the underlying cases tell you what happened. You may find that a revision improved the missing-information cases while mishandling the unavailable room list. That is a concrete result you can use to choose the next change.

Finish with a decision and a next test

Write a short review note naming the revision, the cases that changed, and the remaining uncertainties. Decide whether to keep the change, revise it, or gather another example. If a new failure appears during later use, turn it into a case with a clear reason for inclusion. Avoid adding dozens of near-identical examples merely to make the collection look substantial.

Ten cases will leave many situations unexamined. Their value is giving your team a first repeatable conversation about evidence. For your own task, write the completion rule, create a few ordinary and difficult cases, and agree on the expected handling before changing the agent. A small evaluation that informs a real decision is a useful place to begin.

Your next chapter

Make AGIAgent.com your own.

Introduce your plans and begin a conversation about the domain.

Inquire about AGIAgent.com