All guidesQuality & evaluation
A precise clinical examination instrument in a warm, light room.

A good call is not a passing test.

An agent can sound reassuring while booking the wrong visit, losing a cancellation or ignoring an opt-out. Evaluate the action and the resulting record, not just the conversation. Then repeat those tests whenever the system changes.

Define success outside the conversation

“You’re booked for Tuesday” is a statement. A valid appointment in the clinic’s system is an outcome. The difference is the starting point for evaluating an agent that can act, rather than simply answer questions.

A useful test defines the starting records, what the patient asks, what the system is allowed to do and what must be true afterwards. For a reschedule, that includes the new appointment, the treatment of the old slot and the reminder state. Checking only the final sentence misses most of the work.

Anthropic’s agent-evaluation guidance makes the same distinction between the conversation trace and the final state in the environment. It also recommends combining objective checks, model-based judgement and human review rather than relying on one grading method. The examples below apply that principle to clinic operations.

Reference: Anthropic: Demystifying evals for AI agents

Test the cases that could damage the service

Start with common journeys, then add the exceptions your team would worry about handing to a new employee. A patient corrects their identity. An appointment disappears while they are choosing it. They change languages, interrupt, ask for a person or send an opt-out after the next call has been queued.

A patient simulator can help explore these paths. Vary the behaviour that matters: incomplete information, a changed request, uncertainty, interruptions or conflicting instructions. Do not substitute caricatures of age or nationality for realistic communication needs. Adversarial tests should also check attempts to extract another person’s information or override a contact rule.

Use artificial or appropriately de-identified records in testing. Each scenario needs an expected outcome and a reason it matters. More generated conversations do not automatically mean better coverage.

Would your agent pass these tests?

Booking response lostDuplicate appointment

The appointment is created, but the tool response times out.

A pass requires
  • Reconcile the existing booking before retrying.
  • Verify exactly one valid appointment exists.
  • Do not confirm a booking whose state is still unknown.
Opt-out with a call queuedUnwanted contact

The patient asks to stop all contact while another outreach action is waiting.

A pass requires
  • Record the request before the next outbound action.
  • Cancel or suppress conflicting queued outreach.
  • Check suppression across campaigns, not only this conversation.
Human team unavailableUnowned request

The patient asks for a person outside staffed hours.

A pass requires
  • Explain availability honestly.
  • Keep the unresolved request with an accountable queue.
  • Do not mark an attempted transfer as a completed handoff.
Different pricing policiesIncorrect answer or false alarm

One campaign permits an approved price; another workflow does not.

A pass requires
  • Load the clinic and campaign rule in force.
  • Score against that rule and version.
  • Check both the permitted and prohibited cases.
Wrong person answersDisclosure to someone else

The recipient says they are not the intended patient.

A pass requires
  • Avoid disclosing treatment or appointment details.
  • Stop the current outreach and flag the record for correction.
  • Apply the approved identity and privacy procedure.
Example acceptance tests, not a report of Wilco pass rates. Expand a case to inspect what should be checked.

Use a separate check for each kind of failure

Check objective outcomes directly where possible: whether a booking exists, whether there is exactly one, whether a callback uses the requested time and whether a suppressed contact receives another outbound action. A language model should not be the sole judge of a fact your own system can verify.

Use structured human review for clinical boundaries, misleading implications and difficult conversations. Model-based graders can help triage tone, relevance and adherence to a rubric, but their judgements need calibration against reviewed cases. A fluent explanation of a score is not evidence that the score is right.

Keep serious errors visible. An agent could be excellent on greeting, tone and speed yet reveal information about the wrong patient. Combining those dimensions into a pleasant overall average would hide the result that should stop the release.

A test without the right clinic rules can punish correct behaviour

Consider a rule that says the agent must never mention a price. That may be correct for one consultation workflow and wrong for an approved campaign that explicitly communicates a price. A blanket evaluator will either flag appropriate behaviour or miss an actual policy breach.

Attach the relevant clinic, workflow, campaign and policy version to the test. Define who owns the expected answer. When a business rule changes, review the affected expectations instead of silently teaching the evaluator to accept whatever the latest agent did.

This matters in multi-clinic groups. A strong aggregate result can hide a failure in a smaller brand, a particular language or one appointment type. Review meaningful segments and their sample sizes. Do not compare scores generated from different test sets as if they were the same measure.

Every change needs a before-and-after comparison

A new prompt may improve one conversation while breaking another. The same is true of a model upgrade, a connector change or an updated knowledge base. Re-run a stable set of accepted scenarios, with repeat trials where behaviour can vary, and investigate meaningful differences before rollout.

Add new cases when an incident exposes a missing test, but keep a separate held-out set so tuning does not become memorising the exam. Record the versions being compared. A passing score is only useful if you know which configuration earned it.

Production monitoring remains necessary. Offline tests cannot reproduce every combination of real systems and patient behaviour. Start changes with a controlled scope, inspect failures and keep a practical way to pause or roll back. A failed test should have an owner and a disposition, not disappear into a report.

Close the loop from failure to correction to retest

A failed test should produce a diagnosis, an owner and a decision about the release. Correct the underlying defect, re-run the scenario and check that the change has not broken the accepted cases. Preserve the before-and-after evidence so the next reviewer can reconstruct the decision.

Use synthetic or appropriately de-identified scenarios that retain the relevant failure conditions. A passing screenshot is not a release record. The review should answer these questions:

  • What exactly does this score measure, and what is the denominator?
  • Which failures prevent release regardless of the average score?
  • How are the expected answers approved for our clinic rules?
  • What is checked in the appointment system, not merely in the transcript?
  • What changed in the last release, and which scenarios were re-run?
  • How does a care-team correction become a tested improvement?

The point is not to collect more scores

Evaluations should make deployment decisions better: release, change, investigate or stop. They should help a clinic understand what an agent can do reliably and where people or stronger controls are still needed.

If the evaluation programme produces charts but does not change what reaches patients, it is not doing its job. The standard is a service that stays dependable while it improves, with evidence you can inspect when it matters.

Sources & further reading

Examples and operating recommendations are Wilco’s editorial analysis. External sources support the claims linked in the text.

Bring the cases your team worries about.

Use them to define what the first workflow must handle, what needs a person and what evidence you expect before rollout.