All guidesDeployment
A green notebook and working plans on a sunlit studio table.

Deploy the workflow, not just the agent.

An agent is a capability. A deployment is a commitment to complete a defined piece of work. Confuse the two and you can automate every conversation while leaving the operational problem untouched. The useful unit is the whole workflow: the trigger, the action, the verified outcome and the response when something goes wrong.

The unit of value is completed work

A model can produce a convincing appointment confirmation for very little money. It can do that even when there is no appointment. The important distinction is not between a convincing and an unconvincing voice. It is between a statement and a completed action.

In a clinic, a booking depends on the patient’s identity, visit type, professional, available slot and a successful write to the appointment system. A later cancellation must also change the reminder sequence. The patient experiences all of this as one service, even if different systems perform each step.

This changes the deployment brief. “Add a voice agent” describes a component. “Resolve appointment changes without leaving conflicting bookings or reminders” describes work with a testable outcome. It also makes the dependencies visible before patients discover them.

An agent that resolves its part of a conversation but leaves the record wrong has optimised the wrong boundary. Start with the complete task, then decide which parts need a model, deterministic software or a person.

Autonomy depends on the environment you build

The model needs context, tools and authority. These are separate design choices. Context tells it which appointment exists. A tool lets it request a change. Authority determines whether it may make that change and under what conditions. Giving the model more information does not settle the other two.

Keep consequential controls outside persuasion. A contact policy should prevent an out-of-hours action before it is sent. A booking tool should report whether the write succeeded. A payment state should come from the payment provider. The model should not be able to turn a confident sentence into evidence that any of these things happened.

State must survive interruptions. A callback agreed for Friday cannot disappear when the call ends. A clinic’s manual edit must supersede the old reminder. A request arriving by WhatsApp must not create a second lead simply because the first enquiry came by phone.

a16z’s analysis of enterprise AI describes the last mile as work on business context, integration and unpredictable user paths. In patient operations, those paths include a cancellation arriving during a retry, a changed treatment plan or a request that needs a clinician. The environment determines what the agent can safely do with each.

Reference: a16z: From Demos to Deals, Insights for Building in Enterprise AI

Build vs buy changes ownership, not the work required

Building can be the right choice for a clinic group with sustained engineering capacity, unusual workflows and control of its systems. The commitment extends beyond implementation: connectors need maintenance, model changes need testing and incidents need someone who can diagnose and correct them.

A managed deployment moves some of that responsibility to a specialist. It does not remove the clinic’s authority over care, data access and operating policy. Write down the boundary: who approves a rule, who implements it, who verifies it and who responds when it fails.

A hybrid can be effective. The clinic retains its records and approval rights while a deployment team operates agent workflows. Human follow-up may sit with the clinic, a supplied care team or both. Define coverage, escalation and unresolved-case ownership before deciding how much work the agent should handle.

The label matters less than the operating agreement. If everyone owns quality in principle but no one owns a failed reminder in practice, the deployment has a gap.

Assign ownership before the first patient contact.

ResponsibilityYour clinicsDelivery team
Care and contact policyApproves the rulesImplements and tests the rules
System accessAuthorises access and data useBuilds and maintains agreed connectors
Workflow qualityDefines acceptable outcomesTests changes and investigates failures
Human follow-upAgrees authority and coverageRoutes to clinic care or supplied care team
Clinical decisionsRetains clinical responsibilityEscalates; does not decide treatment
An illustrative division of responsibility for a managed deployment. Confirm the actual scope with your provider.

Design recovery before increasing autonomy

Suppose an appointment API creates a booking but times out before returning the response. Retrying blindly can create a duplicate. Announcing success can mislead the patient. The workflow needs a third path: reconcile the system state, then continue or hand the unresolved case to an owner.

That is an ambiguous outcome, not simply a failed call. Distinguish actions known to have succeeded, actions known to have failed and actions whose result is unknown. Each requires a different retry or recovery policy.

Test interruptions between actions, not just difficult phrases inside a conversation. A patient opts out while a call is queued. Reception reschedules while a reminder is being prepared. The care team closes while a transfer is pending. A fluent transcript cannot prove that the resulting state is correct.

Use test records and controlled failures before release. For each case, define what must be true in the connected systems, what the patient is told and who receives the work if the agent cannot finish it.

  • After an uncertain booking response: reconcile before retrying and verify that exactly one valid appointment exists.
  • After a reschedule: update the diary and invalidate reminders tied to the old appointment.
  • After an opt-out: apply the request across affected queued actions and campaigns.
  • After an unsuccessful human transfer: retain an owned, unresolved request rather than recording a completed handoff.

Lower cost per action can still produce a worse business case

An AI agent is not tokens plus a markup. Inference is one input to a working service, alongside integration, telephony, messaging, evaluation, maintenance and human handling. Optimising one input does not prove that the total cost per resolved task has improved.

Choose the endpoint before comparing costs. A reminder sent, a booking created, an attended visit and a signed treatment plan are different outcomes. If the objective is attendance, the denominator must reach attendance. If it is growth, estimate the visits added beyond the return rate you would have seen anyway.

Include implementation over a stated period and the work still required from clinic staff. A workflow with cheap automated calls but frequent manual repair can cost more than one with a higher inference bill and reliable completion.

Measure contribution as well as conversion. Additional appointments only create economic value if the clinic can serve them at a worthwhile margin. Do not value an unsigned plan as revenue or a signed plan as collected cash. The reactivation analysis linked below works through these distinctions with an editable business case.

Watch where the bottleneck moves

A successful deployment can expose a different constraint. Faster enquiry handling may create more bookings than the relevant professional can serve. Better identification of patient questions may send more work to a care team that was already full. More completed calls can coexist with a slower path to treatment.

Treat this as part of the deployment, not evidence that the agent should become more persistent. Review appointment availability, time to human response and unresolved cases alongside conversion. If those queues grow, increasing contact volume may make the service worse.

Expansion also changes the case mix. The next clinic may use different visit types, permissions or staffing. A broader campaign may reach patients with less immediate need. Do not multiply the first cohort’s result across the whole group without checking those differences.

The useful scaling question is whether the next cohort has the context, capacity and controls required to deliver the same standard. A higher call limit answers none of those questions.

Make the first deployment a decision, not a demonstration

Pick a bounded workflow with a named owner, a known starting state and an outcome the clinic can verify. Agree the eligible cohort, comparison period, contact boundaries, human cover and conditions that would stop the rollout.

Before launch, establish three kinds of evidence: the workflow completes the task, its failure paths behave as intended and the economics are worth testing. None substitutes for the others. A cheap system that leaves records wrong is not ready; a reliable system doing unnecessary work has not established value.

During the rollout, review outcomes and exceptions together. Turn appropriate care-team corrections into reviewed test cases. Re-run them when prompts, tools, models or policies change. Feedback should improve the deployment through a controlled release, not let a live agent rewrite its own authority.

Our work in applied AI since 2018, including through Menhir, repeatedly exposed the gap between predicting an action and making it happen. Closing that gap requires ownership of the surrounding operation. The first deployment should tell you whether to expand, redesign or stop, and give you the evidence for that decision.

Reference: Menhir: The Last Mile of AI

Sources & further reading

Examples and operating recommendations are Wilco’s editorial analysis. External sources support the claims linked in the text.

Bring us the workflow you are considering.

Map the systems, exceptions and responsibilities with the people who would help deploy it. You should leave knowing what needs to be built and what your team would still own.