A CMIO’s view on deploying healthcare AI agents safely

8 min read
A smiling nurse in blue scrubs and a stethoscope shows a tablet to an older man in a light blue shirt and glasses as they sit together on a couch.
  • Healthcare leaders are under pressure to deploy member and patient-facing AI at pace, with no concession on safety or quality.
  • In the back office we measure accuracy. In front of a member or patient we measure harm. That distinction determines what must be evidenced before launch.
  • An agent reaches production, and remains there, only on top of a substantial apparatus of guardrails, escalation pathways, evaluations, and monitoring.
  • The organizations that scale agentic experiences will be those that build one safe foundation and have every agent inherit it.

Every organization I speak with faces the same constraint: costs are rising faster than revenue, and the workforce cannot meet demand. We know member and patient-facing AI can increase productivity without adding headcount, but few organizations have enough agents live and doing that work today.

Most teams are stalled on trust. Once AI moves from the back office into direct contact with an unwell, anxious member or patient, the safety threshold rises considerably. 

I recently hosted a session with Fierce Healthcare alongside Raj Kurup, VP of Engineering and Operations at Point32Health; Jason Wright, VP of Enterprise Business Platforms and Digital Solutions at Cone Health; and Jordan Christensen, SVP of Data and AI Engineering at League. We talked about what it takes to build safe, effective agentic experiences that can go live and stay live.

You can watch the full recording here, and below are the key takeaways from the discussion.

Healthcare-grade isn’t the same as accuracy

In the back office we measure accuracy. In clinical practice we measure harm. An agent that is correct in 99 cases out of 100 is not safe if the hundredth produces serious consequences. What matters is the distribution of the errors: the mechanism by which the agent could harm someone, the severity of that harm, its likelihood, and whether we would detect it. That is clinical hazard analysis, and it is a different discipline from performance testing. Healthcare-grade is a different construct from a higher accuracy score.

In my own practice I now see two kinds of patients: those who arrive with a printout of their ChatGPT conversation and ask whether it is correct, and those for whom the clinician remains the only trusted source. The Harris Poll surveyed 2,057 U.S. adults and found that 62% have used AI tools for medical information, while roughly a third do not trust the medical information those tools provide.¹ That division leaves us a narrow path: providing the immediate answers and support patients now expect, while preserving absolute clinical trust.

Cone Health is addressing this by closing the gap between what patients expect and what the system can deliver today: immediate answers, straightforward navigation, and less friction. What they have deliberately excluded is autonomous clinical decision-making around diagnosis and treatment. As Jason put it, “In healthcare, trust always comes before automation.”

Safety is the harder problem, but utility carries equal weight. An agent that clears every safety review and is not useful will send members straight back to the ungoverned public chatbot.

The iceberg beneath the agent

Modern tools make it easy for organizations to build impressive prototypes, but everything underneath it is where teams stall.

Jordan described this with an iceberg metaphor. Above the waterline sits what people see: the model and the conversation. Below it sits an infrastructure of more than a hundred subsystems that together make the agent safe for a member or patient to use, and the great majority of engineering time is spent there rather than in the conversation layer.

Sustaining that iceberg requires red teams whose sole remit is to break the model, compliance specialists validating against regional regulation, and clinicians confirming that answers meet the clinical standard. Tens of thousands of evaluations are required to evidence that an agent can be trusted in production. All of it must then secure approval from a governance board on which clinicians, compliance, legal, information security, privacy, and operations all sit.

I have sat on an NHS governance committee and on a vendor’s AI governance board. This is consistently the most difficult work to get right, because it never ends. Go-live is one stage. What follows is monitoring the agent in production, remaining accountable for its outputs, and converting each error into the next round of evaluations.

Most organizations struggle to build the iceberg once. The problem compounds when they rebuild it for every new use case. A benefits agent acquires its own escalation rules, guardrails, and governance; the scheduling agent then starts from nothing. We can see why 84% of AI engineering teams spend at least half their time on safety infrastructure.²

See the iceberg, already built

Forge by League is the foundation below the waterline. Every agent inherits built-in guardrails, evaluation frameworks, monitoring, and escalation paths from day one.

What should carry from one agent to the next

It does not need to work that way, and Jordan was specific about why. The guardrails, evaluation framework, escalation pathways, and connectors that bring governed data into an agent apply to every use case. They belong in a foundation that every agent sits on, not inside each agent. Raj made the same point: much of a care navigation build can carry into care gaps, which secures consistency and speed as Point32Health scales their experience.

The analogy I use is the hospital formulary. A hospital does not convene a new committee and devise a new evidence standard each time a drug arrives. It operates one process and scales the scrutiny to the risk. The organizations that scale their agentic experience will be those with a healthcare-grade foundation that every agent inherits. The fifth agent moves faster because the framework is reused, but the safety standard does not move.

Moving faster without compromise

I see this tension daily from both sides. Clinically, demand is rising and staff cannot keep pace. In the governance rooms where we sign off agents, the safety standard does not shift. Clinical trust is not tradable for speed.

When the safety layer is rebuilt for each agent, the cost is paid in full every time, which is how teams spend years getting a handful of agents live. Build a healthcare-grade foundation beneath all of them, and the second agent is faster than the first with no reduction in clinical trust.

Sources

  1. Merck Manual, More Than 3 in 5 Americans Use AI Tools for Medical Information, 2026.
  2. Sinch, 74% of enterprises have rolled back live AI customer communications agents, 2026.
WEBINAR

The fastest path to safe, live agents

Forge by League gives you the healthcare-grade foundation required to get agents into production. Your engineers build the experience. Safety is built in.

Intelligent care—in your inbox

Subscribe now to receive product updates and insights from League.