When AI starts calling patients directly, who's making sure it's safe?
October 6, 2026
Ambient clinical documentation has been revolutionary, fundamentally changing how clinicians work and transforming healthcare through human-AI partnership. But a larger revolution is already in progress.
The next wave of AI in healthcare isn't about helping clinicians document faster or think more clearly. It is about AI acting autonomously, making decisions and contacting patients directly, without a clinician filtering or approving every interaction in real time.
It is about to place outbound calls. It is about to ask a patient to check their blood pressure. It is about to screen a stroke survivor for red-flag symptoms. It is about to talk to a 15-year-old about their insulin. And it is about to navigate, in real time, what that teenager's mom is and isn't allowed to know.
This is patient-facing agentic AI: a fundamentally different category of AI applications with risks that exist in multiple dimensions compared to a word error in a note waiting for a clinician to catch. An autonomous voice agent can reach a patient and act before any human has the chance to intervene.
But there is no guidance built specifically for that scenario. Existing frameworks, including guidance from the Joint Commission and the Coalition for Health AI (CHAI), address AI broadly. None of them spoke directly to agents that talk to patients themselves.
So a group of us decided to build that guidance. We wanted to extend existing AI governance work by taking frameworks like NIST's AI Risk Management Framework and the Joint Commission/CHAI guidance and operationalizing them for the one thing they didn't yet cover: agents that speak with patients directly.
How the taskforce came together
In early 2025, we convened the Healthcare Risk Management for Voice AI Taskforce, an informal, volunteer group with no single institutional home. What made it useful was who was in the room: practicing clinicians alongside health system AI and informatics leaders, a payer executive, industry medical leads from AI companies, and a patient advocate. Multiple academic institutions and professional constituencies with one shared concern: healthcare organizations are about to adopt these tools faster than anyone was building the guardrails for them.
Those of us from healthcare AI companies joined because we could see patient-facing agentic AI emerging in the industry and felt a responsibility to help establish governance frameworks proactively, before these tools proliferated. While Suki is not deploying patient-facing voice agents ourselves, we understand agent-based systems well enough to know that trust in healthcare AI, especially when patients can't see a clinician's face or know they're talking to a machine, depends on safety and transparency being built in from the start, not retrofitted afterward. As a practicing clinician, researcher, and health systems expert at Suki, it was natural for me to participate in this task force and help develop broad guidance.
Step one: ground the work in real use cases, not hypotheticals
Before we could talk about risk, we needed to agree on what these agents actually do. The clinical members of the group each drew on their own practice settings (payer operations, public health, inpatient and outpatient medicine, acute and chronic care) to independently draft use cases that reflected how voice AI is likely to be deployed today, not a speculative future state.
We came up with five specific ways agents could be used in direct patient interactions:
- Outpatient hypertension management. An agent calls patients to walk through a home blood pressure reading and routes the result to the care team.
- Post-surgical follow-up. A 72-hour post-discharge check-in that triages patients by pain, wound healing, and diet tolerance.
- Post-stroke care coordination. A safety net for medication fills, INR monitoring, and red-flag symptom screening in a population especially prone to aphasia and cognitive impairment.
- HEDIS gap closure. Outreach at scale to close preventive-care gaps for health plan members.
- Pediatric diabetes management. An adolescent case layered with minor-consent law, split custody, and confidential medication protections.
We picked these five because together they stretch across acuity, population, and care setting enough to surface the principles that recur. Not because they're the only use cases that matter.
We then defined our terms deliberately, borrowing from patient safety literature: risk is an event or condition that can cause harm or degrade care quality when an agent acts without immediate human oversight; safety is the state achieved when those risks are mitigated to an acceptable level; mitigation is the strategy that gets you there. And we anchored everything to the patient's longitudinal journey: a single call is never really a single call. It's one moment in an ongoing course of care.
Step two: find the patterns across the use cases
With five concrete scenarios in hand, we went looking for what repeated. Four themes emerged immediately, each tied to multiple risks where things can go wrong:
- Agent-level risks. Transcription errors, hallucinations, omissions, social bias, missed escalations, and misuse or prompt injection. This is the software itself: the thing doing the listening, reasoning, and speaking.
- Data-level risks. Identity mis-verification, EHR data errors, fragmented records, model drift, and data over-exposure. Even a well-built agent produces unreliable outputs if what it's working from is wrong, incomplete, or out of date.
- Patient-level risks. Access and equity gaps, clinical inappropriateness, poor symptom recognition, and the over-triage/under-triage trade-off. Not every patient, and not every condition, is well suited to an AI-mediated encounter. How do we need to think of when it is appropriate to use an agent and when it's not, and how we determine if the clinical scenario changes from one to the other?
- Clinician-level risks. Alert fatigue, automation complacency and deskilling, liability ambiguity, workflow disruption, and moral distress. The agent doesn't operate in a vacuum; someone has to trust it, act on it, and answer for it. The answer today is to put the work on the clinicians, but that comes with its own challenges.
For each of the 21 risk categories that came out of those four themes, we wrote specific, evidence-grounded mitigations. Not general reassurances. Things like phonetic-alphabet confirmation for names and dosages, knowledge-base-only responses for restricted clinical topics, dual-factor identity verification, calibrated triage thresholds with a low under-triage tolerance, and explicit clinician override authority at every decision point.
That qualitative framework became our first paper: "Safety of Patient-Facing Agentic AI: A Consensus Framework for Risk Assessment and Mitigation," under review at JMIR.
Step three: don't just assert the risks, measure them
A framework built by the people who wrote it is a reasonable starting point, but it's not evidence. So we went further and put the whole thing to a formal test: a two-round modified Delphi study, following the RAND/UCLA Appropriateness Method, and broadened to 14 participants, reaching out to additional health system, informatics and global health leads. Everyone rated each of the 21 risks against each of the 5 use cases (105 risk-by-use-case combinations) on an ISO 31000 likelihood-and-impact matrix.
Round 1 alone produced 2,940 individual ratings and reached consensus on 95 of 105 cells. Round 2 resolved the rest. We ended with dual consensus on 95 of 105 cells (90.5%) after Round 1 and all 105 (100%) after Round 2.
Two findings stood out enough that they changed how we'd advise a health system to prioritize:
- Clinician-level risks ranked highest. Not agent-level technical failures like hallucination or transcription error, which get most of the attention in the broader AI safety conversation. Alert fatigue, liability ambiguity, and workflow disruption ranked above them. The panel's collective judgment is that the harder problem isn't the model. It's how the model's output lands on the human who has to act on it. This is the urgent place to start for agentic AI governance.
- Risk isn't uniform across use cases. Hypertension monitoring and post-stroke follow-up were rated meaningfully riskier than HEDIS gap closure. That's a direct argument against one-size-fits-all governance: a high-acuity deployment needs different pre-launch investment than a preventive-outreach program does.
That empirical validation became our second paper, a companion piece built specifically to test and ultimately support the framework from the first.
Step four: translate it for the people who have to act on it
A peer-reviewed framework and a Delphi study are exactly what a scientific audience needs. They're not what a hospital CMO evaluating a vendor contract next week needs. So we distilled the whole body of work into a practical white paper for healthcare leaders: providers, payors, pharmacies, and pharma organizations alike.
It boils the four-theme framework down into a four-layer risk matrix leaders can incorporate into a due-diligence checklist. Before signing with any vendor, you can walk through the agent, data, patient, and clinician risks and define a specific mitigation for each.
It also lays out the operational imperatives that matter most going into a launch: allocate resources by use-case acuity rather than uniformly, validate through pre-deployment adversarial red-teaming, define which decisions the agent can make on its own versus which require human review before reaching the patient, disclose the AI clearly and guarantee a path to a human, treat outbound calls as regulated communications under the FCC's 2024 TCPA ruling, shift from periodic to continuous monitoring once live, and establish liability and accountability before you ever go live, not after something goes wrong.
For vendors building agentic AI systems, this means embedding safety mechanisms into the agent itself: not relying on healthcare organizations to solve the problem alone, but designing and building solutions for identity verification, triage calibration, and clinician override directly into the platform, so they can build something truly meaningful for clinical care.
Why we're telling this story
None of this happened because a health system asked us to write it, and none of it happened inside a single institution. It happened because a group of clinicians, technologists, and researchers looked at how fast agentic AI was moving into direct patient contact and decided the guardrails needed to exist before the incidents did. Not after.
Adoption of patient-facing voice AI is still early. That is the opportunity in front of every healthcare organization right now: the window to build safety in is still open. We hope this body of work (the framework, the Delphi validation behind it, and the practical checklist distilled from both) sets out a specific, answerable set of questions for the next voice AI tool health systems consider. Instead of a general sense of unease about a technology that can talk to patients or blind trust that the agent will do the right things, we have a way to think about safety from the beginning.
Read the full body of work:
- Safety of Patient-Facing Agentic AI: A Consensus Framework for Risk Assessment and Mitigation: the qualitative framework paper (JMIR, submitted)
- Safety Risk Prioritization for Patient-Facing Agentic Voice AI in Clinical Care: A Two-Round Modified Delphi Study: the companion empirical validation (JMIR, submitted)
- Voice AI Safety: A Practical Framework for Healthcare Leaders: the white paper distilling both into an operational guide
This work reflects the collective effort of the Healthcare Risk Management for Voice AI Taskforce, an informal multidisciplinary group of clinicians, technologists, and researchers from health systems, academic medicine, health plans, and industry. Suki is proud to be a part of this initiative and contribute to safety standards for the field.

