PRACTITIONER PERSPECTIVE: GOVERNING AGENTIC AI IN HEALTHCARE
Moving beyond approval committees to build deep organizational capacity for governing AI systems that act autonomously in healthcare.
PRACTITIONER PERSPECTIVE
GOVERNING AGENTIC AI IN HEALTHCARE
From Application Review to Rooted Governance
An evidence-informed perspective on applications, organizational readiness, and the wider healthcare system
Sonali Tamhankar, PhD | July 2026
AI governance is often pictured as a committee that can approve, restrict, monitor, or suspend an application. That function remains essential. But as AI moves from predicting and drafting toward planning, using tools, and initiating actions across workflows, the committee's decision becomes only one visible surface of governance, which must be supported through deep organizational roots of AI literacy, sandboxed AI experimentation, and cross-functional collaboration. The more tentacles AI grows, the more roots governance must grow to remain effective.
To maximize benefit and minimize harm in the era of agentic AI, governance means that everyone who designs, purchases, deploys, and uses health AI ensures that a bounded, observable, and revocable AI agency supports the people who deliver and enable care. Patients' agency, health, dignity, values, and choices remain the north star.
This document considers three progressively widening circles - individual AI applications; the organizations and workflows in which they operate; and the wider ecosystem of policy, payment, and healthcare infrastructure. Each circle can strengthen - or constrain - the others.
1 START WHERE YOU ARE
AI is already embedded in care delivery. Most devices on the FDA's non-comprehensive list of AI-enabled medical devices are in radiology, while generative AI is entering documentation, patient-message drafting, summarization, coding, and other clinical and administrative work. Early studies of ambient scribes report improvements in documentation burden and burnout, but effects vary by user, workflow, and outcome. In a national survey using 2023 data, 65% of hospitals reported using predictive models; among those hospitals, 79% used models supplied by their EHR developer, while 61% locally evaluated most or all models for accuracy and 44% for bias. [1-4] Also, AI use beyond expert-designed and validated applications is growing - internal review at a health system found roughly three times as many agents built in approved low-code or no-code tools as the use cases that had passed formal governance.[5]
A 90-DAY STARTING POINT: Healthcare organizations: Conduct a learning-oriented inventory of approved, embedded, home-grown, vendor, low-code/no-code, piloted, and informal AI uses; record the problem, population, data, workflow owner, decision affected, and escalation path. Developers: Test a transparency package using the CHAI Applied Model Card or a comparable framework with a customer governance team to register the product, assign controls, and simulate an update or incident; revise what proves unclear or unusable.
2 THINGS ARE CHANGING - FAST
Consider a hypothetical readmission-reduction tool. A current predictive version may flag a patient as high risk. A more agentic version could investigate risk drivers through explainable AI and RAG-based tools, and initiate processes within existing human-led clinical interventions. For example, if lack of transportation and missed PCP appointments is a risk driver, an agent could trigger social work referral, or bring in care navigation where comorbidities drive risk. Most of the AI technology needed to develop such a tool already exists. The potential value could be substantial if the intervention is matched to patient priorities and real service capacity. A "tentacled" AI solution, one that reaches across data sources, workflows, teams, and downstream actions, could create substantial safety, privacy, and cybersecurity risks. Today's governance should be stress-tested against plausible future workflows of this kind.
A 90-DAY STARTING POINT: Healthcare organizations: Choose one important care outcome or unresolved operational need, imagine what would be possible if the technology existed, and work with patients, IT, and frontline teams to consider benefit, capacity, security, fallback, and harm. Developers: Prototype a layered solution beginning with the smallest useful capability; add agentic or multi-agent components only where they address a defined need, and ensure each layer can be enabled, paused, or removed independently.
3 AI AS A BOUNDED WORKFLOW PARTICIPANT
Agentic systems may function less like passive tools and more like delegated participants in a workflow. This functional resemblance does not transfer moral or clinical responsibility to the system - or resolve the legal accountability of the people and institutions around it. These systems may observe, plan, call tools, communicate, and act across systems. That increases the importance of permissions, provenance, state tracking, privacy, cybersecurity, third-party dependencies, human verification, escalation, and safe failure. CHAI's Agentic AI guidance emphasizes defining an agent's operational scope, read/write guardrails, audit logging, data reconciliation, realistic workflow evaluation, and ongoing monitoring. [5]
A 90-DAY STARTING POINT: Healthcare organizations: Map what an agent may observe, recommend, initiate, communicate, and complete; specify who can stop it, when patients are informed and how they are involved in AI governance, how they reach a human or contest an error, and how work returns safely to people. Developers: Support plain-language disclosure, human escalation, correction, auditable preference handling, and documentation of action boundaries, logs, dependencies, and update effects.
4 AI SOLUTIONS SUCCEED OR FAIL INSIDE ORGANIZATIONAL ECOSYSTEMS
Application-level governance is necessary but insufficient - especially as AI acts across workflows. Organizations need distributed capacity to make, enact, observe, and revise sound AI decisions. Rooted governance includes foundational literacy across the workforce with role-specific AI capability, advanced internal experimentation or learning through shared resources, and cross-functional communication among technical teams, clinicians, operational staff, leaders, patients, and vendors. NIST places trained, empowered, multidisciplinary teams and a critical-thinking culture within AI risk management. Joint Commission-CHAI guidance likewise connects governance with system-wide safeguards, ongoing monitoring, and education and training. [6-8]
A 90-DAY STARTING POINT: Healthcare organizations: Start a role-aware literacy plan, identify and nurture internal or shared expertise, and strengthen communication between IT, research, clinical, operational, and executive domains. Developers: Specify organizational foundational AI literacy expected, and provide additional role-specific education that includes limitations, failure modes, and safe use. Provide a safe sandbox using representative scenarios so intended users can evaluate usefulness, usability, trust, and cognitive burden alongside technical accuracy.
5 HYBRID CARE CREATES NEW WORK
AI-enabled care requires more than a model and a clinician placed "in the loop." Organizations must assess potential benefit from AI for a specific problem, compare potential internal/external, open-source/vendor solutions, curate fit-for-purpose reference datasets, evaluate local performance, redesign workflows, investigate incidents, and decide when to pause or retire a system. This work may fall invisibly to nurses, pharmacists, social workers, navigators, analysts, informaticists, and IT staff. Adding a human to AI may not automatically create superior performance: outcomes depend on task design, expertise, and the division of labor. [9,10] The cost of not budgeting for this work is potential future liability and cybersecurity expenses, or a runaway AI expenditure without commensurate gains.
A 90-DAY STARTING POINT: Healthcare organizations: Map the labor, authority, decision rights, and protected time needed to realistically implement one AI workflow; map over- or under-reliance; do not silently transfer accountability to the last human in the chain. For high-risk use cases, examine whether contractual liability, indemnification, and vendor insurance coverage are proportionate to the risks borne by the healthcare organization. Developers: During one pilot, estimate customer-side implementation and oversight work, provide evaluation tools, and clarify what support is included before and after deployment.
6 LOCAL OPTIMIZATION MAY NOT ADD UP TO GLOBAL OPTIMIZATION
An application can save minutes yet add work elsewhere. It can improve throughput while worsening access, staff burden, or equity. The critical work of AI integration, evaluation, monitoring, incident response, and lifecycle management benefits many patients. In many fee-for-service arrangements, however, this population-level work falls outside a direct per-patient reimbursement pathway. Here organizational governance meets the wider payment environment: a rapid mixed-method evaluation of NHS chest-diagnostics deployment illustrates how resource-intensive AI procurement and implementation can be, while the U.S. Medicare Physician Fee Schedule was not designed for software-based clinical technologies.[11,12] Total cost of ownership must therefore include organizational labor and the cost of failure - not only a license.
A 90-DAY STARTING POINT: Healthcare organizations: build a lifecycle budget and benefit hypothesis for one use case, including workflow labor, monitoring, incident response, and retirement; set continuation and stop criteria. Developers: Separate successful workflow integration from ROI: identify the bottleneck the tool is intended to relieve and test whether it improves that bottleneck without shifting cost, delay, work, or risk elsewhere
7 MEASURE WHAT MATTERS - NOT MERELY WHAT IS EASY
Technical performance metrics like accuracy or positive predictive value are necessary, but they do not always ensure clinical or organizational success. Measures should follow the intended benefit: patient outcomes and goals, safety, access, equity, trust, staff workload, reliability, overrides, unresolved escalations, and recovery from failure. "Time saved" is especially incomplete unless we ask for whom, under what workload, and how that time is reallocated. Organizations can use an 80/20 heuristic to identify the small set of bottlenecks most connected to meaningful outcomes, and continue using older inexpensive solutions where still effective, reserving generative and agentic AI for truly complex use cases.
A 90-DAY STARTING POINT: Healthcare organizations: define success and unacceptable performance before deployment, identify the small set of bottlenecks most likely to improve outcomes that matter, and name what you will not automate. Developers: Agree with one organization on success measures beyond model performance, and provide the traceability and exportable data needed to evaluate them using local users, workflows, and populations.
8 KEEP THE TREE AND THE FOREST IN VIEW
One year from now, technical capabilities may have advanced remarkably while realized value remains uneven. AI will meet the healthcare system we actually have: fragmented data, constrained workforces, uneven digital capacity, reimbursement incentives, regulatory obligations, and services that may already be overextended. Even a sophisticated application can easily become an expensive burden rather than a source of value. Organizations need not be AI research leaders to possess the expertise that matters most: knowledge of their patients, constraints, and workflows.
A 90-DAY STARTING POINT: Healthcare organizations: review the AI portfolio at both use-case and system levels; identify duplication, displaced burden, inequity, and barriers that require policy or payment advocacy rather than another tool. Developers: Conduct a deal-breaker-first assessment in one intended deployment setting; test critical assumptions about data, interoperability, staffing, infrastructure, and service capacity, and state where the product fits, requires adaptation, or does not yet fit.
BUILD GOVERNANCE CAPACITY BEFORE DEPLOYMENT
The central question is not only what AI can do. It is whether the system around it can learn, judge, adapt, and respond with sufficient rigor and speed. Responsible AI frameworks from organizations like CHAI, NIST, URAC[13] and the Joint Commission provide guidance on accountability, data governance, technical controls, patient participation, incident response, and lifecycle monitoring. Our philosophy must evolve beyond a simple approval stamp to a holistic, rooted governance with:
- Foundational as well as role- and application-specific AI literacy that empowers employees to effectively partner in AI risk management and implementation,
- Internal or shared AI expertise with sandboxed learning and experimentation to critically evaluate expensive AI options, and
- Cross-functional communication to align AI solutions with true organizational bottlenecks.
The more tentacles AI grows, the more roots governance must grow to remain effective.
REFERENCES
[1] U.S. Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices. Updated 2026. Source
[2] Olson KD, et al. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout. JAMA Network Open. 2025;8(10):e2534976. doi:10.1001/jamanetworkopen.2025.34976. Source
[3] Garcia P, et al. Artificial Intelligence-Generated Draft Replies to Patient Inbox Messages. JAMA Network Open. 2024;7(3):e243201. doi:10.1001/jamanetworkopen.2024.3201. Source
[4] Nong P, et al. Current Use and Evaluation of Artificial Intelligence and Predictive Models in US Hospitals. Health Affairs. 2025;44(1). doi:10.1377/hlthaff.2024.00842. Source
[5] Coalition for Health AI. Agentic AI Responsible AI Content and Unified Testing & Evaluation Framework. 2026. Source
[6] National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). 2023. Source
[7] NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. 2024. Source
[8] Joint Commission and Coalition for Health AI. Responsible Use of AI in Healthcare. 2025. Source
[9] Liu P, et al. Human-AI Teaming in Healthcare: 1 + 1 > 2? npj Artificial Intelligence. 2025;1:47. doi:10.1038/s44387-025-00052-4. Source
[10] Vaccaro M, Almaatouq A, Malone TW. When Combinations of Humans and AI Are Useful: A Systematic Review and Meta-Analysis. Nature Human Behaviour. 2024;8:2293-2303. doi:10.1038/s41562-024-02024-1. Source
[11] Ramsay AIG, et al. Procurement and Early Deployment of Artificial Intelligence Tools for Chest Diagnostics in NHS Services in England: A Rapid Mixed-Method Evaluation. eClinicalMedicine. 2025;89:103481. doi:10.1016/j.eclinm.2025.103481. Source
[12] Longyear RL, Berenson RA. Artificial Intelligence Payment Policies: Challenges for CMS and the Medicare Physician Fee Schedule. Health Affairs. 2026;45(1):14-21. doi:10.1377/hlthaff.2025.00672. Source
[13] URAC. Health Care AI Accreditation. 2026. Source
Author's note: I contributed to CHAI's Agentic AI work group and the July 28, 2026 panel "The Governance Gap: Preparing Health Systems for Agentic AI." These experiences contributed to my learning, along with my healthcare experience. This independent perspective is written in my personal capacity; the views are my own and do not represent my employer or CHAI. AI was used to assist with literature review and wordsmithing purposes - all the ideas and errors are my own.