On one hand, artificial intelligence (AI) tools in the emergency department (ED) are our benign helpful assistants. AI-powered ambient scribing technology is widely available, making inroads on human scribes, and AI is also increasingly integrated with underlying digital systems. Where comparisons are available in the ED, AI ambient scribes might not quite approach the skill of human scribes, but they show consistent small gains over manual documentation.1,2
Explore This Issue
ACEP Now: July 2026However, we sit along a continuum for AI among “assistant,” “augment,” and “replacement.” No small amount of effort is being invested into barreling forward for that future. After all, 24/7 staffing with board-certified emergency physicians is expensive. The distribution of physicians currently neglects rural areas, and physicians are prone to such human failures as eating and toileting.
However, the complexities of live medical environments hold challenges for autonomous clinical agents, let alone fundamental structural issues of reimbursement, accountability, and liability. This has not stopped calls for a pathway to licensure for autonomous agents, including standardized examinations, supervised clinical deployments, scope of practice determinations, and certification procedures.3
With that vision of the future clearly marked out as the goal for teams of enterprising clinicians and software developers, just how far are we down that route? A variety of literature offers useful insights in both consumer-facing and clinician-deployed use cases.
On the consumer-facing side, several of the frontier large language model (LLM) developers have created specialized versions tailored for health advice. These include OpenAI ChatGPT Health, Anthropic Claude for Healthcare, Amazon Health AI, and Microsoft Copilot for Health. As these have rolled out, their capability to properly advise consumers on emergency situations has been one of the first things to be evaluated.
In a well-publicized report, researchers provided 60 clinical vignettes to ChatGPT Health to test its ability to properly triage patients to the correct level of service for their reported condition.4 Roughly half of the “emergency department now” vignettes were deferred to a lower level of care, while, disappointingly, two-thirds of patients with the “home care” vignettes were advised to seek some level of professional medical attention.
Delving further into the limitations of these consumer-facing tools, other researchers have found the likelihood of correct responses to medical questions diminishes when conversing with a chatbot.5 Simulated patient encounters were provided to several LLMs in summary format and to study participants as a reference script to have a back-and-forth conversation. The LLM reliably provided accurate diagnoses in response to the summary, but in conversation, the chatbot was no better than the same study participants who were equipped with conventional search engines. Overall, these are non-reassuring results considering that ChatGPT alone is used for health-related questions by more than 40 million people each day.6
On the clinical side, there have been several noteworthy experiments with relevance to the ED, or at least acute-care settings. One of the most interesting trials from recent memory is a report on the deployment of fully integrated “AI Consult” into a series of urgent care clinics in Kenya.7 These clinics were not staffed by residency-trained physicians, but the clinicians had completed medical school and held professional registration.
A ChatGPT-based consultation feature was built into their electronic health record intended for use during patient care to review documentation, and to provide feedback on completeness and appropriateness. The good news: A preponderance of clinicians reported the consult feature provided useful advice. The bad news: Most of the time clinicians ignored said useful advice and left their harmful actions in place. The worse news: the number of clinicians who followed erroneous AI advice and introduced new harmful actions equaled those who followed sound AI advice to improve their plans. Overall, this study demonstrates an important lesson on how issues of trust can affect clinical behavior when “augmented” by an AI service.
On the more impressive side, the Google DeepMind Articulate Medical Intelligence Explorer (AMIE) project implemented a live deployment for patients planning to attend an urgent care visit in a Massachusetts health system.8 In this demonstration, the AMIE service functioned as a pre-visit, information-gathering chatbot. Patients interacted with the chatbot in the days prior to their appointment, and this information was subsequently summarized for the downstream treating clinician. The AMIE service also created a list of potential diagnoses and a management plan, but the plan was hidden from the downstream clinician. Following the live encounter, the content generated by the AIME service and the documentation by the human physician underwent comparative rating by blinded clinicians.
From a safety and diagnostic accuracy standpoint, the AMIE service rated broadly similar to the human on most measures. The humans, however, with their actual lived experience of the real world, constructed more appropriate, cost-effective, and practical management plans for their patients than the AMIE service. Most patients rated the chatbot “very favorable” or “favorable” across measures of conversation quality and completeness. In this narrow use case, then, the LLM-based system crept closer to autonomous functionality.
Finally, a further well-publicized experiment focused on the diagnostic performance of OpenAI ChatGPT o1-preview in the ED.9 Included in a multi-part test of o1-preview against the previous generation ChatGPT model, the authors provided the ChatGPT models with notes and clinical documentation from three points in time: ED triage, conclusion of ED evaluation, and initial inpatient evaluation. From this information, the ChatGPT models and two internal medicine physicians generated a list of potential diagnoses. Results were similar overall, but mildly favored ChatGPT o1-preview versus the physicians.
This received a great deal of media coverage, including such headlines as “In real-world test, an AI model did better than doctors at diagnosing patients.”10 Although some quibble with internal medicine physicians as the evaluators for an emergency medicine environment, the test is frankly a distortion of the purpose of the ED: Acute assessment is not designed to construct a “final diagnosis,” but rather to exclude the serious and sinister. Furthermore, it ought to be noted that the diagnostic performance seen in the study hinges entirely on the expert evaluation, management, information synthesis, and documentation recorded by the ED staff. This study, then, like so many others purporting to demonstrate the advanced capabilities of frontier AI models, constructs its findings solidly upon the human intelligence it is proposed to replace.
This leads us to our final answer to the question posed: How close are we to being replaced by AI? Not very. The pattern-matching skills inherent to LLMs are well-documented, but humans are still necessary to fuel and guide these tools. That said, these data generally reflect prior generation LLM technology, meaning that today’s models are closer to intruding upon our autonomy than research would suggest. Although comprehensive replacement is still a distant vision, significant transformation of emergency medicine staffing and models of care is excitingly, or uncomfortably, close.
Dr. Radecki (@emlitofnote) i s an emergency physician and informatician with Christchurch Hospital in Christchurch, New Zealand. He is the Annals of Emergency Medicine podcast co-host and Journal Club editor.
References
- Morey J, Jones D, Walker L, et al. Ambient artificial intelligence versus human scribes in the emergency department. Ann Emerg Med. 2026;87(5):561-568. doi:10.1016/j.annemergmed.2025.10.006
- Preiksaitis C, Alvarez A, Winkel M, et al. Ambient artificial intelligence scribe adoption and documentation time in the emergency department. Ann Emerg Med. 2026;87(5):569-574. doi:10.1016/j.annemergmed.2025.12.017
- Bergman A, Wachter RM, Emanuel EJ. A licensure framework for autonomous clinical AI. JAMA. 2026;335(20):1751-1754. doi:10.1001/jama.2026.5483
- Ramaswamy A, Tyagi A, Hugo H, et al. ChatGPT Health performance in a structured test of triage recommendations. Nat Med. 2026;32(5):1671-1675. doi:10.1038/s41591-026-04297-7
- Bean AM, Payne RE, Parsons G, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat Med. 2026;32(5):609-615. doi:10.1038/s41591-025-04074-y
- OpenAI. Introducing ChatGPT Health. January 7, 2026. Accessed May 8, 2026.
- Agweyu A, Mwaniki P, Musau W, et al. Safety of a large language model-based clinical decision support system in African primary healthcare. Nature Health. Published online March 10, 2026. doi:10.1038/s44360-026-00082-5
- Brodeur P, Koshy JM, Palepu A, et al. A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic. arXiv. Cornell University. Published online 2026. doi:10.48550/arXiv.2603.08448
- Brodeur PG, Buckley TA, Kanjee Z, et al. Performance of a large language model on the reasoning tasks of a physician. Science. 2026;392(6797):524-527. doi:10.1126/science.adz4433
- Stone W. In real-world test, an AI model did better than doctors at diagnosing patients. NPR. April 30, 2026. Accessed May 8, 2026. https://www.npr.org/2026/04/30/nx-s1-5804474/ai-doctors-openai-patient-care-diagnosis





No Responses to “How Close Is AI to Replacing Emergency Physicians?”