The proliferation of artificial intelligence (AI) throughout the healthcare landscape is proceeding rapidly, with tremendous potential to revolutionize medical education and patient care. The transition from deep learning algorithms to generative AI tools, specifically large language models (LLMs), within the last decade has led to a frantic effort to better understand the underlying technology and its limitations, develop robust regulatory frameworks, and determine their effects on the training and development of physicians.1,2 Generative AI tools are already being deployed to perform diverse tasks, both patient- and physician-facing, with the promise of improving workflow efficiency and augmenting clinical decision-making.3
However, the integration of these tools remains fragmented and institution-specific as each entity grapples with how to best integrate AI into various tasks and data sources.4 At the individual level, adoption and usage of AI is prevalent despite the absence of uniform organizational governance frameworks, which may result in misuse or harmful outcomes.5–7 In addition to health system efforts, there are initiatives underway to develop strategies to supervise learner usage such as the DEFT-AI framework, given the concern about impact on medical education.8 Around the country, medical schools and residency training programs are trying to integrate AI training into curricula but actual implementation remains limited.9 Insufficient faculty knowledge regarding AI tools is a large barrier, and educators and learners are both simultaneously grappling with the same issues on how to build AI literacy and utilize LLMs judiciously.
The educational stakes for medical trainees are particularly high as they are vulnerable to “deskilling,” “mis-skilling,” or “never-skilling” given the ease with which AI can be utilized for cognitive off-loading.8 AI tools are prone to automation bias given the fluency and confidence with which they provide responses, and this can be particularly impactful for less-experienced clinicians who may internalize inaccurate or biased information.10 Learners develop their clinical acumen over time and usage of AI shapes this longitudinal development, with long-lasting ramifications beyond any individual patient encounter for which they might be utilizing generative AI. As efforts continue to crystallize AI governance at institutional levels and to develop baseline AI literacy across clinicians of all experience levels, having a practical point-of-care cognitive tool for clinician-educators and trainees is critical. Supervising learner usage of AI tools involves teaching them to identify appropriate use-cases for AI, curbing reflexive usage purely for ease and convenience, and modeling how to critically appraise outputs. There is no standardized evaluation framework for AI outputs that is currently widely adopted, and many of the current appraisal tools (e.g. APPRAISE-AI, ABCDEFG framework) being introduced are intended to evaluate AI research studies, not real-time clinical outputs.11,12
AI Timeouts
We propose “AI Timeouts” (Figure 1) – a pair of structured and deliberately sequenced cognitive pauses intended to safeguard against reflexively reaching for AI tools without thoughtful intention, and the tacit acceptance of outputs without careful analysis. The AI Intent Timeout is a pause that is done before engaging with AI, and the Appraisal Timeout is done after an AI output is obtained but before any further action is taken. This cognitive framework is designed specifically for AI-driven clinical decision support systems, which are intended to augment the clinical decision-making process by serving as a rapid cognitive aid at the point-of-care.13 The intention of these systems is to equip clinicians with the ability to rapidly analyze data, scour medical literature, and offer tailored responses to clinical queries. However, without appropriate recognition of their pitfalls and limitations, users run the risk of making clinical decisions on unsound information.14 Thus, users need a deliberate approach to decide when to utilize these tools, and then subsequently determine whether the outputs are trustworthy and useful.
Several existing frameworks address AI use in clinical practice, such as DEFT-AI and APPRAISE-AI. The suggested conceptual framework can be complementary to the existing tools at hand as they are deployed and used in a slightly different way, as described in Table 1.
Intent Timeout
The Intent Timeout is intended to address the core question of: “Should I use AI for this?” and includes prompts related to privacy considerations and data stewardship as well as the potential impact on clinical reasoning from cognitive offloading. It first prompts the user to consider whether AI is the correct tool for the task, guarding against impulsive use of AI tools purely for convenience. If deemed appropriate, the clinician is then asked to consider any potential concerns related to privacy or data stewardship. This is to prevent identifiable patient information or sensitive data from being input into AI tools inadvertently unless they are explicitly HIPAA compliant and have appropriate institutional agreements in place. The final component of the Intent Timeout asks the clinician to assess the potential consequences of choosing to utilize AI in that moment, specifically the impact on clinical reasoning skills.
Appraisal Timeout
The Appraisal Timeout tackles the issue of critical appraisal of outputs: “Should I trust this output?”. AI tools are known for their sycophancy15,16 and fluent confidence, generating a “veneer of veracity”17 that leaves clinicians at risk of making clinical decisions based on flawed reasoning. This timeout first asks to consider the fit of output, namely whether the output is a lengthy, unfocused response that is too general, or whether it appropriately targets the situation at hand. It then prompts consideration of whether the output has included the rationale for how a particular conclusion was reached and whether that rationale is truly logical. This forces the clinician to utilize their own clinical judgment and expertise, yet another layer to protect against de-skilling or never-skilling. Lastly, the priority consideration impresses upon the clinician to thoughtfully apply the relevant information from the output in an actionable manner.
Theoretical Foundation
The AI Timeouts are inspired by the pre-procedural/pre-surgical timeout, a standardized practice which formalizes a deliberate pause prior to a consequential action.18,19 Our framework is deliberately centered around AI usage in the clinical reasoning and medical education domains and is also based off cognitive forcing strategies which help to avoid diagnostic error and other biases.20 Concepts such as “the medical pause,” which categorizes pauses into two phases (a decision-making and an executive phase) and the pursuit of “endpoint diagnoses” are examples of metacognitive approaches aimed toward systematizing cognitive steps in a more concrete way.21,22 By developing a shared approach and vocabulary, it provides flexibility in how it can be used; the AI Timeouts can be applied by a singular clinician at the bedside, an attending supervising junior learners, or an institution wanting to implement a standardized approach to using AI. This type of framework does not require deep AI literacy or expertise and only requires thoughtful intention on the part of the clinician. It is also anchored in concepts that are already recognizable to all clinicians: at some point during medical training, clinicians will have encountered a pre-surgical timeout and will also have been taught about specific clinical reasoning skills to safeguard against cognitive biases.
Use Cases for the AI Timeout
The addition of any framework into the clinical environment potentially introduces friction by creating checklist-fatigue. To demonstrate how the AI Timeouts function as rapid heuristics, we outline three common scenarios to illustrate how the AI Timeouts can protect independent reasoning, model critical appraisal of output, and complement the DEFT-AI framework (Table 2).
Limitations
As it currently stands, this is purely a conceptual framework, and no empirical data exists regarding its uptake, fidelity, and impact on learning outcomes. While the barrier to use is likely low, it still requires an inherent degree of clinical expertise to execute effectively. The framework is intended to function as a heuristic, taking seconds to use as opposed to a cumbersome checklist. Given that there is limited bandwidth in a busy clinical workflow, using this framework for higher-stakes clinical tasks might be the most pragmatic approach. Furthermore, the Appraisal Timeout is scaled with clinical experience; as experienced clinicians will more readily recognize when an output might be outdated, biased, or incorrect, whereas less experienced trainees may require senior guidance to evaluate the output.
Conclusion
The rapid adoption of generative AI in the absence of guidance from governance structures leaves individual clinicians and trainees to navigate this new tool on their own. The AI Timeouts can be immediately deployed on rounds and at the bedside. This framework will safeguard against reflexive AI use and uncritical trust in outputs while simultaneously allowing attendings to model clinical reasoning for their trainees. The framework encourages purposeful pause, regardless of how much the capabilities of AI evolves. After all, the care of patients is not in the hands of a tool but in the hands of a clinician exercising clinical judgment.
Disclosures/Conflicts of Interest
Declaration of Generative AI and AI-Assisted Technologies in the Writing Process
During the preparation of this work, the authors used Google Gemini to format tables and brainstorm the organization of clinical use cases. Claude (Sonnett 5, Anthropic) was also utilized to assist with figure generation and manuscript outlining and editing for grammatical clarity. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the final content and the integrity of the publication. The authors have no conflicts of interest to disclose.
Corresponding author
Satya Patel, MD, FACP
Associate Clinical Professor,
David Geffen School of Medicine at University of California, Los Angeles
Hospitalist, Greater Los Angeles Veterans Affairs Health care System
11301 Wilshire Blvd Bld 500 Mail Code 111 Los Angeles, CA 90073
E-mail: satya.patel2@va.gov
