The proliferation of artificial intelligence (AI) throughout the healthcare landscape is proceeding rapidly, with tremendous potential to revolutionize medical education and patient care. The transition from deep learning algorithms to generative AI tools, specifically large language models (LLMs), within the last decade has led to a frantic effort to better understand the underlying technology and its limitations, develop robust regulatory frameworks, and determine their effects on the training and development of physicians.1,2 Generative AI tools are already being deployed to perform diverse tasks, both patient- and physician-facing, with the promise of improving workflow efficiency and augmenting clinical decision-making.3

However, the integration of these tools remains fragmented and institution-specific as each entity grapples with how to best integrate AI into various tasks and data sources.4 At the individual level, adoption and usage of AI is prevalent despite the absence of uniform organizational governance frameworks, which may result in misuse or harmful outcomes.5–7 In addition to health system efforts, there are initiatives underway to develop strategies to supervise learner usage such as the DEFT-AI framework, given the concern about impact on medical education.8 Around the country, medical schools and residency training programs are trying to integrate AI training into curricula but actual implementation remains limited.9 Insufficient faculty knowledge regarding AI tools is a large barrier, and educators and learners are both simultaneously grappling with the same issues on how to build AI literacy and utilize LLMs judiciously.

The educational stakes for medical trainees are particularly high as they are vulnerable to “deskilling,” “mis-skilling,” or “never-skilling” given the ease with which AI can be utilized for cognitive off-loading.8 AI tools are prone to automation bias given the fluency and confidence with which they provide responses, and this can be particularly impactful for less-experienced clinicians who may internalize inaccurate or biased information.10 Learners develop their clinical acumen over time and usage of AI shapes this longitudinal development, with long-lasting ramifications beyond any individual patient encounter for which they might be utilizing generative AI. As efforts continue to crystallize AI governance at institutional levels and to develop baseline AI literacy across clinicians of all experience levels, having a practical point-of-care cognitive tool for clinician-educators and trainees is critical. Supervising learner usage of AI tools involves teaching them to identify appropriate use-cases for AI, curbing reflexive usage purely for ease and convenience, and modeling how to critically appraise outputs. There is no standardized evaluation framework for AI outputs that is currently widely adopted, and many of the current appraisal tools (e.g. APPRAISE-AI, ABCDEFG framework) being introduced are intended to evaluate AI research studies, not real-time clinical outputs.11,12

AI Timeouts

We propose “AI Timeouts” (Figure 1) – a pair of structured and deliberately sequenced cognitive pauses intended to safeguard against reflexively reaching for AI tools without thoughtful intention, and the tacit acceptance of outputs without careful analysis. The AI Intent Timeout is a pause that is done before engaging with AI, and the Appraisal Timeout is done after an AI output is obtained but before any further action is taken. This cognitive framework is designed specifically for AI-driven clinical decision support systems, which are intended to augment the clinical decision-making process by serving as a rapid cognitive aid at the point-of-care.13 The intention of these systems is to equip clinicians with the ability to rapidly analyze data, scour medical literature, and offer tailored responses to clinical queries. However, without appropriate recognition of their pitfalls and limitations, users run the risk of making clinical decisions on unsound information.14 Thus, users need a deliberate approach to decide when to utilize these tools, and then subsequently determine whether the outputs are trustworthy and useful.

Figure 1
Figure 1.AI timeouts for clinician educators for intent and appraisal

Several existing frameworks address AI use in clinical practice, such as DEFT-AI and APPRAISE-AI. The suggested conceptual framework can be complementary to the existing tools at hand as they are deployed and used in a slightly different way, as described in Table 1.

Table 1.Comparison of existing AI frameworks with the AI Timeout
Framework Core Purpose Trigger Operator
AI Timeout Metacognitive pause to prevent reflexive AI use and appraise AI output Before and after AI use Individual or educator + learner
DEFT-AI Educator-led debriefing of a learner’s clinical reasoning with AI After AI use is identified Educator + learner
APPRAISE-AI Critical appraisal of AI clinical-decision support studies During appraisal of an AI study Researchers

Intent Timeout

The Intent Timeout is intended to address the core question of: “Should I use AI for this?” and includes prompts related to privacy considerations and data stewardship as well as the potential impact on clinical reasoning from cognitive offloading. It first prompts the user to consider whether AI is the correct tool for the task, guarding against impulsive use of AI tools purely for convenience. If deemed appropriate, the clinician is then asked to consider any potential concerns related to privacy or data stewardship. This is to prevent identifiable patient information or sensitive data from being input into AI tools inadvertently unless they are explicitly HIPAA compliant and have appropriate institutional agreements in place. The final component of the Intent Timeout asks the clinician to assess the potential consequences of choosing to utilize AI in that moment, specifically the impact on clinical reasoning skills.

Appraisal Timeout

The Appraisal Timeout tackles the issue of critical appraisal of outputs: “Should I trust this output?”. AI tools are known for their sycophancy15,16 and fluent confidence, generating a “veneer of veracity”17 that leaves clinicians at risk of making clinical decisions based on flawed reasoning. This timeout first asks to consider the fit of output, namely whether the output is a lengthy, unfocused response that is too general, or whether it appropriately targets the situation at hand. It then prompts consideration of whether the output has included the rationale for how a particular conclusion was reached and whether that rationale is truly logical. This forces the clinician to utilize their own clinical judgment and expertise, yet another layer to protect against de-skilling or never-skilling. Lastly, the priority consideration impresses upon the clinician to thoughtfully apply the relevant information from the output in an actionable manner.

Theoretical Foundation

The AI Timeouts are inspired by the pre-procedural/pre-surgical timeout, a standardized practice which formalizes a deliberate pause prior to a consequential action.18,19 Our framework is deliberately centered around AI usage in the clinical reasoning and medical education domains and is also based off cognitive forcing strategies which help to avoid diagnostic error and other biases.20 Concepts such as “the medical pause,” which categorizes pauses into two phases (a decision-making and an executive phase) and the pursuit of “endpoint diagnoses” are examples of metacognitive approaches aimed toward systematizing cognitive steps in a more concrete way.21,22 By developing a shared approach and vocabulary, it provides flexibility in how it can be used; the AI Timeouts can be applied by a singular clinician at the bedside, an attending supervising junior learners, or an institution wanting to implement a standardized approach to using AI. This type of framework does not require deep AI literacy or expertise and only requires thoughtful intention on the part of the clinician. It is also anchored in concepts that are already recognizable to all clinicians: at some point during medical training, clinicians will have encountered a pre-surgical timeout and will also have been taught about specific clinical reasoning skills to safeguard against cognitive biases.

Use Cases for the AI Timeout

The addition of any framework into the clinical environment potentially introduces friction by creating checklist-fatigue. To demonstrate how the AI Timeouts function as rapid heuristics, we outline three common scenarios to illustrate how the AI Timeouts can protect independent reasoning, model critical appraisal of output, and complement the DEFT-AI framework (Table 2).

Table 2.Various use cases for the AI Timeouts and how they can complement the DEFT-AI framework
Scenario Intent Timeout Appraisal Timeout How DEFT-AI can be used
Medical student pre-rounding on a patient with new hyponatremia The student attempts to create their own differential before using AI to broaden the list The student notices several rare diagnoses suggested and is unsure of their likelihood in this hyponatremic patient. The student asks the attending on rounds about how they are prioritizing the differential diagnoses. The attending notices the broad differential, asks if AI was used, and then uses the DEFT-AI framework to debrief how the student reconciled the AI output
Intern and resident managing sepsis due to urinary tract infection in a patient with multiple antibiotic allergies Given cross-reactivity of antibiotics, the dyad agrees that AI can suggest antibiotic options based on existing literature that might be outside of the typical alternatives The output is verified with pharmacy to confirm the risk of any potential cross-reactivity The attending probes about the antibiotic choice, asks if AI was used, and then uses the DEFT-AI framework to debrief on how the trainees handled the AI output
Attending assessing an elderly patient in clinic with polypharmacy and new-onset confusion AI can suggest obscure drug-drug interactions outside of the typical ones that might be typically identified by the attending AI flags a minor drug-drug interaction, and the attending uses their clinical judgment to determine the risk-benefit of holding particular medications N/A

Limitations

As it currently stands, this is purely a conceptual framework, and no empirical data exists regarding its uptake, fidelity, and impact on learning outcomes. While the barrier to use is likely low, it still requires an inherent degree of clinical expertise to execute effectively. The framework is intended to function as a heuristic, taking seconds to use as opposed to a cumbersome checklist. Given that there is limited bandwidth in a busy clinical workflow, using this framework for higher-stakes clinical tasks might be the most pragmatic approach. Furthermore, the Appraisal Timeout is scaled with clinical experience; as experienced clinicians will more readily recognize when an output might be outdated, biased, or incorrect, whereas less experienced trainees may require senior guidance to evaluate the output.

Conclusion

The rapid adoption of generative AI in the absence of guidance from governance structures leaves individual clinicians and trainees to navigate this new tool on their own. The AI Timeouts can be immediately deployed on rounds and at the bedside. This framework will safeguard against reflexive AI use and uncritical trust in outputs while simultaneously allowing attendings to model clinical reasoning for their trainees. The framework encourages purposeful pause, regardless of how much the capabilities of AI evolves. After all, the care of patients is not in the hands of a tool but in the hands of a clinician exercising clinical judgment.


Disclosures/Conflicts of Interest

Declaration of Generative AI and AI-Assisted Technologies in the Writing Process

During the preparation of this work, the authors used Google Gemini to format tables and brainstorm the organization of clinical use cases. Claude (Sonnett 5, Anthropic) was also utilized to assist with figure generation and manuscript outlining and editing for grammatical clarity. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the final content and the integrity of the publication. The authors have no conflicts of interest to disclose.

Corresponding author

Satya Patel, MD, FACP
Associate Clinical Professor,
David Geffen School of Medicine at University of California, Los Angeles
Hospitalist, Greater Los Angeles Veterans Affairs Health care System
11301 Wilshire Blvd Bld 500 Mail Code 111 Los Angeles, CA 90073
E-mail: satya.patel2@va.gov