Feed aggregator

Thinking and organising in systems: reframing the long problem of learning from incidents

Quality and Safety in Health Care Journal -

Learning from safety incidents is one of the most common and widespread improvement strategies in healthcare. It is also one of the most problematic. Healthcare systems around the world expend enormous time and effort investigating large numbers of incidents, writing reports and issuing recommendations and a wide array of policies, frameworks, tools and methods surround and support these efforts.1–3 To take just two examples: the English NHS now collects around 3 million patient safety incident reports each year4; and between just 2020 and 2023, England’s Maternity and Newborn Safety Investigation body conducted around 3000 incident investigations and issued over 4620 recommendations.5 The remarkable scale of these activities is increasingly matched by growing frustrations at the limited return on these investigative investments: patients continue to be harmed by the same types of incidents in the same ways, while investigations find the...

We need a new paradigm to think about generative AI

Quality and Safety in Health Care Journal -

Numerous studies have now shown that large language models (LLMs) have remarkable similarities to physician cognition in silico, showing both human-like—or even superhuman—performance on cognitive medical tasks as well as reflecting human biases.1–6 Reasoning models, which were publicly introduced in September 2024, have since become ubiquitous in commercial products such as ChatGPT (GPT-5, OpenAI) and Gemini (Gemini 3, Google). They use chain-of-thought processing during model inference, allowing them to decompose complex clinical scenarios into intermediate steps, verify logic, correct errors prior to providing output and drastically increase their performance on a variety of cognitive tasks. The recent piece by Wang and Redelmeier in this issue of BMJ Quality and Safety shows that these powerful models, similar to the base models that preceded them, continue to show human cognitive biases when tested on clinical vignettes.7

This...

Do patient safety incident investigations align with systems thinking? An analysis of contributing factors and recommendations

Quality and Safety in Health Care Journal -

Background

Globally, up to 17% of hospitalised people suffer a patient safety incident. Learning from adverse events through patient safety investigation is critical to prevention; however, their utility is still questioned. Two key investigation outputs include identifying contributing factors (CFs) and proposing recommendations to prevent future occurrences. Criticisms of current methods include incomplete analysis of CFs and weak incident prevention strategies. A proposed solution is systems thinking analysis, which recognises healthcare complexity. However, it is not clear whether such methods are being applied in practice.

Objective

This study aimed to assess current use of systems thinking-based strategies by examining a set of Australian patient safety incident investigations.

Methods

Investigations (n=300) from 56 different Australian health services were deductively analysed. Identified CFs were classified by healthcare system level using a framework combining Systems Engineering Initiative for Patient Safety (SEIPS) principles and AcciMap’s hierarchical structure. Recommendation sustainability and effectiveness were classified as weak, medium or strong using US Department of Veteran Affairs’ criteria.

Results

51% of incidents were issues with clinical processes and procedures. The investigations identified CFs that disproportionally focused on the people involved in those processes (n=677, 47%) rather than other system levels and as a consequence, most recommendations were of medium (n=665, 51%) and weak (n=560, 43%) strength. Notably, 10% of investigations lacked any CFs or recommendations.

Conclusion

The focus on individual actions highlighted that simple linear thinking persists in patient safety incident investigations. This study proposes five key areas of effective incident analysis and investigation: a sociotechnical focus; improved data collection techniques; investigative independence; the professionalisation of investigators; and the aggregation of data. Learning from incidents is key to maximising their preventative effectiveness, especially in an increasingly complex healthcare system.

Artificial intelligence chain-of-thought reasoning in nuanced medical scenarios: mitigation of cognitive biases through model intransigence

Quality and Safety in Health Care Journal -

Background

Artificial intelligence large language models (LLMs) are increasingly used to inform clinical decisions but sometimes exhibit human-like cognitive biases when facing nuanced medical choices.

Methods

We tested whether new chain-of-thought reasoning LLMs might mitigate cognitive biases observed in physicians. We presented medical scenarios (n=10) to models released by DeepSeek, OpenAI and Google. Each scenario was presented in two versions that differed according to a specific bias (eg, surgery framed in survival vs mortality statistics). Responses were categorised and the extent of bias was measured by the absolute discrepancy between responses to different versions of the same scenario. The extent of intransigence (also termed dogma or inflexibility) was measured by Shannon entropy. The extent of deviance in each scenario was measured by comparing the average model response to the average practicing physician response (n=2507).

Results

DeepSeek-R1 mitigated 6 out of 10 cognitive biases observed in practicing physicians by generating intransigent all-or-none responses. The four biases that persisted were post hoc fallacy (34% vs 0%, p<0.001), decoy effects (44% vs 5%, p<0.001), Occam’s razor fallacy (100% vs 0%, p<0.001) and hindsight bias (56% vs 0%, p<0.001). In every scenario, the average model response deviated substantially from the average response of practicing physicians (p<0.001 for all). Similar patterns of persistent specific biases, intransigent responses and substantial deviance from practicing physicians were also apparent in OpenAI and Google.

Conclusion

Some biases persist in chain-of-thought reasoning LLMs, and models tend to produce intransigent recommendations. These findings highlight the role of clinicians to think broadly, respect diversity and remain vigilant when interpreting chain-of-thought reasoning artificial intelligence LLMs in nuanced medical decisions for patients.

Impact of COVID-19 on incidence and trends of adverse events among hospitalised patients in Calgary, Canada: a retrospective chart review study

Quality and Safety in Health Care Journal -

Background

While the incidence of hospital adverse events appeared to be declining before 2019, the COVID-19 pandemic may have changed its course. This study aimed to evaluate adverse event incidence rates and trends during the pandemic and analyse differences in patient outcomes.

Methods

This retrospective electronic chart review included a random sample of adult patients admitted to four acute care hospitals in Calgary between 2017 and 2022. 18 adverse events and patient information were extracted. We calculated the observed and risk-standardised incidence rates of adverse events. Interrupted time series analysis was employed to determine the impact of COVID-19 on adverse events trends. Outcome differences were evaluated using mixed-effects logistic regression and negative binomial models.

Results

Among 10 673 patient admissions, 2310 adverse events were identified, resulting in an incidence rate of 21.64 (95% CI 20.77 to 22.54) per 100 patient admissions, or 26.85 (95% CI 25.77 to 27.97) per 1000 patient days. After adjusting for patient characteristics, seasonal variations and overall trends, the adverse event incidence rate increased by 14% (incidence rate ratio (IRR) 1.14, 95% CI 1.01 to 1.29) during the COVID-19 pandemic. In multivariable mixed-effects models, adverse events were associated with significantly longer hospital stays (IRR 3.13, 95% CI 2.97 to 3.30), increased odds of 30-day readmission (OR 1.4, 95% CI 1.17 to 1.68) and in-hospital death (OR 1.72, 95% CI 1.43 to 2.08).

Conclusion

The incidence of adverse events was high but relatively stable in acute healthcare settings before the COVID-19 pandemic and increased during the pandemic. Strengthening healthcare resilience and prioritising patient safety initiatives are crucial as we transition into the post-pandemic era.

Grand rounds in methodology: four key things to know about the reliability of measurement

Quality and Safety in Health Care Journal -

Reliability is the most reported measurement characteristic used as evidence of the quality of a measurement. However, researchers often miss opportunities to design studies in a way that allows reliability to be calculated and report a calculation that is not germane to the purpose of the measurement.

Reliability is a number that quantifies the ability of a measurement to distinguish between the members of a population with respect to a measured quantity. It is a simple function of the signal and noise in a measurement. In this paper, I review four key attributes. First, reliability has a single technical definition, but there are many statistics that purport to quantify reliability that do not fit this definition or obscure their relationship to it. This undermines the fundamental simplicity of the concept and its useful implications. Second, researchers sometimes do not appreciate that the relevant calculation of reliability changes with the purpose and conditions of measurement and then report the wrong number. Third, reliability is a summary measure with several components that may be as or more relevant to report than reliability. Fourth, reliability is specific to a population, for example, a patient satisfaction score that is highly reliable in one population could have abysmal reliability in a different population when using the same survey instrument.

Reliability is an important part of evaluating and improving the measurements that form the foundations of scientific research in healthcare quality and safety. Understanding these four key attributes of reliability will improve the description and use of healthcare quality and safety measurements.

Bridging diagnostic safety and mental health: a systematic review highlighting inequities in autism spectrum disorder diagnosis

Quality and Safety in Health Care Journal -

Introduction

There is increased recognition that diagnostic errors disproportionately affect marginalised and underserved patient populations in the USA. However, evidence on diagnostic inequities in mental disorders is sparse and not well integrated into the overall diagnostic safety literature.

Objective

We systematically reviewed and narratively synthesised evidence on inequities in diagnosis of mental disorders, guided by the Diagnostic Process Framework developed by The National Academies of Sciences, Engineering, and Medicine.

Methods

We conducted a systematic review and a narrative synthesis. Medline, Embase, PsycInfo and CINAHL were searched for studies published between 2015 and 2024. Studies were eligible if they reported on inequities in the diagnosis of mental disorders and applied a quantitative, qualitative or mixed-methods design. Studies had to be peer reviewed, US based and published in English. The Mixed-Methods Appraisal Tool was used for quality appraisal. Data were analysed with a descriptive intent, and inequities were mapped into the diagnostic process.

Results

20 studies of varying methodological quality were included. Though not the initial focus, autism spectrum disorder (ASD) emerged as the most studied mental disorder (n=17). Of the diagnostic errors identified, most fell into the category of delayed diagnosis. 11 factors emerged as contributors to diagnostic inequities. Limited health literacy among patients and caregivers was the leading cause of diagnostic error in symptom recognition. Insurance coverage issues delayed patient engagement with the healthcare system. Provider bias during clinical history-taking and interviewing was seen as a key cause of delays and misdiagnoses. Within diagnostic testing and interpretation, culturally inequivalent assessment measures might cause misdiagnosis and delayed diagnosis for Black/African American and Hispanic/Latino patients. The use of medical jargon and lack of qualified language interpreters during communicating the diagnosis were associated with diagnostic errors impacting patients with limited health literacy and low English language proficiency.

Conclusions

Diagnostic inequities in ASD and other mental disorders persist across US patient populations. Multiple factors such as parental health literacy, provider bias and limited access interact and impact the diagnostic process. Addressing these interconnected barriers is essential to ensure timely, accurate and equitable care.

PROSPERO registration number

CRD42024581271.

Pages

Subscribe to Medication Safety Officers Society- MSOS aggregator