What does the published evidence show about the accuracy, citation reliability, and clinical impact of AI large language model tools used to answer clinical questions at the point of care, including h

What does the published evidence show about the accuracy, citation reliability, and clinical impact of AI large language model tools used to answer clinical questions at the point of care, including hallucinated or unsupported references, guideline concordance, and effects on clinician decision-making?
BioSkepsis

Bottom line. Large language models exhibit substantial medical reasoning capabilities but remain prone to citation fabrication, guideline deviations, and prompt sensitivity, requiring retrieval-augmented generation and physician oversight for safe point-of-care use.

Diagnostic and Management Accuracy Across Clinical Specialties

Published evaluations demonstrate that large language model (LLM) performance varies substantially depending on whether the task involves static multiple-choice examinations, curated case challenges, or open-ended clinical workflows. When evaluated on standardized licensing benchmarks such as the USMLE (MedQA) and MultiMedQA, scaled models such as Flan-PaLM and Med-PaLM achieved benchmark accuracies of 67.6% and above, demonstrating significant knowledge encoding (PMID: 37438534). Similarly, in rheumatology case assessments, ChatGPT-4 provided the correct top diagnosis in 35% of cases compared to 39% for rheumatologists (p = 0.30), and listed the correct diagnosis among its top 3 in 60% compared to 55% for physicians (p = 0.38) (PMID: 37742280). In emergency department triage and diagnostic prediction, GPT-4 scored 1.76 out of 2 points compared to 1.59 points for resident physicians across internal medicine emergencies (p = 0.01) (PMID: 38976865).

However, when deployed in autonomous, multi-step clinical environments requiring active information gathering, diagnostic accuracy declines substantially. In an evaluation of 2,400 real patient encounters from the MIMIC-IV database across four acute abdominal conditions (appendicitis, cholecystitis, diverticulitis, and pancreatitis), leading open-access models (Llama 2 Chat, OASST, and WizardLM) achieved mean diagnostic accuracies of only 45.5% to 54.9% when required to autonomously request physical examinations, laboratory tests, and imaging (PMID: 38965432). When tested as second readers with all information provided upfront, these models achieved mean diagnostic accuracies of 58.8% to 67.8%, performing significantly worse than human hospitalists (p < 0.001), with a mean diagnostic gap of 16 to 25 percentage points (PMID: 38965432). Similarly, on an unbalanced real-world sample of 1,000 emergency department visits, GPT-4-turbo demonstrated lower accuracy than resident physicians for hospital admission recommendations (0.43–0.58 vs. 0.83) and radiology request status (0.74 vs. 0.79), despite exceeding physician accuracy on antibiotic prescribing status (0.82–0.83 vs. 0.78) (PMID: 39379357).

Bibliographic Citation Reliability, Fabrication, and Falsification

A prominent vulnerability across unaugmented generative LLMs is the spontaneous fabrication and corruption of bibliographic citations. In a systematic evaluation of 636 citations across 84 academic papers generated by GPT-3.5 and GPT-4, 55% of GPT-3.5 citations and 18% of GPT-4 citations were completely fabricated (PMID: 37679503). Fabricated citations frequently synthesized genuine author rosters, real journal names, and plausible article titles that did not correspond to any published manuscript (PMID: 37679503; PMID: 40206627). In an evaluation of 115 references in medical content generated by ChatGPT-3.5, 47% were fabricated, 46% were authentic but contained substantive errors, and only 7% were authentic and fully accurate (PMID: 37337480). Across evaluated reference elements, incorrect PubMed Identifier (PMID) numbers occurred in 93% of papers, while errors in volume numbers (64%), page ranges (64%), and publication years (60%) were common (PMID: 37337480).

Similar error rates have been reported across diverse clinical domains. In an analysis of ChatGPT responses to 20 medical questions evaluated against PubMed and journal repositories, 69% (41/59) of provided citations were fabricated, despite 95% listing authors with established publication records in the respective fields (PMID: 40206627). In a dedicated literature search experiment in psychiatry, ChatGPT provided 35 citations of which only 2 (6%) were real, while Google Bard generated 8 citations with 0% accuracy (PMID: 37499282). In radiology-specific information retrieval, 63.8% (219/343) of references generated by ChatGPT-3 were unidentifiable, and only 37.9% of the verifiable references provided sufficient evidence to support the generated statements (PMID: 37078489). Furthermore, among verifiably real citations generated by GPT-3.5 and GPT-4, 43% and 24%, respectively, contained substantive bibliographic errors, such as incorrect publication years (16% to 22%) or incorrect volume, issue, and page numbers (13% to 34%) (PMID: 37679503).

Guideline Concordance, Numerical Reasoning, and Treatment Recommendations

Studies benchmarking LLM therapeutic outputs against established clinical practice guidelines reveal frequent discordance and critical omissions. In a study evaluating ChatGPT (GPT-3.5) against National Comprehensive Cancer Network (NCCN) guidelines across 104 prompts for breast, prostate, and lung cancer, 34.3% (35/102) of outputs recommending treatment included at least one nonconcordant recommendation, and 12.5% (13/104) recommended completely hallucinated treatments, such as localized therapies for advanced disease (PMID: 37615976). When tested against American Academy of Orthopedic Surgeons (AAOS) guidelines for osteoarthritis, default web-based GPT-4 achieved 62.9% overall concordance, with performance ranging from 30% to 77.5% depending on recommendation strength and prompting strategy (PMID: 38378899). In a multi-institutional evaluation of 10 fictional precision oncology cases presented to a Molecular Tumor Board, four unaugmented LLMs achieved low F1 scores (0.04 to 0.19) when compared to expert human recommendations, with the MTBs noting that 27 of 85 NCT trial identifiers generated by ChatGPT were entirely hallucinated (PMID: 37976064).

A fundamental barrier to safe clinical deployment is the inability of base LLMs to reliably interpret numerical laboratory test results. In the MIMIC-IV acute abdomen study, models failed basic laboratory classification tasks when provided with explicit reference ranges, achieving accuracies as low as 24.1% to 50.1% for abnormally high results and 26.5% to 70.2% for abnormally low results (PMID: 38965432). In addition, LLMs demonstrated severe undertreatment in high-acuity cases: models failed to consistently recommend emergent colectomy for patients with perforated diverticulitis or surgical drainage for infected pancreatic necrosis, while omitting required antibiotic coverage for appendicitis and future colonoscopy surveillance for diverticulitis (PMID: 38965432).

Impact on Clinician Decision-Making, Efficiency, and Cognitive Workload

Randomized controlled trials evaluating physician interaction with LLMs demonstrate nuanced effects on reasoning quality, workflow duration, and potential automation bias. In a randomized trial of 50 attending and resident physicians diagnosing six complex clinical vignettes, access to GPT-4 in addition to conventional resources did not significantly improve diagnostic reasoning scores compared with conventional resources alone (median score 76% vs. 74%, adjusted difference 2 percentage points, 95% CI -4 to 8, p = 0.60), nor did it significantly alter case completion time (519 s vs. 565 s, difference -82 s, 95% CI -195 to 31, p = 0.20) (PMID: 39466245). Notably, standalone GPT-4 scored 92% (16 percentage points higher than the control group, p = 0.03), illustrating that unguided clinician interaction did not translate the model's standalone capability into decision-making gains (PMID: 39466245).

Conversely, in a prospective trial of 92 physicians evaluating 400 complex clinical management cases, physicians randomized to use GPT-4 scored significantly higher in management reasoning than those using conventional resources alone (43.0% vs. 35.7%, mean difference 6.5%, 95% CI 2.7 to 10.2, p < 0.001) (PMID: 39910272). However, LLM-assisted physicians spent significantly more time per case (801.5 s vs. 690.2 s, difference 119.3 s, 95% CI 17.4 to 221.2, p = 0.022) (PMID: 39910272). The authors observed that engaging with the LLM served as a structured 'time out' that prompted deeper reflection on contextual and patient factors, resulting in greater empathy without altering the rate of severe clinical harm (7.6% vs. 7.5%) (PMID: 39910272).

Clinician oversight remains challenging because LLMs exhibit extreme sensitivity to minor variations in input formatting. On MIMIC-IV clinical cases, minor phrasing modifications (e.g., requesting 'main diagnosis' or 'primary diagnosis' instead of 'final diagnosis') caused diagnostic accuracy to swing between +8.7% and -10.6% (PMID: 38965432). Furthermore, models paradoxically degraded in accuracy when provided with complete clinical data compared to single diagnostic tests, and altering the presentation order of identical physical, laboratory, and imaging findings caused diagnostic shifts of up to 18.0% (PMID: 38965432).

Mitigation Strategies: Retrieval-Augmented Generation, Structured Prompting, and Knowledge Graphs

To counter hallucinations, ungrounded outputs, and data poisoning, several architectural and workflow-level mitigations have been validated. Retrieval-Augmented Generation (RAG) grounds LLM outputs in verified external sources. The Almanac framework, which couples GPT-4 with curated databases (PubMed, UpToDate, BMJ Best Practice, MDCalc) via vector retrieval, achieved a 91% valid citation rate, significantly outperformed base models (ChatGPT-4, Bing, Bard) in factuality, completeness, and clinical preference (p < 0.01), and demonstrated 100% resilience to adversarial prompt injection (PMID: 38343631). In preoperative surgical fitness assessments across 14 clinical scenarios, a GPT-4 RAG system incorporating 23 international guidelines achieved 96.4% accuracy compared to 86.6% in human anesthesiologists (OR = 4.84, p = 0.016), with hallucination rates remaining between 0% and 2.9% (PMID: 40185842). In diabetes education, the RISE RAG framework increased the proportion of accurate responses across GPT-4, Claude 2, and Bard by an average of 12% (PMID: 39046096).

Knowledge graphs (KGs) provide deterministic factual verification and token efficiency. The SPOKE KG-RAG framework reduced token consumption by 53.9% relative to Cypher-RAG and improved open-source Llama-2-13b performance on challenging multiple-choice biomedical questions by 71% (PMID: 39288310). When addressing data-poisoning vulnerabilities—where corrupting as little as 0.001% of pretraining tokens increased harmful medical outputs—a defense algorithm verifying extracted named-entity triplets against the BIOS knowledge graph successfully captured 91.9% of harmful content (F1 = 85.7%) (PMID: 39779928).

Prompt engineering and input pre-filtering provide additional non-computational gains. In diagnostic reasoning tasks, chain-of-thought (CoT) prompts structured around differential diagnosis and intuitive reasoning provided interpretable rationales: 65% of incorrect GPT-4 answers exhibited logic errors in their rationales compared to only 18% of correct answers (PMID: 38267608). In acute abdominal evaluations, rule-based filtering that removed normal laboratory values and automated context summarization protocols consistently improved model diagnostic performance by eliminating extraneous numerical noise and preventing context window overflow (PMID: 38965432).

Synthesis and Outlook

Moving forward, the field requires rigorous prospective clinical trials and standardized reporting frameworks—such as the QUEST evaluation principles—to systematically evaluate safety, real-world utility, and patient outcomes before autonomous or semi-autonomous LLMs can be integrated into point-of-care workflows (PMID: 39333376).

Research notebook

What is the overall diagnostic, management, and point-of-care question-answering accuracy of LLMs across clinical specialties?

Across clinical specialties including emergency medicine, radiology, oncology, and rheumatology, LLM diagnostic and triage performance varies substantially by case complexity, information format, and disease prevalence.

Status: verified • Confidence: high

  • In a simulated real-world emergency workflow involving 2,400 clinical cases from MIMIC-IV across four acute abdominal conditions, leading open-access LLMs (Llama 2 Chat, OASST, WizardLM) demonstrated diagnostic accuracies of only 45.5% to 54.9% when required to autonomously gather history, laboratory, and imaging findings, performing significantly worse than human hospitalists who achieved 87.5% to 92.5% accuracy (p < 0.001). (PMID 38965432, results)
  • Standalone GPT-4 evaluated on 6 complex validated clinical vignettes scored 92% on a structured reflection diagnostic reasoning rubric, demonstrating high capability on well-curated case representations, yet physician interaction with the tool did not translate into equivalent performance gains in practice. (PMID 39466245, results)

What is the frequency, nature, and reliability of references and citations generated by LLMs, including hallucinated, fabricated, or non-supporting references?

Even when LLM-generated citations correspond to real published scientific papers, a substantial proportion contain substantive bibliographic errors or fail to support the specific factual claims made in the text.

Status: verified • Confidence: high

  • Evaluation of 636 bibliographic citations across 84 generated academic reviews revealed that 55% of references generated by GPT-3.5 and 18% generated by GPT-4 were entirely fabricated, with fabricated entries frequently synthesizing real researcher names, genuine journal titles, and fictitious article titles or DOIs. (PMID 37679503, results)
  • Among real (non-fabricated) citations generated by GPT-3.5 and GPT-4, 43% and 24% respectively contained substantive errors including incorrect publication years (16-22%), wrong volume/issue/page numbers (13-34%), or misattributed author rosters and journal titles. (PMID 37679503, results)
  • Standard commercial LLMs tested on clinical QA tasks frequently produced citations to unrelated webpages, blogs, or irrelevant biomedical abstracts that did not contain the evidence required to validate the clinical recommendations. (PMID 38343631, results)

How well do LLM recommendations align with clinical practice guidelines and evidence-based medicine standards, and what are common clinical errors or omissions?

LLMs struggle with numerical laboratory interpretation and guideline-mandated diagnostic testing sequences, posing safety risks when deployed without clinical filtering.

Status: verified • Confidence: high

  • In an evaluation of ChatGPT (GPT-3.5) recommendations for breast, prostate, and lung cancer against NCCN guidelines across 104 clinical prompts, 34.3% of outputs recommended at least one nonconcordant treatment and 12.5% recommended completely hallucinated treatments, such as inappropriate localized therapies for advanced metastatic cancer or non-indicated immunotherapies. (PMID 37615976, results)
  • Across 2,400 acute abdomen cases, leading LLMs failed to recommend critical guideline-mandated treatments, drastically undertreated severe conditions (e.g. omitting emergent colectomy in perforated diverticulitis or necrosectomy in infected pancreatic necrosis), and frequently failed to recommend essential supportive therapies and surveillance colonoscopies. (PMID 38965432, results)
  • When tested on categorizing laboratory values as below, within, or above provided reference ranges, LLMs exhibited severe comprehension failures (accuracies as low as 24.1% to 50.1% on abnormally high or low values) and consistently failed to order essential diagnostic test panels (e.g. pancreatic enzymes and severity markers) mandated by guidelines. (PMID 38965432, results)

How does the use of LLMs impact clinician decision-making, diagnostic accuracy, efficiency, cognitive workload, and potential overreliance at the point of care?

Clinician interaction with LLMs introduces unique cognitive dynamics, where unguided prompting, conversational tone, and confident hallucinations risk automation bias and necessitate substantial cognitive effort for manual verification.

Status: verified • Confidence: high

  • In a randomized clinical trial of 50 attending and resident physicians diagnosing complex clinical vignettes, the availability of GPT-4 did not significantly improve diagnostic reasoning compared with conventional resources (median score 76% vs 74%, difference 2 percentage points, 95% CI -4 to 8, p = 0.60; time difference -82 s, p = 0.20), despite GPT-4 alone scoring 92% (p = 0.03). (PMID 39466245, results)
  • In a prospective randomized controlled trial of 92 physicians across 400 complex management cases, physicians using GPT-4 scored significantly higher in management reasoning than those using conventional resources (43.0% vs 35.7%, difference 6.5%, 95% CI 2.7% to 10.2%, p < 0.001), though physicians with LLM access spent 111.3 seconds longer per case (p = 0.022). (PMID 39910272, results)
  • Physicians given unguided access to LLMs frequently underutilized the model's diagnostic capabilities due to suboptimal conversational prompting and lack of structured integration, illustrating that simply providing LLM access does not overcome human cognitive friction. (PMID 39466245, discussion)
  • Engaging with GPT-4 functioned as a cognitive 'time out' that prompted physicians to slow down and consider broader contextual factors, resulting in longer encounter times and more empathetic management framing without increasing the rate of severe clinical harm. (PMID 39910272, discussion)
  • Because LLM outputs vary based on subtle prompt changes, information order, and disease context, clinicians using ungrounded models face substantial cognitive overhead to supervise model reasoning and filter erroneous suggestions. (PMID 38965432, discussion)

What point-of-care architectures and mitigation strategies (e.g., Retrieval-Augmented Generation [RAG], fine-tuning, structured prompting, clinician verification) improve LLM clinical accuracy and reference integrity?

Workflow-level mitigation strategies, including structured diagnostic reflection prompts, automated context summarization, and rule-based pre-filtering of abnormal laboratory results, substantially improve model stability and diagnostic reasoning.

Status: verified • Confidence: medium

  • The Almanac RAG framework, which couples GPT-4 with curated databases (PubMed, UpToDate, BMJ Best Practice, MDCalc) and vector retrieval, achieved a 91% valid citation rate, significantly outperformed standard LLMs in clinical factuality, completeness, and physician preference, and provided 100% resilience against adversarial prompting. (PMID 38343631, results)
  • Automated summarization protocols prevented context degradation over long clinical trajectories, while rule-based pre-filtering that removed normal laboratory results directly improved LLM diagnostic accuracy by preventing attention saturation from extraneous numerical data. (PMID 38965432, results)
Want to take this research further?
Sign up free and the thread will land in your workspace so you can refine the question, ask follow-ups, or branch into related searches.

Top cited papers

PMID 38965432
View
PMID 39466245
View
PMID 37679503
View
PMID 38343631
View
PMID 37615976
View