Patients increasingly consult large language models (LLMs) before specialist review and may do so in widely differing emotional registers. We examined whether re-casting an unruptured intracranial aneurysm (UIA) case from a third-person vignette into a first-person patient account, with anxiety about treatment or non-treatment, alters the recommendations LLMs return. Sixty-seven UIA cases previously discussed at our multidisciplinary team (MDT) were presented to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking under four conditions: third-person vignette (3P), neutral first-person (1PN), and first-person accounts anxious about treatment (1PAT) or non-treatment (1PANT). Each case was submitted five times (4,020 prompts). Majority recommendations were benchmarked against MDT consensus (Cohen's κ, McNemar), tested for directional asymmetry (Bowker), classified as progressive or regressive, and linguistically analysed. Concordance with MDT was fair-to-moderate (71.6-80.0%; κ 0.34-0.51). ChatGPT over-treated under all conditions (p = 0.001-0.013) and Gemini over-treated only at 3P baseline (p = 0.001), abolished under first-person framing. Claude showed no directional preference. First-person framing drove significant clipping-to-coiling migration in Gemini and ChatGPT (p = 0.001 and 0.015). Claude alone hedged under treatment-directed anxiety (p = 0.002). Anxiety-induced shifts were predominantly regressive. Gemini under rupture-anxiety combined the highest empathy-opener rate (85%), highest confidence-marker density and lowest hedging density of any cell, a sycophantic phenotype invisible to categorical analysis. Patient persona and affective framing measurably and model-specifically alter LLM recommendations for UIAs, with anxiety-induced shifts moving models away from rather than toward expert consensus. Clinicians should anticipate AI-shaped expectations varying with a patient's emotional state and counsel against framing-induced advice. This study was retrospectively registered on the Open Science Framework (https://doi.org/10.17605/OSF.IO/HCU6D).
This study was aimed to compare the efficacy of three most popular large language models (LLMs)-Claude Opus 4.6, ChatGPT Thinking 5.4 and DeepSeek v3.2 in answering frequently asked questions (FAQs) about scoliosis. 20 scoliosis related questions (four categories, five questions in each category) were submitted to each LLM. A panel of 9 experts (two spine surgeons, two pediatric orthopedic surgeons and five physical therapists, all blinded to the LLMs and responses) rated independently each response generated by LLMs on a 6 points Likert scale (1 as strongly disagree to 6 as strongly agree). 540 total ratings were collected. Intergroup comparisons were conducted by Kruskal Wallis test and Mann Whitney U pairwise tests. Paired question level analysis was achieved by Friedman test and Wilcoxon signed rank comparisons. Claude's score was 5.53 ± 0.76 much higher than both ChatGPT (4.84 ± 0.86, p < 0.001) and DeepSeek (4.86 ± 0.84, p < 0.001), but no difference was found between ChatGPT and DeepSeek (p = 0.749). Claude performed on top for 19 of 20 questions (95%) and was favored by 7 of 9 reviewers. Consistency of Claude was also highest [CV = 13.8% vs. 17.9% (ChatGPT) and 17.2% (DeepSeek)]. Although all three LLMs achieved favorable overall ratings (>4.8/6), Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency. Within the scope of the present evaluation, Claude demonstrated the strongest overall performance among the three LLMs tested.
Careful evaluation of research methodology is fundamental to scientific progress but represents a significant burden on human experts. The complexity of functional MRI (fMRI) methods makes transparent reporting, as suggested by OHBM COBIDAS guidelines, particularly critical. Large Language Models (LLMs) present a potential solution for rapid, scalable methodological assessment. We evaluated three state-of-the-art LLMs (Gemini 2.5 Pro, Claude 4 Sonnet, ChatGPT-o3-pro) against human expert ratings. Fifty fMRI articles (taken from 2016 to 2025) were independently evaluated by ten human experts and three LLMs using an 82-item COBIDAS based rubric. Human raters demonstrated excellent inter-rater reliability (ICC = 0.801), while LLMs showed poor internal agreement (ICC = 0.254). When comparing total scores across papers, Gemini showed strong positive correlation with human consensus (r = 0.693, p < 0.0001), Claude showed moderate positive correlation (r = 0.394, p = 0.004), while ChatGPT showed negative correlation (r = -0.172, p = 0.233). Gemini maintained high reliability when added to human raters (combined ICC = 0.811), achieving 85.3 % exact agreement and 98.8 % within-1-point agreement. Domain-specific analysis revealed Gemini's consistently high agreement across all six COBIDAS sections (experimental design: 0.915, statistical modeling: 0.880), while ChatGPT and Claude showed weaker, more variable performance. Obvious differences emerged in determining non-applicable items: humans marked 40.5 % as not applicable versus 32.3 % for Gemini, 9.2 % for ChatGPT and 21.1 % for Claude. ChatGPT exhibited extreme score volatility, with papers ranging from 0 to 121 points compared to humans' 44.2-77.7 range. LLM scoring required 1-7 min versus 30-35 min for humans. This proof-of-concept study demonstrates that LLM-assisted methodological evaluation is feasible for complex neuroimaging research and could likely be applied to other research fields.
The current study aimed to quantify the diagnostic accuracy of commonly utilized chatbots including Gemini, Copilot, Claude, and specialized architectures like Manus in the detection and differential diagnosis of various jaw lesions, while concurrently evaluating the clinical safety and fidelity of the information they provide. Cone beam computed tomography (CBCT) dataset from 97 patients presented with jaw lesions were collected and anonymized. Panoramic 2D views were reconstructed from Digital Imaging and Communication in Medicine (DICOM) of all cases using Bluesky Plan software and provided to 4 chatbots (Gemini 2.5 Pro, Copilot, Claude and Manus). Moreover, the DICOM data was provided to Manus followed by prompting. The reports generated were evaluated for accuracy, relevance and feasibility. Statistically significant differences were detected between the chatbots in all measured parameters. In all evaluated parameters Manus CBCT showed the most accurate results (95% of lesions were detected and correctly diagnosed,). Gemini 2.5 pro ranked second where 80% of lesions were detected and 56% were correctly diagnosed. Manus Pan showed less accurate results. The least accurate results were detected in Copilot and Claude. Significant discrepancies exist among artificial intelligence (AI) chatbots regarding their diagnostic accuracy in reporting jaw lesions. Notably, the integration of raw 3-dimensional CBCT data substantially optimizes chatbot performance in lesion detection and diagnosis, as demonstrated by Manus architecture.
To evaluate the performance of four artificial intelligence (AI) systems (ChatGPT 4o, Claude 3.7, Gemini 2.0, and Grok 2) in analysing nasal deformities. The artificial intelligence chatbots were compared to experts in terms of their capacity to analyse nasal deformities. A quantitative analysis compared AI-generated MIRA scores with expert MIRA scores using error measures, Bland-Altman analysis, concordance metrics, and intraclass correlation coefficients to evaluate agreement and systematic bias. A qualitative evaluation was conducted using a 5-point Likert scale to characterise the major nasal type (tension nose, saddle nose, deviated nose, etc.). Fifty adult patients seeking rhinoplasty were evaluated by the chatbots and two experts based on standardised photographs. The evaluations by the two surgeons demonstrated very strong concordance (ICC = 0.997) for nasal analysis using the MIRA scale. Only Claude 3.7 and the experts had comparable total MIRA score evaluations (p > 0.05). Detailed analysis of MIRA sub-scores showed a significant difference between chatbots and experts across all models (p < 0.05), including Claude. Grok 2 (p < 0.001) demonstrated the poorest performance. The qualitative description of the nose by ChatGPT 4o achieved the best results, with an accuracy rate reaching 70%. No model achieved significant performance on MIRA sub-scores in the quantitative analysis of nasal deformity. The qualitative assessment shows that ChatGPT4o could assist, under supervision, with rhinoplasty assessments to analyse major nose types. However, it was effective in only two-thirds of cases. To date, AI tools are not reliable for analysing nasal deformities. This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .
Validated measures of pain catastrophizing primarily assess catastrophizing as a stable trait. However, emerging evidence suggests catastrophizing fluctuates with context, highlighting a need for ecologically valid methods to capture it. This study evaluated large language models (LLMs) as implicit markers of catastrophizing from free-text responses from ninety-one adults with chronic pain receiving long-term opioid therapy (57.3% Female; mean age = 60.5 years). Patients completed baseline measures, including the trait pain catastrophizing scale (PCS), followed by a 10-minute writing task after random assignment to a negative, positive, or neutral pain-coping condition. State affect and pain were assessed before and after writing tasks and again after a cold pressor task (4°C; ≤ 2 minutes). A state PCS followed the cold pressor task. Free-text responses were analyzed using four LLMs (Claude Opus 4; GPT Mini 4o; Llama 4 Maverick; and Gemini 2.5 Pro). ANOVA-based results supported discriminant validity, as all four LLM-derived pain catastrophizing scores differentiated negative from positive and neutral pain-coping conditions. Convergent validity was model-dependent; only Gemini-derived scores correlated with state catastrophizing (r = .22) and pain unpleasantness (r = .23). Divergent validity was mixed. LLM-derived scores were unrelated to pain intensity, but Gemini and Claude-derived scores showed small correlations with trait PCS (r's = .21; 28, respectively). All LLM-derived scores also correlated with negative affect (r's = .29-.41), comparable in magnitude to state PCS, suggesting limited specificity. These findings provide preliminary evidence that certain LLMs may serve as implicit markers of state pain catastrophizing, but further study is needed.
Artificial intelligence (AI) chatbots or large language models (LLMs) are adept at generating language, but their increasing use in the healthcare field, including endodontics, raises concerns about their accuracy. The potential of LLMs to assist clinicians in their decision-making processes regarding vital pulp therapy (VPT) is worth exploring. This study aims to evaluate and compare the responses provided by OpenAI GPT-5.1 Instant, DeepSeek-R1, Claude, Google Gemini, Comet, and Perplexity to clinically relevant questions related to VPT according to the guidelines set by the American Association of Endodontists, European Society of Endodontics, and Indian Endodontic Society. Twenty-three open-ended questions covering various aspects of VPT were developed and presented to OpenAI GPT-5.1 Instant, DeepSeek-R1, Claude, Google Gemini, Comet, and Perplexity. Two experienced endodontists, who were blinded to the different chatbots, evaluated the answers on a 3-point Likert scale. To assess the reproducibility of these answers, the same questions were presented again after 1 month and subsequently saved in a separate Microsoft Word file. The findings were recorded in an Microsoft Excel Sheet, and then statistical analysis was performed. All the LLMs were able to answering all the questions on VPT with almost similar reproducibility across two different intervals. Most tested LLMs, regardless of whether they are free or subscription-based, demonstrated high accuracy and reproducibility when evaluated on guidelines-based questions related to VPT.
Enrollment in phase I oncology trials remains low largely because potentially eligible patients are not identified and evaluated quickly enough. Current clinical trial matching systems can identify candidate patients from the electronic health record, but cases with missing or uncertain eligibility data are often routed for offline manual review. This delay impedes clarification and prolongs the final eligibility determination. This study evaluated TrialTriage, a semiautonomous system built on the n8n platform and designed to resolve eligibility ambiguity during prescreening for phase I oncology trials. When eligibility information is missing or uncertain, TrialTriage emails the investigator, captures the reply, and reruns classification within the same workflow. TrialTriage combined large language model-based variable extraction from free-text clinical narratives and investigator email replies with a deterministic rule engine applying a prespecified 7-criterion protocol. Each case was classified as eligible, not eligible, or ambiguous. Ambiguous cases triggered a structured email query to the investigator, followed by reclassification after a reply. Two requests were sent at 24-hour intervals; after 48 hours without a reply, the case was referred for manual review. The system was tested on 90 synthetic patient cases generated independently by Claude Sonnet 4.6, Gemini 3.1, and Grok 4, with 30 cases per model and balanced distributions of eligible, not eligible, and ambiguous cases. Answer keys were reviewed for accuracy before system execution. Five independent reviewers classified the Claude dataset using a uniform survey form. TrialTriage's classifications were 100% concordant with the author-confirmed ground truth in all 90 synthetic cases (95% CI 96.0%-100.0%). All ambiguous cases were correctly escalated to investigator query. The mean processing time was 2.3 (SD 0.5) minutes per 30-case dataset (range 1.8-2.8 min, approximately 3.5-5.5 s per case). The 5 reviewers achieved a mean accuracy of 96.7% (SD 3.3%), with a Fleiss κ of 0.910, and required a mean of 9.8 (SD 4.8) minutes to review 30 cases. In a subset test of 6 first-pass ambiguous cases, 4 of 6 were reclassified definitively after investigator response, while 2 remained ambiguous because the replies lacked actionable information. TrialTriage demonstrates the feasibility of a semiautonomous prescreening workflow in which ambiguous cases trigger an immediate investigator email query and are reclassified after reply capture with new information within the same system. The main contribution is the integration of email ambiguity resolution into the workflow rather than immediate deferral to offline manual review. Because the evaluation used synthetic cases and label definitions aligned with the same protocol rules used to design the rule engine, these findings should be interpreted as proof of concept and implementation fidelity rather than evidence of real-world clinical performance. Prospective validation using data from real-world electronic health records would be a plausible next step.
Background Clinical histories accompanying imaging orders guide protocol selection and diagnostic focus. However, they are often incomplete, potentially compromising diagnostic accuracy and workflow efficiency. Purpose To evaluate whether large language models (LLMs) can improve the clinical utility of provided imaging indications by leveraging clinical notes. Materials and Methods This retrospective study curated a dataset from deidentified electronic health records at the University of California San Francisco (January 2012 to August 2024), consisting of radiology reports with paired referring clinician-provided and radiologist-curated indications linked to clinical notes. The dataset was stratified across five body systems and five pathophysiologic categories to derive LLM selection and reader study internal test sets. For the reader study, 20 radiologists with 2-25 years of experience compared indications from the referring clinician, radiologist, and best-performing LLMs. Readers scored comprehensiveness, factuality, and conciseness and ranked indications for usefulness in protocoling, usefulness in interpretation, and overall ranking. Models and clinicians were compared using cumulative link mixed models with Tukey-adjusted post hoc comparisons. Results From 28 313 patients (mean age, 59 years ± 20.6 [SD]; 14 912 women), 250 examinations from 247 patients were sampled for the reader study. After nine exclusions, 241 examinations were analyzed, yielding 482 reader-examination evaluations. Indications from the best-performing proprietary (Claude 3.5 Sonnet; Anthropic) and open-source (Qwen 2.5-7B Instruct; Alibaba) LLM were rated as more comprehensive (Likert rating of 5: 37.14% and 28.42%, respectively; both P < .001) and factual (68.05% and 59.75%; both P < .001) than referring clinician indications. The proprietary LLM ranked most useful in protocoling (rank 1: 40.87%; all P < .001), useful in interpretation (44.61%; all P < .001), and overall ranking (44.19%, all P < .001). Comprehensiveness (65.77% of ratings; both P < .001) most strongly influenced overall rankings. Conclusion LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by Yilmaz and Cardoza-Ochoa in this issue.
Accurate assessment of distal radius fracture stability is essential for appropriate triage and timely referral to hand specialists. The LaFontaine criteria provide a structured radiographic framework for predicting instability but are not routinely reported. Recent advances in multimodal large language models (LLMs) capable of direct image interpretation have generated interest in their potential role as adjunctive diagnostic tools. However, their real-world performance in structured radiographic assessment remains unclear. A cross-sectional diagnostic accuracy and agreement study was performed using 20 distal radius fracture radiographs. Five hand surgeons independently assessed each case for the 5 LaFontaine criteria, with majority agreement serving as the reference standard. Two publicly available multimodal LLMs (ChatGPT and Claude) were evaluated using a standardized, single-prompt approach designed to approximate real-world use. Both models were provided identical radiographs and clinical prompts then asked to classify each criterion and overall fracture stability. Agreement and diagnostic performance were calculated relative to surgeon consensus. Hand surgeons demonstrated high consistency in identifying LaFontaine criteria. Agreement between LLMs and clinician consensus varied across individual features, with several criteria showing limited agreement. Both models achieved similar overall accuracy for fracture stability classification (0.75; 95% CI, 0.56-0.94). ChatGPT demonstrated moderate agreement with surgeon consensus, while Claude showed fair agreement. Agreement between clinician-determined instability and independent operative recommendations was moderate. Multimodal LLMs demonstrated variable and generally limited agreement with clinician consensus in classifying fracture stability. Performance across several criteria approached chance levels, and specificity was limited. These findings should be interpreted as exploratory and hypothesis-generating; current models are not yet reliable for clinical decision-making and require substantial validation before potential use as adjunctive triage tools. Diagnostic Level III.
Large language models (LLMs) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear. We systematically reviewed LLM performance across clinical tasks in inflammatory arthritis. We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating LLM performance on clinical tasks in inflammatory arthritis. Two reviewers (Y.A., A.G.) screened 113 records. Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4). Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions. Over 20 distinct LLMs were evaluated, including ChatGPT-3.5 to ChatGPT-4o, Gemini 2.0, DeepSeek-R1/V3, Claude, and Perplexity; ChatGPT/GPT variants were the most frequently tested models (16 of 18 studies), so the current evidence base is predominantly GPT/ChatGPT-based. Findings spanned patient education (n=11), guideline adherence (n=6), clinical reasoning (n=3), and other applications (n=1). All readability assessments exceeded recommended thresholds. Guideline concordance ranged from 48% to 96%. Accuracy was lower for case-based clinical scenarios (4.24/6) than FAQ and guideline-based questions (5.32-5.36/6; p=0.044). When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0). LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions. None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows. Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.
Large language models (LLMs) are increasingly used in higher education, but multi-country evidence on dental students' use, verification, and integrity practices is limited. To compare senior dental students' LLM use, perceived time and academic impact, reliability judgements, verification practices, and integrity safeguards across five countries. An anonymous cross-sectional online survey was administered to final-year dental students in the United Arab Emirates (UAE), Jordan, Malaysia, Oman, and Brazil. Measures included tools used, frequency and motivations, learning activities, perceived time and academic impact, verification frequency and strategies, guideline awareness, and integrity safeguards. Analyses used Kruskal-Wallis and chi-square tests with Benjamini-Hochberg adjustment, effect sizes, Spearman correlations, and ordinal logistic models. In total, 454 students participated (UAE 160, Jordan 101, Malaysia 75, Oman 62, Brazil 56; mean age 22.9; 74.9% female). ChatGPT predominated (95.9%), followed by Gemini, formerly Bard (18.0%), DeepSeek (16.4%), and Claude (7.4%). Tool diversity varied across country-based cohorts, with Oman showing greater multi-tool uptake. Use was frequent (several times/week 39.2%, daily 28.6%). Key motivations were saving time (73.0%), clarifying concepts (56.9%), and summarising (54.1%). Common activities included understanding complex concepts (75.3%), summarising lecture notes (70.0%), exam preparation (61.5%), and assignment research (53.2%); exam-time assistance was reported by 25.6%. Verification was 'always' 20.0% and 'often' 34.1%, varying across country-based cohorts, with Oman verifying less frequently than other cohorts. Guideline awareness was 40.3% overall (UAE 61.3% vs Brazil 8.3%). Integrity safeguards commonly involved paraphrasing (69.6%), citations (39.2%), and plagiarism checks (38.0%); disclaimers were uncommon (9.2%). LLM-use frequency correlated with broader academic use (ρ = 0.289) but not with integrity concern (OR = 0.963). LLM use is widespread and heterogeneous across settings, including non-trivial higher-stakes use. Dental programmes should implement explicit training in verification, evidence traceability, and disclosure, supported by clear, enforceable guidance and assessment designs aligned with real-world LLM practices.
Timely and highly accurate diagnoses by physicians play a crucial role in improving the quality and effectiveness of patient treatment outcomes. Currently, the use of artificial intelligence capabilities in this area has garnered the attention of many health science researchers. Therefore, the main goal of this preliminary study was to compare the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in definitive and differential diagnoses using standardized clinical vignettes. This descriptive comparative study evaluated the diagnostic accuracy of 10 emergency medicine physicians and 4 large language models (LLMs)-ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1-using 10 standardized clinical vignettes. All LLMs were accessed via official web interfaces. Clinical vignettes were developed from emergency department presentations, validated by an expert panel, and presented as text-only inputs to all evaluators. Diagnostic accuracy was assessed using a standardized scoring protocol: definitive diagnoses required an exact match with expert-derived reference standards, while differential diagnoses required ≥ 3 matches. Overall diagnostic accuracy was the primary outcome. Data were analyzed using Pearson's Chi-square test, McNemar's test, and Generalized Estimating Equation (GEE) logistic regression with Bonferroni correction (SPSS version 28). A total of 280 diagnostic evaluations (10 clinical cases assessed by 14 evaluators across 2 diagnosis types) were analyzed. Overall diagnostic accuracy was 61.79%. AI models demonstrated significantly higher overall accuracy (73.75%) compared to emergency medicine physicians (57.00%, p = 0.014). Across all evaluators, definitive diagnoses were more accurate than differential diagnoses (70.71% vs. 52.86%). Generalized Estimating Equation (GEE) analysis revealed a significant interaction between evaluator group and diagnosis type (p = 0.036). Specifically, physicians experienced a significant decline in accuracy when providing differential diagnoses compared to definitive diagnoses (45.0% vs. 69.0%, p = 0.003; remained significant after Bonferroni correction). In contrast, AI models maintained consistently high accuracy across both diagnosis types, with no significant difference between definitive (75.0%) and differential (72.5%) diagnoses (p = 1.000). Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks. While human physicians struggled significantly with differential diagnoses, AI models maintained high and stable performance regardless of the diagnostic complexity. These findings indicate that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool, particularly in complex scenarios requiring differential diagnostic reasoning. Due to study limitations, such as the small number of clinical scenarios and assessors, these findings should be interpreted with significant caution.
Large language models (LLMs) have revolutionized biomedical research, yet they remain prone to hallucinations and struggle with the precise, multi-hop reasoning required for biomedical analysis. To bridge this gap between generative capability of AI model and factual rigor, this article introduces GeneGenie, a model-agnostic, multi-agent framework built upon a directed acyclic graph architecture. Unlike static prompting strategies, GeneGenie implements a deterministic five-node pipeline that orchestrates query planning, intelligent retrieval-augmented generation across curated databases (GenCC, HGNC, and UniProt), and the dynamic execution of bioinformatics tools, including NCBI E-Utilities and local BLAST+. We evaluated the system using the updated 16-module GeneTuring benchmark, comprising 1600 question-answer pairs. The experimental design compared six state-of-the-art models-including GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Pro-operating in a standalone "Direct Mode" versus the agentic "Graph Mode." The results demonstrate that the graph-based architecture consistently outperforms single-model baselines across all metrics. Notably, among the six selected LLM models we explored, Gemini 2.5 Pro achieved the highest performance, correctly answering 1158 questions (72.375% accuracy), compared with the best baseline score of only 15.8%. Furthermore, our evaluation utilized an "LLM-as-Judge" semantic assessment, revealing that the agentic approach significantly enhances not only lexical accuracy but also the completeness and factual grounding of responses. While limitations remain in named entity recognition for protein-coding genes, GeneGenie establishes a robust, reproducible paradigm for future biomedical AI systems, proving that tool-augmented orchestration is superior to reliance on raw model scale alone.
To validate DT-RAG, a curated retrieval-augmented generation system for dental traumatology decision support, against eight commercial large language models. A knowledge base of 250 curated text units ("chunks") from five authoritative sources (IADT 2020, ESE 2021, Krastl 2021, AAE 2013, Cochrane) was built and deployed with Gemini 2.5 Flash as base model. In Study 1, 99 binary clinical questions were submitted in three runs to DT-RAG and eight LLMs; modal accuracy was compared by McNemar exact tests with Holm correction. In Study 2, seven blinded specialists scored DT-RAG against three frontier LLMs on ten clinical scenarios using a 92-point rubric; differences were estimated by linear mixed-effects regression. DT-RAG achieved 96.0% modal accuracy (95% CI 90.1-98.4), significantly exceeding every commercial LLM (best comparator GPT-5.5 87.9%; paired difference +8.1 pp, 95% CI +2.9 to +14.9; p_Holm = 0.013). The curated knowledge base elevated the base model from 49.5% to 98.0% valid rationale rate, eliminating confabulations in this evaluation (0 vs 21). In Study 2, DT-RAG achieved the highest mean score (82.1/92; 89.3%), significantly exceeding Claude Opus 4.5 (72.6; paired difference +9.6 points, exact Wilcoxon p = 0.016), Gemini 2.5 Pro (60.3) and GPT-4.1 (50.4); all seven evaluators ranked DT-RAG first (Kendall's W = 0.97). DT-RAG, a curated retrieval-augmented configuration, outperformed frontier-tier general-purpose LLMs on this dental traumatology benchmark, with no confabulated rationales observed among the responses assessed. This approach may be applicable to other well-defined clinical domains with authoritative guidelines. Curated retrieval augmentation produced accurate, source-traceable and reproducible responses in dental traumatology under benchmark and simulated-scenario conditions, making every error auditable against its source. Clinical safety requires prospective evaluation.
The subcutaneous implantable cardioverter-defibrillator (S-ICD) is an established therapy for sudden cardiac death prevention, but sex-specific outcomes remain incompletely characterized. This study evaluated sex-related differences in baseline characteristics, appropriate shocks, complications, reinterventions, and mortality among S-ICD recipients. The nationwide HONEST (S-ICD French Cohort Study) cohort enrolled all patients who received an S-ICD in France between 2012 and 2019. Clinical endpoints were centrally adjudicated. Sex-specific associations with outcomes were assessed by using propensity score-based inverse probability weighting. Among 4,924 S-ICD recipients, 1,148 were women (23.3%). Compared with men, women were younger (47.3 ± 15.6 years vs 50.6 ± 14.7 years; P < 0.001), less frequently received an implant for primary prevention (57.6% vs 65.1%; P < 0.001), less often had coronary artery disease (38.9% vs 56.5%; P < 0.001), and more often had electrical heart disease (26.3% vs 20.4%; P < 0.001). After adjustment, women had a lower 5-year risk of appropriate shocks (HR: 0.85; 95% CI: 0.74-0.98; P = 0.023) and similar overall complication and reintervention rates but a distinct complication profile, with higher risks of chronic pain (HR: 2.63; 95% CI: 1.61-4.29; P < 0.001) and lead dislodgment (HR: 1.79; 95% CI: 1.09-2.95; P = 0.022) and a lower risk of inappropriate shocks (HR: 0.64; 95% CI: 0.50-0.82; P < 0.001). All-cause mortality was lower in women, whereas S-ICD-unresponsive sudden death and device-related mortality were similar. Women receiving an S-ICD experienced fewer appropriate shocks, with similar overall complication and reintervention rates, but a distinct complication profile. These findings support sex-informed S-ICD selection and follow-up. (S-ICD French Cohort Study [HONEST]; NCT05302115).
Microbial communities were recently revealed in the biliary tract of pancreaticobiliary disorders. However, evidence is limited and comparative data are lacking. We aimed to characterize the biliary microbiota in patients with naïve papilla affected by obstructive jaundice eligible for endoscopic treatment. 222 consecutive patients undergoing ERCP were prospectively enrolled from July 2022 to August 2023. Bile was sampled before and after sphincterotomy,then stored for cultures and resistance profiles. Pre-sphincterotomy (66,6%) and post-sphincterotomy samples (67,5%) revealed bacterial growth, with similar components. Gram-positive bacteria, as Enterococcus spp, were mainly identified. Age ≥60 years, Charlson Comorbidity Index (CCI) ≥4, fever, ongoing antimicrobial therapy and positive blood cultures were associated with positive bile cultures. Positive C-reactive protein was independently related to positive cultures. Multidrug Resistand (MDR) strains, according to international standardized definition, were detected (18%), with a higher prevalence of ESBL bacteria and E. faecium VRE. Antimicrobial therapy was an independent risk factor for MDR biliary bacteria in the multivariate analysis. Positive cultures, polymicrobial flora, and MDR bacteria were similar in malignant and benign disease. Multiple clusters and MDR bacteria were detected in patients with obstructive jaundice. We identified clinical and biochemical risk factors for bacteriobilia and MDR commensals.
Machine learning (ML) has transformed oncological risk prediction by enabling personalized therapeutic strategies. Local tumor control remains a critical endpoint in anal cancer management. This study aimed to develop and validate an explainable ML model for predicting local recurrence at 3 years in patients with anal cancer. We analyzed data from the prospective multicentric FFCD-Anabase cohort, comprising 1,015 patients with anal cancer treated with chemoradiotherapy across 60 French centers between January 2015 and April 2020. The endpoint was local recurrence at 3 years. An extreme gradient boosting model with an Accelerated Failure Time extension was developed to handle time-to-event data. Model inputs combined routinely available clinical, biological, and treatment variables, selected on the training set only. Model development incorporated cross-validation for hyperparameter optimization, followed by calibration to ensure reliable probability estimates. Performance was assessed on an independent test set using discrimination, calibration metrics, and clinical utility measures, with right-censored patients retained through inverse-probability-of-censoring weighting. Model interpretability was enhanced using Shapley Additive exPlanations values, providing global feature importance and individual patient-level prediction explanations. The model demonstrated a C-index of 0.735 (95 % CI 0.65-0.82) and a time-dependent AUC at 3 years of 0.755 (0.66-0.85). The model achieved a sensitivity of 64 % and specificity of 74 %, with positive and negative predictive values of 39 % and 89 %, respectively. All 3-year classification metrics were computed on the full test set using inverse-probability-of-censoring weighting, so that censored patients were not discarded. Overall accuracy was 72 % with an F1-score of 0.49. Calibration performance was assessed using Brier score (0.135) and integrated Brier score (0.131). Calibration plots demonstrated good agreement between predicted and observed probabilities. Decision curve analysis revealed net clinical benefit across a range of threshold probabilities, with optimal risk stratification at a 37 % probability threshold for distinguishing low- and high-risk patients. Kaplan-Meier survival analysis confirmed statistically significant differences between risk groups (p = 2.78 × 10⁻5). Model interpretability analysis using SHAP value (SHapley Additive exPlanations, a method that quantifies each feature's contribution to a model's prediction) identified World Health Organization (WHO) performance status as the most influential predictor, followed by tumor size and age. Our model yielded good performances on real-world data to predict the risk of local recurrence at 3 years for anal cancer.
Nosocomial transmission of respiratory infections poses a major threat to patient safety, while also affecting healthcare workers' (HCW) health, generating substantial costs for hospitals. These infections spread through both close-proximity interactions at short distances, and via aerosols that remain suspended in the air, enabling long-range transmission within a room when a susceptible individual is at a distance from an infectious individual. The relative contribution of each transmission route is pathogen-dependent. However, models distinguishing them remain scarce, limiting the design of effective intervention strategies. Here, we propose a novel agent-based stochastic model of respiratory pathogen transmission in a hospital ward that integrates both transmission routes together with contact patterns and individual movements. After informing our model with real close-proximity interaction data collected in two French intensive care units, we simulate a range of combinations of short- and long-range transmission levels to investigate their differences. Selecting parameter values that keep overall ward transmission intensity stable across combinations, the model is used to further evaluate the impact of intervention strategies on incidence risk. We find that the predominance of one route over another has little effect on overall outbreak dynamics, though the impact on individuals varies markedly. Patients are mostly at risk of short-range transmission from HCWs, while HCWs are mostly affected by whichever route is predominant. This directly influences intervention effectiveness. Universal masking emerges as the most effective strategy, reducing both transmission routes. Its stringency can be relaxed with limited loss of effectiveness when combined with ventilation in relevant rooms. Importantly, interventions targeting HCWs, notably ventilation in rooms not accessible to patients, indirectly reduces incidence in patients, with a stronger effect when coupled with relaxed masking interventions. Finally, intervention ranking remains robust across parameter values, as confirmed by a sensitivity analysis. This new model highlights the importance of explicitly considering physical mechanisms of transmission, and the need for interventions that remain effective irrespective of pathogen characteristics and ward organization.
Injection-related bacterial infections represent a major but under-recognised health issue among people who inject drugs (PWID). Harm reduction interventions (HRIs) like needle and syringe programmes (NSP) and opioid agonist treatment (OAT) could mitigate their burden. This review aimed to identify and synthesise evidence on the effectiveness of HRIs in preventing bacterial infections among PWID. Systematic review with meta-analysis of studies with more than 40 participants from Medline, Embase, Cochrane Library and Web of Science, published between 1990 and 2023 in English or French. We included interventional and observational studies that reported a quantitative effect measure for an HRI conducted in community, harm reduction, healthcare and outreach settings, targeting bacterial infections among PWID, defined as individuals who had injected drugs at least once within the previous year. Twelve studies met the inclusion criteria, for a total of n = 11 611 participants. The primary outcome was the impact of HRI on the prevalence or incidence of bacterial infections, measured as a risk difference, relative risk, number needed to treat, relative risk reduction, odds ratio, incidence rate ratio, hazard ratio or preventable fraction among the unexposed. Risk of bias was assessed using the Newcastle-Ottawa Scale for observational studies and the Cochrane RoB 2 tool for randomised trials. Meta-analysis was performed when at least 3 comparable estimates were available. Overall, the available evidence was sparse and heterogeneous, with substantial variability in study design, intervention definitions and outcome measurement across the 12 included studies. Sterile injecting equipment provision was found protective in 2/6 studies (n = 1938), 1/6 (n = 5209) found increased risk and 3/6 (n = 1323) reported no statistically significant association. OAT was protective in 2 studies (n = 2934) when comparing current PWID or those who had never used OAT to past PWID. Only 1 study (n = 1876) evaluated a combination of these interventions, showing a statistically significant reduction in skin and soft tissue infections. Among hygiene interventions, 1 of 2 studies (n = 59) reported a statistically significant protective effect, and the same for drug consumption rooms (n = 665). Overall, 8/10 studies assessed were judged to be at high risk of bias. A random-effect meta-analysis of crude odds-ratios (ORs) associated with NSPs yielded a pooled OR of 1.25 (95% confidence interval = 1.07-1.47). Evidence on the effectiveness of harm reduction interventions in preventing bacterial infections among people who inject drugs is limited and inconsistent, as most studies are observational, focus on skin and soft tissue infections and present substantial methodological limitations.