BackgroundPosttraumatic stress disorder (PTSD) is common yet frequently underdiagnosed, in part due to barriers to systematic screening and the reliance on self-report instruments. Large language models (LLMs) have shown promise in extracting clinically relevant information from unstructured language, but their ability to infer item-level PTSD symptom severity from clinical interviews remains unclear.MethodsUsing the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WoZ), we analyzed 100 semi-structured clinical interview transcripts paired with item-level PTSD Checklist-Civilian Version (PCL-C) scores. Six LLMs (DeepSeek 3.1, Claude Sonnet 4, LLaMA 4 Scout, GPT-4o, GPT-5, and Gemini 2.5 Flash) used zero-shot prompting to predict all 17 PCL-C items. Performance was assessed for binary symptom endorsement (≥3 vs. < 3), 5-point Likert prediction, and DSM-IV symptom-cluster analyses using accuracy, F1 score, and Matthews correlation coefficient (MCC).ResultsFor binary prediction, Claude 4 achieved the highest mean accuracy (0.705; 95% CI, 0.681-0.728), followed by DeepSeek 3.1(0.699; 95% CI, 0.675-0.724) and Gemini 2.5 (0.698; 95% CI, 0.677-0.718). For Likert prediction, DeepSeek 3.1 performed best (accuracy = 0.438; 95% CI, 0.401-0.475), only modestly above the majority-class baseline (0.399; 95% CI, 0.355-0.443). Performance varied by symptom domain, with re-experiencing and hyperarousal symptoms generally predicted more accurately than avoidance/numbing symptoms. Across models, predicted item-level symptom patterns showed a meaningful alignment with observed PCL-C responses despite reduced accuracy in fine-grained severity estimation.ConclusionZero-shot LLMs' performance was insufficient for clinical application in predicting PTSD symptoms from semi-structured interview transcripts. While models showed some ability to capture overall symptom patterns, performance varied across domains and remained limited for fine-grained severity estimation. Given these constraints and the non-trauma-specific nature of the dataset, findings should be interpreted as preliminary, with only modest differences observed between models.Plain Language Summary TitleCan Artificial Intelligence Identify PTSD Symptoms from Conversations? A Study Using Clinical Interview TranscriptsPlain Language SummaryPost-traumatic stress disorder (PTSD) is a common mental health condition, but it is often missed in clinical settings. Screening usually relies on questionnaires that patients must complete themselves, which may not always happen due to time, stigma, or discomfort discussing trauma. Researchers are exploring whether artificial intelligence (AI) could help identify PTSD symptoms from conversations instead.In this study, we tested several advanced AI systems, known as large language models, to see if they could estimate PTSD symptoms based on written transcripts of clinical interviews. These interviews were not specifically designed to assess trauma, which makes the task more challenging but closer to real-world situations. We compared the AI predictions to participants' own questionnaire responses about their symptoms.We found that the AI models were somewhat able to recognize general patterns of PTSD symptoms, especially more visible ones like sleep problems or distressing dreams. However, they struggled with more internal or less obvious symptoms, such as avoidance or emotional numbness. Overall, their accuracy was moderate and not reliable enough for clinical use, particularly when trying to estimate how severe symptoms were.Importantly, differences between the AI models were small, and none performed well enough to replace existing screening methods. These findings suggest that while AI may have future potential as a supportive tool, it is not yet ready to be used for diagnosing or screening PTSD on its own.Further research using better data, improved methods, and real clinical settings is needed before this approach could be considered for practical use. Le trouble de stress post-traumatique (TSPT) est fréquent, mais il est souvent sous-diagnostiqué, en partie en raison d'obstacles au dépistage systématique et de l'utilisation d'instruments d'autodéclaration. Les grands modèles de langage (LLM) se sont révélés prometteurs pour ce qui est d'extraire des renseignements cliniquement pertinents d'un langage non structuré, mais leur capacité à déduire la gravité des symptômes de TSPT selon des éléments à partir d'entrevues cliniques demeure incertaine. À l'aide du corpus d'entrevues Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WoZ), nous avons analysé 100 transcriptions d'entrevues cliniques semi-structurées appariées au score pour les éléments de la liste de vérification pour le TSPT - Version civile (PCL-C). Six GML (DeepSeek 3.1, Claude Sonnet 4, LLaMA 4 Scout, GPT-4o, GPT-5 et Gemini 2.5 Flash) ont utilisé des requêtes sans exemple pour prédire les 17 éléments de la PCL-C. Le rendement a été évalué en ce qui concerne l'approbation binaire des symptômes (≥ 3 p/r à < 3), la prédiction de l'échelle Likert à 5 points et les analyses du groupe des symptômes dans le DSM-IV selon l'exactitude, le F-score et le coefficient de corrélation de Matthews (MCC). Pour la prédiction binaire, Claude 4 a obtenu l'exactitude moyenne la plus élevée (0,705; IC à 95 % : de 0,681 à 0,728), suivi de DeepSeek 3.1 (0,699; IC à 95 % : de 0,675 à 0,724) et de Gemini 2.5 (0,698; IC à 95 % : de 0,677 à 0,718). Pour la prédiction de l'échelle de Likert, DeepSeek 3.1 a obtenu les meilleurs résultats (exactitude = 0,438; IC à 95 % : de 0,401 à 0,475), légèrement supérieurs aux valeurs initiales de la classe majoritaire (0,399; IC à 95 % : de 0,355 à 0,443). Le rendement variait en fonction du domaine de symptômes, les symptômes de reviviscence et d'hypervigilance étant généralement prédits avec plus de précision que les symptômes d'évitement ou d'émoussement. Dans l'ensemble des modèles, les schémas des symptômes prédits au niveau des éléments ont montré une harmonisation significative avec les réponses observées au questionnaire PCL-C, malgré une exactitude réduite dans l'estimation fine de la gravité. Le rendement des LLM sans exemple était insuffisant pour une application clinique dans la prédiction des symptômes de TSPT à partir de transcriptions d'entrevues semi-structurées. Bien que les modèles aient montré une certaine capacité à saisir les tendances globales des symptômes, le rendement variait d'un domaine à l'autre et demeurait limité pour l'estimation fine de la gravité. Compte tenu de ces contraintes et de la nature non traumatique de l'ensemble de données, les résultats doivent être considérés comme préliminaires, avec de légères différences observées entre les modèles.
After-discharge instructions often fail due to poor usability and language misalignment. We evaluated a clinician-supervised method for generating instructions for common emergency department presentations using a clinician-supervised method using large language models. Eight common ED presentations were identified via a physician needs assessment. After-discharge instructions were generated using three publicly accessible large language models (ChatGPT‑4.0, Claude‑3.5 Haiku, and Gemini‑2.0 Flash Thinking) and iteratively refined by expert physicians. After-discharge instructions were assessed for clinical accuracy, completeness, readability (Flesch Reading Ease scale), semantic similarity, and understandability. This analysis was repeated after the inclusion of reviewer edits. Five AI-simulated personas based on local marginalised patient profiles were used to identify comprehension barriers. We applied Bag‑of‑Words and BioClinical BERT similarity metrics to objectively quantify the semantic and contextual consistency of LLM outputs beyond what readability scores alone capture. All large language models produced clinically accurate instructions. Physician edits improved accuracy but paradoxically reduced objective readability scores. Whilst Claude was preferred for simpler language after revisions, persona reviews revealed persistent medical jargon and vague instructions that could hinder understanding for marginalised groups. Large language models with expert clinician supervision can create clinically accurate after-discharge instructions. However, clinician-led refinements decreased readability, increasing the risk of poorer post-visit patient understanding. AI-simulated personas may offer a scalable method to surface potential comprehension barriers in patient instructions, but must be followed by validation with real patients. RéSUMé: OBJECTIF: Les instructions après déchargement échouent souvent en raison d’une mauvaise utilisation et d’un mauvais alignement linguistique. Nous avons évalué une méthode supervisée par un clinicien pour générer des instructions pour les présentations courantes aux services d’urgence à l’aide d’une méthode supervisée par un clinicien utilisant de grands modèles de langage. MéTHODES: Huit présentations courantes de TCA ont été identifiées par une évaluation des besoins des médecins. Les instructions après la sortie ont été générées à l’aide de trois grands modèles de langage accessibles au public (ChatGPT‐4.0, Claude‐3.5 Haiku et Gemini‐2.0 Flash Thinking) et affinées de manière itérative par des médecins experts. Les instructions après la sortie ont été évaluées pour leur précision clinique, leur exhaustivité, leur lisibilité (échelle de facilité de lecture Flesch), leur similarité sémantique et leur compréhensibilité. Cette analyse a été répétée après l’inclusion des modifications du réviseur. Cinq personnes simulées par l’IA, basées sur des profils locaux de patients marginalisés, ont été utilisées pour identifier les obstacles à la compréhension. Nous avons appliqué les métriques de similarité BERT Bag of‐Words et BioClinical pour quantifier objectivement la cohérence sémantique et contextuelle des résultats LLM au-delà de ce que les scores de lisibilité à eux seuls capturent. RéSULTATS: Tous les grands modèles de langue ont produit des instructions cliniquement précises. Le médecin améliore la précision mais paradoxalement réduit les scores de lisibilité objective. Bien que Claude ait été préféré pour un langage plus simple après les révisions, les revues de personnes ont révélé un jargon médical persistant et des instructions vagues qui pourraient entraver la compréhension pour les groupes marginalisés. CONCLUSIONS: Des modèles à large langage supervisés par des cliniciens experts peuvent créer des instructions cliniquement précises après la sortie. Cependant, les améliorations apportées par les cliniciens ont réduit la lisibilité, augmentant ainsi le risque d’une mauvaise compréhension du patient après la visite. Les personnes simulées par l’IA peuvent offrir une méthode évolutive pour mettre en évidence les obstacles potentiels à la compréhension dans les instructions des patients, mais doivent être suivies d’une validation avec de vrais patients.
The emergence of AI technology has sparked curiosity regarding the capabilities of large language models (LLMs) in the field of medicine. Minimal research exists regarding the proficiency of various AI models in ethics scenarios, specifically in specialty-based scenarios. This study aimed to compare the performance of GPT-4o and Claude Sonnet 4 on ethics questions with that of medical students and orthopedic residents. A total of 200 ethical or legal scenario questions were randomly selected from question banks targeted for third- and fourth-year medical students (UWorld, AMBOSS) and orthopedic residents (OrthoBullets). Questions at the medical student level were exclusively text-based, while resident-level questions included text-based questions accompanied by images. Each question was entered identically into each AI model 3 separate times. If answers varied between trials, the answer provided most frequently by the model was used as the selected answer. GPT-4o correctly answered 140 (70%) of 200 questions, which was similar to the average human test taker score of 71% (~142/200 questions). Claude correctly answered 180 (89%) questions, a score greater than that of human test takers and significantly better than GPT-4o (P<.001). Claude scored significantly higher than GPT-4o in almost all question categories. GPT-4o provided different responses to identically worded trials for 27 (21%) of 130 general questions and 3 (4%) of 70 orthopedic questions (P=.002), while Claude did not have a significant difference in variability between these 2 groups (general: 16/130, 12% vs orthopedic: 3/70, 4%; P=.06). GPT-4o selected the incorrect response for 60 (30%) total questions and chose the incorrect response most commonly selected by humans significantly more frequently on UWorld interpersonal-specific questions (30/40, 75%) than on UWorld all social sciences (27/40, 68%; P=.03). Claude showed no significant difference in the rate of most common incorrect response selection between question categories. These results suggest that GPT-4o can potentially answer both general and specialty-specific ethical questions with similar proficiency to sample groups of both medical students and orthopedic residents, while Claude AI performs significantly better than both humans and GPT-4o. Variables such as AI model framework and training data may drive the observed difference in performance, but the exact cause cannot be definitively isolated without intentional testing. Therefore, further research is needed to ensure safety by minimizing output variability before integrating AI as a patient-facing resource.
The indication for surgery for isolated lateral malleolar fracture (AO/OTA 44B1) is debatable and in many cases, relies upon radiographic assessment of fracture stability. Artificial intelligence chatbots with visual analysis capabilities offer a potential attribute for radiographic assessment. This study compared the diagnostic performance of a commercially available AI chatbot with that of three fellowship-trained orthopedic trauma surgeons in evaluating equivocal isolated lateral malleolar fractures. A retrospective study. 50 patients with isolated lateral malleolar injury at the level of the syndesmosis (AO/OTA 44B1) were evaluated by three blinded fellowship-trained orthopedic trauma surgeons. Each rater measured standardized radiographic ankle parameters, medial clear space, tibiofibular clear space, and tibiofibular overlap, on anteroposterior and mortise views and determined a surgical versus nonoperative treatment recommendation. Subsequently, the same sets of radiographs were independently evaluated by an AI chatbot (Claude, Anthropic). The observers and AI decisions were compared to the actual outcome of the patients (operative vs. nonoperatives). All raters recommended surgery at lower rates (34.0-46.0%) than the actual operative rate (56.0%). The difference in outcomes between the actual treatment and the observers varied and ranged between 67.3-86% with the AI within the same ranges. The AI's radiographic measurements differed systematically from all surgeons across five of six parameters. Inter-rater agreement between the AI and surgeons was slight, while inter-surgeon agreement was moderate (κ = 0.457-0.589). ROC analysis showed comparable AUC values (0.63-0.67) for all raters. The AI chatbot demonstrated diagnostic accuracy comparable to orthopedic trauma surgeons in directing treatment for isolated lateral malleolar fractures, despite using a systematically different measurement strategy. All raters exhibited conservative bias as comparted with the actual outcome with modest discriminatory ability, reflecting the inherent difficulty of this clinical issue. These findings support a potential complementary role for AI in ankle fracture triage, while final clinical management decisions should remain in the hands of the orthopedic surgeon.
Surgical consent documents are frequently written at reading levels exceeding average health literacy. Large language models (LLMs) may offer a scalable approach to generating clearer, procedure-specific consent forms. This study evaluated the clarity, clinical accuracy, and acceptability of consent forms generated by GPT-4 and Claude for common otolaryngologic procedures. Twenty AI-generated consent forms (10 GPT-4.0, 10 Claude-2.1) were produced using standardized prompts. In a survey-based, non-clinical setting, five board-certified otolaryngologists independently rated each form for medical accuracy, readability, comprehensibility, legal/ethical sufficiency, and usability using a 4-point scale. A cross-sectional cohort of 300 English-speaking adults (15 raters per form) evaluated perceived clarity and signing comfort on 5-point Likert scales, and perceived trust using a binary (Yes/No) item, and completed eight binary quality assessments. A blinded subgroup (n = 10) compared AI-generated and official national health system templates across five Likert domains. Readability was assessed using Flesch-Kincaid Grade Level (FKGL). Mean lay ratings for clarity across AI-generated forms were high overall. Claude demonstrated numerically higher scores than GPT-4 for clarity (4.72 vs. 4.68), perceived trust (reported as proportions), and signing comfort (4.40 vs. 4.27). However, when analyzed at the form level, differences between models were not statistically significant for clarity (mean difference 0.04; t(9) = -0.80; p = 0.44) or signing comfort (mean difference 0.13, t(9) = -1.68, p = 0.13). Across binary domains, ≥ 95% of participants affirmed adequate explanation of risks, benefits, and alternatives. Experts rated GPT-4 more accurate than Claude (2.2 vs. 1.5, p = 0.034). Mean Flesch-Kincaid Grade Level was lower for AI-generated forms compared to official templates. Although prompts targeted a 6th-8th grade reading level, achieved readability scores were slightly higher (8.8-9.4). In a non-clinical evaluation, AI-generated consent forms were perceived as clear and clinically complete, with model-specific trade-offs between perceived clarity and clinical detail. These perception-based findings-reflecting participant ratings of clarity, perceived trust, and willingness to sign rather than objective comprehension-are hypothesis-generating, and prospective clinical and legal validation in more representative patient populations is required.
Auditory event-related potential (ERP) brain-computer interfaces (BCIs) offer communication support for individuals with amyotrophic lateral sclerosis (ALS) who eventually progress to completely locked-in states. However, individual-specific BCI pipeline optimization is technically demanding and time-consuming, leaving substantial room for performance improvement in practice. A central challenge is increasing selection speed while maintaining reliable classification accuracy, since slower selections reduce the sense of agency and undermine the motivational and feedback dynamics essential for sustained BCI use. We investigated whether an AI coding assistant could address this challenge for individual patients. A three-class auditory ERP-BCI was optimized for a single ALS patient using Claude Code (Anthropic, Inc.), which iteratively generated and evaluated 23 optimization scripts over approximately 24 hours with minimal human-in-the-loop oversight. The resulting AI-Designed ERP classifier (AIDE) was evaluated on 189 EEG trials spanning 3.5 years using five cross-validation strategies. For the baseline models, halving the stimulus repetitions to shorten selection time degraded classification accuracy; AIDE prevented this degradation, achieving 85.03% mean cross-validation accuracy (selection time 17 s; ITR 2.92 bits/min). This doubled the information transfer rate from 1.43 to 2.92 bits/min. Accuracy exceeded 84% across four of five cross-validation strategies. Feature space visualization revealed that the AI autonomously selected and combined EEG features established in prior studies into an effective discriminative architecture, without domain-specific algorithmic guidance from the human researcher. In addition, online test confirmed 66.7% accuracy for AIDE versus 50.0% for the baseline model. These findings provide proof of concept that single-subject BCI performance can be improved via a single prompt, offering an efficient pathway to individualized optimization in clinical and research settings.
Artificial intelligence (AI) is increasingly integrated into scientific publishing workflows, yet no study has formally evaluated the ability of large language models (LLMs) to reproduce human editorial desk-review (R0) decisions in a general orthopaedic surgery journal. We investigated whether three commercially available LLMs could accurately replicate the editorial decisions of the Editorial Board of Orthopaedics & Traumatology: Surgery & Research (OTSR). The study addressed four questions: (1) Is the concordance between LLM and human R0 decisions satisfactory for editorial use? (2) Do LLMs exhibit a severity bias? (3) Do LLMs generate decision letters of acceptable quality, and do they reproduce the specific critiques of human reviewers? (4) Does prompt complexity influence LLM decision-making? LLMs used without task-specific fine-tuning or prior exposure to the study corpus would demonstrate at least moderate concordance (κ ≥ 0.40) with human editorial decisions. A corpus of 32 manuscripts randomly selected from submissions to OTSR between 2025 and 2026 (n = 10 outright rejected at R0: 3 out of scope, 3 plagiarism/dual submission, 4 direct desk rejection; n = 11 accepted for peer review; n = 11 rejected after full peer review) was anonymised and independently evaluated, without task-specific fine-tuning or prior exposure to the study corpus, by ChatGPT (GPT-5.5, OpenAI), Gemini (3.1, Google), and Claude (Sonnet 4.6, Anthropic) using a structured prompt incorporating the OTSR guidelines. The primary outcome was assessed using Cohen's kappa between LLM and human binary decisions. Secondary outcomes included accuracy, inter-LLM agreement, domain-specific scoring, ARCADIA quality scoring of 115 eligible decision letters by two blinded raters with ICC, human-performed content concordance analysis, sensitivity analysis (structured vs. minimal prompt), and test-retest reproducibility at 24 hours. Overall accuracy (i.e. the decision was similar for LLM and editorial decision) was 59.4% for ChatGPT (19/32) and Claude (19/32), and 62.5% for Gemini (20/32). Cohen's kappa was near-zero for ChatGPT (κ = -0.05) and Gemini (κ = 0.00), and low for Claude (κ = 0.15). All LLMs showed systematic over-rejection of accepted manuscripts. No LLM identified plagiarism or simultaneous dual submission as a rejection motive. Test-retest concordance was 84.4 - 90.6% across models. ARCADIA quality scoring (n = 115 scorable letters, inter-rater ICC = 0.86, 95% CI 0.80-0.90) showed Claude achieved the highest scores (4.46 ± 0.32 /5), significantly above the human OTSR letters (4.01 ± 0.44, p < 0.001), ChatGPT (3.81 ± 0.49, p = 0.002), and Gemini (3.28 ± 0.49, p < 0.001). LLMs reproduced 30 - 41% of human reviewer-specific critiques, with Claude achieving the highest match (40.5%) without hallucinations. Switching to a minimal prompt markedly increased acceptance rates for Gemini (87.5%) and Claude (75%), while ChatGPT remained largely insensitive to prompt simplification (15.6%). The principal finding of this study is the dissociation between formal review quality and true editorial reliability. Although modern LLMs generated persuasive and methodologically structured decision letters, they failed to achieve meaningful concordance with real editorial decisions and displayed stable architecture-specific biases that were highly sensitive to prompt design. These results indicate that current LLMs reproduce the surface features of peer review more successfully than its underlying scientific and contextual reasoning. Consequently, LLMs may represent valuable supervised assistants for editorial workflows, but not reliable autonomous substitutes for human editorial expertise in orthopaedic scientific publishing. IV; Observational pilot study, concordance analysis.
Consumer artificial intelligence chatbots are now accessed by hundreds of millions of users seeking health information, yet systematic evaluation of their safety boundary maintenance under real-world caregiver pressure remains scarce. We evaluated PediatricSafetyBench-v2, a benchmark of 600 pediatric health queries comprising 300 authentic caregiver queries sourced from the HealthCareMagic-100k-en physician consultation corpus and 300 matched adversarial variants incorporating six operationalized caregiver pressure patterns, across four consumer AI systems (GPT-4o-mini, Gemini-2.0-Flash, Claude-3.5-Haiku, and Llama-3.1-8B). Safety boundary maintenance was assessed using a validated five-component Safety Composite Score (maximum 15 points; safety-appropriate threshold of 10 or above), validated against independent human raters prior to full-corpus application (mean weighted kappa 0.76; Pearson r = 0.88). The overall safety-appropriate rate was 95.5%. Safety-oriented system prompt deployment improved safety-appropriate rates by 5.9 percentage points across all four models. Counter-intuitively, adversarial caregiver pressure was associated with higher rather than lower Safety Composite Score values for all four models across all ten topic categories and severity levels. False expertise claims were the most vulnerability-inducing pressure pattern, whereas emotional escalation was associated with the highest scores. Consumer AI systems maintain safety boundaries in the large majority of pediatric health interactions. PediatricSafetyBench-v2 is publicly released for longitudinal safety monitoring.
(1) Background: Clinical trial data extraction from registries such as ClinicalTrials.gov remains labor-intensive and error-prone, often missing critical details hidden in unstructured protocol descriptions. Large Language Models (LLMs) offer potential to automate this process, yet systematic multi-model comparisons on real clinical trial data remain scarce. (2) Methods: Four LLMs (OpenAI o4-mini-high, Anthropic Claude-Sonnet-4, Google Gemini 2.5-Pro, and Meta Llama-4-Maverick) extracted brain stimulation parameters from 67 transcranial direct current stimulation (tDCS) trials in Parkinson's disease via a structured JSON schema. Pairwise inter-model agreement was quantified with Cohen's Kappa and percentage agreement across binary, categorical, and multi-component task tiers. (3) Results: Under exact-string matching, agreement was near-perfect for binary classifications (non-invasive classification: 100%; brain stimulation presence: 99.3%, κ = 0.50) and substantial for categorical extractions (primary stimulation type: 96.4%, κ = 0.70), but fell to 48.6% (κ = 0.43) for complex anatomical targets. Numeric parameters revealed model-specific strengths: o4-mini-high and Claude-Sonnet-4 achieved perfect duration agreement (r = 1.000, n = 19) while Llama-4-Maverick diverged substantially (r < 0.12). Validation against an expert gold standard (100% inter-annotator agreement on a 20-trial overlap) confirmed high extraction accuracy across all features (mean 93.7-98.9%). Crucially, the low agreement on anatomical targets proved to be an artifact of exact-string scoring: under the same semantic matching used to measure accuracy, inter-model agreement rose to 97.0%, coinciding with the 95.5% expert accuracy. Inter-model agreement therefore tracks accuracy once both are measured on a common basis. (4) Conclusions: Exact-string inter-model agreement decreases with task complexity, but this decline largely reflects interchangeable free-text wording rather than reduced accuracy. Evaluated semantically, agreement and expert accuracy are both high and closely aligned. A residual risk is not low accuracy but the rare error shared across all models, which agreement cannot detect, and which overall accuracy can itself mask when one class dominates. These findings inform hybrid human-AI systematic review pipelines in which targeted expert oversight focuses on shared-error and minority-class detection.
Childhood accidents are among the leading causes of injury during early childhood. This study aimed to evaluate and compare the accuracy, clarity, and comprehensiveness of pediatric first aid information generated by LLMs. A cross-sectional comparative evaluation design was employed. Twenty standardized pediatric first aid questions were developed based on international guidelines and expert consensus. Responses generated by ChatGPT, Claude, Gemini, and Copilot were independently evaluated by a pediatric nurse and a physician using a 5-point Likert scale. Inter-rater reliability was assessed using Cohen's kappa coefficient, and differences among models were analyzed using one-way analysis of variance. Moderate inter-rater agreement was observed across all evaluation domains. Statistically significant differences were identified among the four LLMs. Claude demonstrated the highest overall performance across all evaluation domains. Gemini demonstrated relatively high accuracy but lower clarity and comprehensiveness scores. Copilot performed well in clarity but showed limited depth of clinical content. ChatGPT received the lowest scores across all assessed domains. The findings reveal considerable variability in the quality of pediatric first aid information generated by LLMs. While certain models may serve as supportive educational tools, none should be considered a substitute for professional medical assessment or emergency care.
To compare the quality of tinnitus-related information generated by multiple generative artificial intelligence (GenAI) systems and web search using expert evaluation. Cross-sectional comparative study. Digital platforms evaluated in their native public interfaces. Not applicable. Thirty commonly searched tinnitus-related questions derived from United States Google Trends data (2020-2025). Questions were submitted to 6 GenAI systems (OpenEvidence, Claude, DeepSeek, GPT-5, Gemini, GPT-4) and Google Search (first organic result). Responses were independently rated by 6 experts using the QAMAI framework. Mean expert-rated quality scores across 5 domains (accuracy, clarity, relevance, completeness, and usefulness). Overall quality differed significantly across systems (Friedman P<0.001; Kendall W=0.34). OpenEvidence achieved the highest mean score (4.45±0.72; 95% CI: 4.40-4.49), followed by Claude (4.00±1.02), DeepSeek (3.92±1.13), GPT-4 (3.89±0.84), Gemini (3.62±0.98), and GPT-5 (3.30±1.11). Google Search scored lowest (2.27±1.12; 95% CI: 2.20-2.35). Completeness was the lowest-performing domain across systems (range: 1.70-4.41). Pairwise comparisons showed significant differences between OpenEvidence and all other systems (effect size r=0.49-0.86). Inter-rater reliability was high (ICC=0.82). Readability demonstrated an inverse pattern relative to expert-rated quality. OpenEvidence demonstrated the lowest readability (Flesch-Kincaid Grade Level 17.5; Flesch Reading Ease 2.2), corresponding to a postgraduate reading level, whereas general-purpose LLMs produced more accessible responses at a sixth-seventh grade reading level. The quality of tinnitus information varies substantially across digital platforms. While GenAI systems generally outperform web search, deficiencies in completeness persist. Readability analysis revealed an inverse relationship between expert-rated quality and response accessibility, suggesting that clinician and patient assessments of informational value may not always align. These findings highlight the need for continued evaluation and clinician oversight to ensure safe, comprehensive, and accessible patient-facing information.
Since the 1994 "Back to Sleep" campaign, pediatricians have promoted evidence-based infant safe sleep practices to reduce sleep-related infant deaths. However, caregivers increasingly seek guidance online. We sought to determine the accuracy of large language model (LLM) responses to caregiver questions about infant safe sleep, compared with the American Academy of Pediatrics' (AAP) 2022 recommendations. Nine caregiver questions adapted from Reddit New Parents forum were mapped to core AAP safe sleep topics. Each was entered into three LLMs: ChatGPT 5, Gemini 2.5 Flash, and Claude Sonnet 4.5, three times within the same day to assess stability. Three reviewers scored responses on a 0-2 scale for accuracy (primary outcome), completeness, and empathy. Stability reflected similarity across repeated responses. Readability was calculated using the Flesch-Kincaid grade level. Mean scores were compared using descriptive statistics, analysis of variance, and post hoc testing. Mean accuracy varied significantly across models. Gemini had the highest accuracy score (mean 1.85), followed by Claude (1.44), and ChatGPT (1.30). Gemini was significantly more accurate than ChatGPT (p=0.01). All models scored high in empathy (2). There were no significant differences in completeness and stability between models. ChatGPT had the lowest average readability, with all models' reading levels between grades seven to nine (7.64 vs 9.15 vs 8.82, p=0.01). Direct guideline questions yielded higher accuracy than nuanced questions. LLMs offer inconsistently accurate but empathetic infant safe sleep advice with frequent deviations from AAP recommendations. Pediatric oversight and collaboration with technology developers are essential to ensure safe, evidence-based information for families.
How do the three most widely accessible large language models perform in terms of accuracy, consistency, and reliability when answering clinically relevant questions derived from the 2022 ESHRE guideline on endometriosis? Model A achieved the highest accuracy scores, model C demonstrated significantly superior consistency across repeated queries, and all three models showed comparable but suboptimal reliability under free-tier access, while under official API access median accuracy converged across providers and reliability increased. Large language models are increasingly consulted by both clinicians and patients as readily accessible sources of medical information. In reproductive medicine, preliminary evidence suggests that individual platforms may retrieve endometriosis-related content with mixed fidelity. However, no study has simultaneously benchmarked multiple models against a single, internationally recognized endometriosis guideline. The extent to which these tools can be trusted to faithfully reproduce evidence-based recommendations on endometriosis diagnosis and management remains largely unexplored. Cross-sectional, multi-platform comparative study. Fifty clinically relevant questions covering the diagnostic and therapeutic domains of the 2022 ESHRE endometriosis guideline were simultaneously submitted to all three models during December 2025. Each question was entered in duplicate using independent sessions to assess response consistency and reliability. Subgroup analyses were carried out by submitting the questions to the API version and to the free-tier version of the platforms available in May 2026. The three models were accessed through their respective free web interfaces using new accounts, without prompt engineering, prior training, or retrieval-augmented generation. A zero-shot prompting approach was adopted. Accuracy was evaluated by two independent experts using the Global Quality Score (GQS); consistency was defined as identical responses across the three iterations; reliability was defined as the alignment of each response with the ESHRE guideline. Discrepancies were settled by a third reviewer. Significant differences in accuracy were observed across models (Kruskal-Wallis H = 37.10, P < 0.001). Model A achieved the highest median GQS (5, interquartile range [IQR] 4-5), followed by model C (4, IQR 3-5) and model B (3, IQR 3-4). Post-hoc analysis confirmed that model A significantly outperformed model B (P < 0.001) but not model C (P = 0.123), while model C also scored significantly higher than model B (P < 0.001). For consistency, model C demonstrated a significantly higher rate of reproducible responses (92.0%) compared with model A (72.0%, P = 0.028) and model B (68.0%, P = 0.008). No significant between-model differences were found for reliability (χ2 = 1.029, P = 0.598), with rates of 76.0% for model C, 68.0% for model A, and 68.0% for model B. In API and free-tier May 2026 subgroup analyses, median GQS converged to 4 across all three providers, and reliability rose to at least 76% for each provider. Model outputs may change with subsequent updates. GQS retains a degree of subjectivity despite expert adjudication. The study tested factual recall rather than complex clinical reasoning, limiting generalizability to real-world decision-making scenarios. These findings provide the first multi-platform benchmark of large language models against the latest endometriosis guideline. While models A and C retrieved guideline-concordant information with acceptable fidelity, none of the models achieved a level of reliability required for unsupervised clinical use. The dissociation between accuracy and consistency at the free tier, and its attenuation under API and more recent accesses, indicates that both model capability and commercial tier shape user-facing outputs. The dissociation between accuracy and consistency underscores that a model producing high-quality answers does not necessarily do so in a reproducible manner. Expert oversight remains crucial when interpreting their output regarding endometriosis care. Future research should extend this framework to additional endometriosis guidelines and to longitudinal monitoring of models' performance with their upgrades. No external funding was received for this study. The authors declare no competing interests. N/A.
Large Language Models (LLMs) are emerging as potential tools for health communication and patient education. However, their ability to translate complex medical guidelines into accessible, safe, and accurate messages for older adults remains insufficiently evaluated. To assess the capacity of ChatGPT-4, Claude AI, and Deepseek to generate exercise awareness messages for older adults, aligned with international expert consensus. This was a cross-sectional observational study conducted between April and August 2025. Using the ICFSR 2021 Global Consensus on optimal exercise recommendations as the gold standard, we evaluated messages generated by three LLMs for 13 common chronic conditions in older adults. A standardized prompt was used to generate messages addressing exercise prescription considerations, disease progression, and recommended modalities. Two independent expert evaluators; a geriatrician and a sports medicine physician assessed five dimensions using a 5-point Likert scale; accuracy, clarity, safety, behavioral relevance, and absence of fabrication. Readability was measured using Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG). Inter-rater agreement was assessed using intraclass correlation coefficient (ICC). Claude achieved the overall scores from both evaluators at 4.63 ± 0.55 and 4.60 ± 0.55, followed by ChatGPT-4 at 4.55 ± 0.50 and 4.48 ± 0.53 and Deepseek at 4.15 ± 0.87 and 4.17 ± 0.86. Inter-rater agreement was moderate was 0.668. Claude demonstrated accuracy scores at 4.58 ± 0.64, while ChatGPT-4 excelled in clarity with 5 ± 0. All models achieved perfect scores for absence of fabrication (5 ± 0). Readability indices revealed high complexity across all LLMs, with median FKGL values ranging from 10.84 to 10.93, corresponding to 10th-11th grade reading level, exceeding recommended levels for older populations. A significant correlation was found between Claude's accuracy scores and FKGL (r = 0.599, p = 0.030). LLMs, particularly Claude, ChatGPT-4 and Deepseek demonstrate strong potential for generating accurate, safe, and hallucination-free exercise awareness messages for older adults. However, readability remains above recommended levels, requiring optimization.
While AI coding tools have demonstrated potential to accelerate software development, their use in scientific computing raises critical questions about code quality and scientific validity. In this paper, we provide twelve practical tips for AI-assisted coding that balance the capabilities of AI with the demands of scientific and methodological rigor. We address how AI can be leveraged strategically throughout the development cycle with four key themes: problem preparation and understanding, managing context and interaction, testing and validation, and code quality assurance and iterative improvement. These principles serve to emphasize maintaining human agency in coding decisions, establishing robust validation procedures, and preserving the domain expertise essential for methodologically sound research. These tips are intended to help researchers harness AI's transformative potential for faster software development while ensuring that their code meets the standards of reliability, reproducibility, and scientific validity that research integrity demands.
The rapid adoption of Large Language Models (LLMs) in educational assessment has reshaped scoring practices, yet evaluation remains tethered to aggregate reliability metrics like Quadratic Weighted Kappa, which obscure discrimination and rater effects. This study applies Signal Detection Theory to evaluate eight state-of-the-art LLMs (including Claude 3.5 Haiku, DeepSeek-V3, Gemini 3 Flash, GPT-4o, and Grok 4.1) against expert human raters across 1,726 essays. By decoupling discrimination from response criteria, I provide a diagnostic analysis of AI scoring behavior. Results indicate that human raters exhibit significantly superior evaluative precision, with average discrimination estimates approximately double those of the AI models. Furthermore, LLMs are prone to pronounced centrality effects and score compression, systematically failing to award the highest rubric tiers. These findings demonstrate that low human-machine agreement stems from both a deficit in discriminative accuracy and systematic shifts in response criteria. Ultimately, this research provides a robust framework for calibrating and selecting AI scoring systems based on specific pedagogical goals and fairness requirements.
The potential for large language models (LLMs) to standardize clinical decision-making in ARF, particularly in optimizing high-flow nasal cannula (HFNC) therapy application, is gaining interest due to the condition's complex etiology. This study aimed to evaluate the clinical utility of artificial intelligence (AI) models in aligning with the European Respiratory Society (ERS) clinical practice guidelines for the use of HFNC in acute respiratory failure (ARF). Eight advanced LLMs were assessed across eight representative clinical scenarios related to HFNC use in ARF. Each model's responses were independently scored by three independent reviewers across four domains: accuracy, overconclusiveness, supplementary value, and completeness, using a 5-point Likert scale. Inter-rater reliability was measured by Fleiss' Kappa. Readability metrics, including Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL), were also analyzed. No differences were detected in accuracy or overconclusiveness (both P>0.99). Significant differences emerged in supplementary value among models (P<0.0001). DeepSeek-V3.1 ranked highest, significantly outperforming Claude and ChatGPT models (P<0.05). For completeness, ChatGPT-4o scored lower than both DeepSeek-V3.1 (P=0.006) and DeepSeek-R1 (P=0.01). Readability analyses revealed that DeepSeek-V3.1 achieved the highest FRE score, while DeepSeek-R1 had the lowest FKGL, reflecting better textual accessibility. Inter-rater reliability was substantial (κ=0.781). Evaluations of LLMs for HFNC guideline interpretation show that while all are accurate, DeepSeek-V3.1 and R1 excel in completeness, supplemental detail, and readability, marking them as more suitable for ARF clinical decision support. Their integration into critical workflows still demands expert supervision.
This study evaluated whether otolaryngologists can distinguish between human- and machine-written abstracts. The primary question was whether large language models (LLMs) produce abstracts comparable in clarity and usefulness to human-authored work, and whether reviewers can identify authorship with accuracy. A blinded cross-sectional design was used. Forty-eight abstracts were evaluated, consisting of twenty-four human-authored abstracts and 24 generated by four LLMs. Human abstracts were drawn from articles published after July 2025 to minimize overlap with LLM training data. Twenty otolaryngologists independently reviewed all abstracts. Using a structured rubric, raters classified authorship, rated clarity, usefulness, and confidence on 5-point scales, and provided optional free-text explanations. Group comparisons were performed using chi-square and Mann-Whitney tests, with Kruskal-Wallis tests for model-level analyses. Overall recognition accuracy was 44.7%. Human-written abstracts were more often misclassified as AI than AI-generated abstracts were mistaken for human. Human abstracts received significantly higher clarity and usefulness scores than LLM abstracts, though effect sizes were small. Confidence did not correlate with correctness, indicating miscalibration of rater judgments. Model-level performance varied. Grok-generated abstracts were most easily identified as AI, whereas GPT-5 and Claude 3.5 more frequently resembled human writing. Free-text rationales commonly referenced style, vagueness, or lack of detail when AI authorship was suspected. LLMs generate abstracts that increasingly resemble human scientific writing, yet still lag in perceived usefulness and credibility. Clinicians were only moderately successful at detecting authorship and were frequently confident in incorrect classifications. These findings highlight both the promise and risks of AI-assisted scientific communication.
Robot-assisted radical cystectomy (RARC) is a complex procedure that requires patients to understand surgical indications, urinary diversion, perioperative treatment, complications, recovery, and long-term functional outcomes. Although artificial intelligence (AI) chatbots are increasingly used to obtain medical information, their suitability for RARC patient education remains unclear. We conducted a cross-sectional comparative evaluation of four contemporary AI chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro. A set of 20 core patient-education questions on RARC was developed by three senior urologic experts. Chatbot responses were assessed using DISCERN, the Ensuring Quality Information for Patients tool, the Global Quality Scale, and JAMA benchmark criteria. Readability was evaluated using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, and SMOG. Reliability scores differed significantly across models for DISCERN, EQIP, and GQS, while JAMA benchmark criteria were summarized descriptively as transparency signals. DeepSeek-V4 achieved the highest mean scores for DISCERN, EQIP, and GQS, while ChatGPT-5 and DeepSeek-V4 showed the strongest JAMA benchmark performance. Gemini 3.5 Pro generally had the lowest reliability and transparency scores. Readability also varied across models. DeepSeek-V4 produced the most readable responses overall, whereas Gemini 3.5 Pro generated the most complex text. However, all models exceeded the recommended sixth-grade reading level, and FRES scores remained below the recommended threshold. Contemporary AI chatbots generated responses with variable presentation quality, transparency, and readability for common RARC patient-education questions. Because factual accuracy was not directly assessed, these tools should not be interpreted as validated sources of clinical guidance and should not replace individualized counseling by urologists.
Clinical decision-making requires integrating history, physical examination, laboratory, and imaging data. In the emergency department (ED), workload, time pressure, and cognitive burden may impair this process and affect decision quality. This study compares the diagnostic outputs of ChatGPT, Claude, and Gemini with those of emergency physicians in real-world ED cases. This prospective, single-centre observational diagnostic agreement study compared the stage-wise outputs of four Large Language Models (LLMs) (ChatGPT-4o, ChatGPT-5, Claude Opus 4.1, and Gemini 2.5 Pro) with those of emergency physicians in critically ill ED patients. Between 10 August and 10 September 2025, de-identified clinical data were entered into the models via their official web interfaces using standardised prompts. In the first stage, physicians and LLMs each generated five preliminary diagnoses based on vital signs and medical history. In the second stage, following physical examination and laboratory and imaging results, both refined their lists into three differential diagnoses. In the third stage, the physicians' final diagnosis was accepted as the reference, and each LLM was prompted to provide a final diagnosis. LLM preliminary and differential diagnoses were compared with those of the physicians at the corresponding stage, and LLM final diagnoses with the reference; the inclusion of the final diagnosis within earlier lists was also evaluated. Agreement was quantified using Cohen's κ; analyses were performed in R. Of 389 screened patients, 180 were included (56.1% male; mean age 67 ± 15.9 years). Physicians contained the reference diagnosis within their top-5 preliminary and top-3 differential lists in 83.9% and 98.3% of cases, respectively, significantly exceeding every LLM (all p < 0.001). Final-diagnosis match rates were 67.2% [60.3-73.5] for ChatGPT-4o, 65.6% [58.7-71.9] for ChatGPT-5, 63.3% [56.3-69.9] for Claude Opus 4.1, and 59.4% [52.3-66.1] for Gemini 2.5 Pro (p = 0.16). Cohen's κ ranged from 0.575 (Gemini 2.5 Pro) to 0.656 (ChatGPT-4o), indicating moderate-to-substantial agreement, with no pairwise difference reaching significance. The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases. Despite final-diagnosis match rates of 59%-67%, their current diagnostic role in the ED remains limited.