With the growing integration of generative Artificial Intelligence (AI) into healthcare, the DeepSeek large language model has emerged as a versatile tool for medical students, offering adaptive learning solutions tailored to various medical scenarios. However, the adoption of AI in medicine also raises ethical concerns that require warrant consideration. This study aimed to investigate medical students' usage patterns, attitudes, and influencing factors related to DeepSeek. In July 2025, a cross-sectional survey was conducted using random sampling among 1000 medical students from institutions affiliated with Anhui Medical University. Based on previous AI attitude scales, a validated self-administered questionnaire was used to collect data on students' attitudes, DeepSeek usage patterns. Among the 937 valid responses (response rate: 93.7%), 874 participants were aware of DeepSeek prior to the survey and 765 had used it. The Cronbach's alpha value for the questionnaire was 0.83. 42.0% of students doubted the accuracy of information provided by DeepSeek, 48.6% were apprehensive about potential plagiarism accusations, and 43.0% worried about over-reliance on AI. Additionally, 46.6% found DeepSeek interesting and appealing, with 37.7% expressing enthusiasm for learning new AI technologies. Several factors were significantly associated with usage experience, including being female [β = 6.1(1.5-10.7), (p = 0.031)], holding a Master's degree [β = 10.5(2.1-18.1), (p = 0.044)], and engaging in clinical [β = 7.4(1.2-13.6), (p = 0.021)] or basic research [β = 9.2(2.9-15.4), (p = 0.012)]. Regarding attitudes, significant predictors included a Master's degree [β = -2.9(-6.2 to-0.5), (p = 0.011)], having a stomatology background [β = 8.0(2.3-13.7), (p = 0.035)], engaging in clinical [β = -2.6(-5.0 to -0.1), (p = 0.036)] or basic [β = -2.7(-5.2 to -0.3), (p = 0.042)] research. Among the medical students surveyed in this study, 94.8% reported that they would be willing to use and recommend DeepSeek, but they maintain a cautious attitude toward its privacy, reliability, and dependency. Despite these concerns, 90.2% of participants said they would be willing to use DeepSeek to help them complete their tasks.
Esophageal cancer remains a significant global health issue. ChatGPT-4o and DeepSeek V3 can provide the public with health-related knowledge about esophageal cancer. This study aimed to evaluate the accuracy of DeepSeek V3 and ChatGPT-4o in responding to health knowledge questions related to esophageal cancer. Fifty-two questions related to esophageal cancer were classified into themes of basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis and patient frequently asked questions (FAQs). These questions were entered into DeepSeek V3 and ChatGPT-4o to obtain responses, and 2 experienced gastroenterologists independently evaluated the accuracy and temporal stability of each response. Overall, the scores of DeepSeek V3 and ChatGPT-4o on all questions were 4 (3-4), and there was no statistically significant difference between the 2 groups. The final scores of DeepSeek V3 in basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis, and FAQs were 4 (3-4), 4 (3-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively, while the scores of ChatGPT-4o were 4 (3-4), 3 (2-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively. For temporal stability across 2 independent test runs, DeepSeek V3 presented inconsistent responses on 2 questions, and ChatGPT-4o on 1 question; no statistically significant differences were found in overall and subgroup scores between the 2 runs for both models (all P > .05). ChatGPT-4o and DeepSeek V3 showed favorable accuracy and comprehensive responses to most of the 52 esophageal cancer-related questions in this study, but our findings do not confirm their general reliability for esophageal cancer health information in routine clinical or public use.
Large language models are increasingly applied in medical education and clinical decision support. However, comparative evaluations of their performance on pediatric content-spanning multiple subspecialties and question complexities-remain limited. This study sought to assess the performance of two advanced large language models, DeepSeek-R1 and GPT-4o, using a comprehensive set of Pediatrician Licensing Exam practice questions. We administered 280 expert-validated pediatric questions covering 11 subspecialties and three levels of clinical complexity (A1: direct, A2: simple cases, A3: complex cases) to DeepSeek-R1 and GPT-4o. Each model was tested twice, with a 4-week interval. Performance metrics, including per-run accuracy, consistent accuracy (correct in both runs), aggregate accuracy (correct in either run), and inter-run consistency, were compared using chi-square tests and Bonferroni correction. DeepSeek-R1 outperformed GPT-4o in both runs (1st run: 90.7% vs. 81.8%, P=0.003; 2nd run: 86.8% vs. 79.3%, P = 0.02). It also showed higher aggregate accuracy (92.5% vs. 85.7%, P=0.01) and consistent accuracy (85.0% vs. 75.4%, P = 0.01). Performance declined with increasing question complexity for both models, with DeepSeek-R1 demonstrating numerically higher accuracy across all question types. For consistent accuracy, DeepSeek-R1 achieved 89.2% vs. 79.2% on A1 questions (P=0.05), 83.8% vs. 71.6% on A2 questions (P=0.11), and 80.2% vs. 73.3% on A3 questions (P=0.36). At the subspecialty level, DeepSeek-R1 showed more uniform performance, achieving >90% consistent accuracy in four subspecialties: Neonatology (94.7%), Infectious Diseases (93.8%), Nephrology (91.9%), and Cardiovascular diseases (90%). In comparison, GPT-4o reached >90% consistent accuracy in three subspecialties-Genetics (100%), Infectious Diseases (100%), and Cardiovascular diseases (90%)-with greater variability across other domains. DeepSeek-R1 demonstrated higher accuracy than GPT-4o on pediatric questions. The observed variability across different runs warrants attention, particularly in contexts that require highly consistent and reproducible performance.
Modern large language models (LLMs) like ChatGPT (based on the GPT-4 architecture) and DeepSeek offer unprecedented capabilities for generating scientific text. However, their performance in replicating structured, high-quality scientific writing, especially compared to human-authored abstracts, remains insufficiently evaluated. To compare the abstract quality produced by human authors, ChatGPT/GPT4, and DeepSeek across: six evaluation criteria Clarity, Coherence, Conciseness, Accuracy, IMRaD Structure, and Language Quality, using blinded expert ratings and non-parametric statistical methods, specifically the Kruskal-Wallis test followed by pairwise Wilcoxon rank-sum tests with false discovery rate correction. We selected 23 medical and healthrelated research topics, each yielding three abstracts (human, ChatGPT, DeepSeek), for a total of 69 abstracts. Three raters scored each abstract. Kruskal-Wallis tests assessed group differences; Cliff's Delta (δ) was calculated as a nonparametric effect size for each comparison, suitable for ordinal data. Across criteria, ChatGPT and DeepSeek significantly outperformed human authors in Clarity, Coherence, IMRaD Structure, and Language Quality. In contrast, Conciseness and Accuracy showed negligible effect sizes (|δ| <0.10), suggesting parity across all three sources. ChatGPT and DeepSeek achieved significantly higher scores in clarity, coherence, structure, and language quality, while showing comparable performance in conciseness and accuracy. These findings complement recent evaluations showing competitive medical and reasoning performance of DeepSeek models compared to proprietary LLMs. While shortform abstracts, expert oversight, and domain expertise remain critical, the results suggest that LLMs-particularly GPT4 and DeepSeek-can serve as effective tools in drafting scientific abstracts.
The integration of artificial intelligence (AI) tools like DeepSeek into scientific research offers new opportunities for efficiency and innovation. However, attitudes and adoption among medical students remain underexplored. This study aims to investigate medical students' attitudes, usage patterns, and interest in learning with DeepSeek. A cross-sectional online survey was conducted among medical students from various academic levels and fields. The questionnaire assessed demographics, attitudes, usage frequency and purposes, and learning interests. Data were analyzed using descriptive statistics and independent sample t-tests. Among 589 respondents, most were familiar with DeepSeek (86.08%), and 70.29% used it. A majority held positive attitudes toward DeepSeek, agreeing that it is useful for academic success (77.91%), makes research easier (78.37%), and will play an important future role (74.51%). However, concerns included reliability (61.34%), overreliance (48.51%), and privacy risks (51.88%). Usage varied: 28.23% used DeepSeek frequently, while 29.76% had never used it. Common uses included problem-solving (406 users) and literature search (256 users). Over 80% expressed interest in learning to use DeepSeek more effectively. No significant differences in attitudes were found across demographic groups. Medical students view DeepSeek as a promising tool for academic and research support but have significant concerns regarding reliability and ethical use. There is strong interest in structured learning opportunities. Integrating AI literacy into medical education and providing targeted training are recommended to promote responsible and effective use.
Background: Large language models (LLMs) are increasingly used by patients for obtaining medical information; however, concerns remain regarding their clinical safety, reliability, and appropriateness of patient guidance. Evidence evaluating LLM performance in hemorrhoid-related patient questions remains limited. Objective: To compare the clinical accuracy, safety, and overall clinical adequacy of responses generated by ChatGPT, Gemini, and DeepSeek to hemorrhoid-related patient questions. Methods: In this cross-sectional comparative study, 25 hemorrhoid-related patient questions were developed and categorized into three predefined subgroups: basic informational questions, clinically significant scenarios, and misleading/risky patient statements. Responses generated by ChatGPT (GPT-5.3), Gemini (3.1), and DeepSeek (R1) were evaluated by two experienced surgeons using a consensus-based expert assessment approach and a structured 5-point scoring system assessing clinical accuracy, safety, appropriateness of patient guidance, and overall clinical adequacy. Critical errors and qualitative communication characteristics were also analyzed. Friedman and post hoc Conover tests with Bonferroni correction were used for statistical comparisons. Results: Overall response quality differed significantly among models (χ2(2) = 29.119, p < 0.001, Kendall's W = 0.582). ChatGPT achieved the highest overall scores (5.00 ± 0.00), followed by Gemini (4.80 ± 0.41) and DeepSeek (4.12 ± 0.67). Significant differences were primarily observed between DeepSeek and the other models, whereas ChatGPT and Gemini showed comparable performance. Model divergence became more pronounced in clinically significant scenarios involving alarm symptoms, rectal bleeding, persistent symptoms, and acute anorectal pain. No model generated directly harmful medical recommendations or explicit guidance likely to result in substantial diagnostic delay. However, qualitative assessment demonstrated differences in communication style and risk communication. ChatGPT generally produced more balanced and context-appropriate responses, Gemini generated more explanatory responses, whereas DeepSeek showed a tendency toward disproportionately urgent or alarmist language in some higher-risk scenarios. Conclusions: Large language models demonstrated generally high clinical accuracy in answering hemorrhoid-related patient questions; however, notable model-specific differences were observed in clinical guidance, communication style, and risk communication, particularly in high-risk clinical scenarios. These findings suggest that LLMs may serve as useful supportive tools for patient education and health information delivery, although they should currently be regarded as systems that support rather than replace human clinical judgment.
Generative AI is becoming widely used in medical education, but what drives medical students to continue using domestic tools such as DeepSeek after initial adoption remains poorly understood. This study examined the psychological mechanisms underlying medical students continued use of DeepSeek, with particular attention to how cognitive appraisals, satisfaction, and task fit shaped continuance in real learning contexts. An explanatory sequential mixed-methods design was used. In the quantitative phase, 630 valid questionnaires were analyzed using structural equation modeling to test a continuance pathway centered on cognitive appraisal, satisfaction, and behavioral intention, while interview data were used to explain unexpected and nonsignificant quantitative findings. System quality and subjective norm positively affected perceived ease of use, while subjective norm and expectation confirmation positively affected perceived usefulness. Perceived ease of use and perceived usefulness both increased satisfaction, and satisfaction was the strongest predictor of continuance intention. Task-technology fit also positively influenced continuance intention, which strongly predicted actual continued use. Technology characteristics and task characteristics both improved task-technology fit. By contrast, information quality negatively affected perceived ease of use, subjective norm negatively affected satisfaction, and privacy concerns and expectation confirmation did not significantly affect continuance intention or satisfaction. Students mainly continued using DeepSeek because it was easy to access, helpful for academic writing and exam preparation, and suited to some specialized tasks; their primary concerns were unstable performance, inaccurate outputs, future pricing, and data security. Continued DeepSeek use followed a cognitive-affective-behavioral sequence: perceived ease of use and usefulness drove satisfaction, which in turn predicted continuance intention (β = 0.769), while task-technology fit provided an independent behavioral pathway (β = 0.157), together accounting for actual continued use (β = 0.732).
Large language model (LLM) are increasingly explored for oncology decision support, yet their alignment with real-world clinical practice across varying disease complexities remains insufficiently characterized. This study aimed to evaluate and compare the accuracy, stability, and concordance of two advanced LLMs-DeepSeek V3.1 and ChatGPT-5-against experienced oncologists in generating breast cancer treatment plans within a specific clinical setting. This retrospective study compared the performance of DeepSeek V3.1 and ChatGPT-5 with senior oncologists using de-identified records from 213 breast cancer patients (Stages I-IV). To assess effectiveness, we implemented a multidimensional evaluation framework: accuracy was measured using a 5-point Likert scale by three independent, blinded expert reviewers; internal consistency was quantified via variance and coefficient of variation; and clinical concordance was evaluated using a structured five-level scoring system. Statistical analyses, including ANOVA and ordinal regression, were used to examine the impact of disease stage on AI-human agreement. Under standardized retrospective review conditions, LLM-generated recommendations demonstrated higher expert-rated guideline concordance and lower variability than historical real-world oncologist plans. Specifically, DeepSeek V3.1 achieved the highest expert-rated accuracy scores with minimal internal variance (4.91 ± 0.36), outperforming both ChatGPT-5 (4.65 ± 0.62) and clinicians (3.82 ± 0.63, P < 0.001). While AI outputs exhibited high mutual consistency (74.2%), expert evaluations revealed a significant decline in AI-clinician agreement as disease stage advanced (P < 0.001), particularly in Stage IV cases where clinicians prioritized real-world constraints such as financial toxicity. Advanced LLMs, particularly DeepSeek V3.1, demonstrated strong performance in generating standardized, guideline-concordant breast cancer treatment plans, showing superior consistency over human specialists in protocol-driven scenarios. However, the widening gap in complex late-stage cases highlights limitations in accounting for clinical context and socioeconomic factors. These findings support the role of LLMs as robust clinician-supervised decision-support tools while emphasizing the necessity of human judgment for individualized care.
Artificial intelligence (AI) is already showing enormous potential in the healthcare sector. Generative AI, particularly, is accelerating the sector's digital transformation by delivering intelligent decision support, automated diagnosis, and optimized resource allocation. DeepSeek-R1, a large-language model with a strong performance-to-cost ratio, has gained popularity as a foundation model in Chinese hospitals. However, the deployment of generative AI remains challenging, and hospitals continue to lack clear guidance on how to select among deployment architectures and how to balance computational demand with cost. Deeper, data-driven analysis is therefore warranted to inform future roll-outs. This study surveyed AI deployment across Chinese hospitals, with a focus on DeepSeek's potential applications under national policies. Data were collected from the top 20 hospitals, regional centers, and township hospitals via official WeChat platforms. The survey examined deployment strategies, model versions, and platform choices, while keeping in account hospital needs, data resources, and technological-economic decisions. The study highlights DeepSeek's impact on diagnostic accuracy, personalized treatment, medical documentation automation, and resource management optimization. Among the 17 surveyed hospitals, 6 hospitals employed detailed model versions, 5 used the 671B model, and 1 used the 32B version. Among 10 hospitals of different levels, 2 selected the 671B, 3 selected the 70B, and 4 selected the 32B model. All hospitals preferred local deployment. Different needs and applications were observed across the studied hospitals. Selection of the right AI model requires balancing computational power with cost. Larger models offer higher accuracy, but incur higher costs, whereas distilled models suit smaller hospitals with fewer resources. Future development should therefore focus on selecting deployment strategies based on the hospital size while addressing data quality disparities to bridge the regional healthcare gaps. As such, coordination among government, hospitals, and doctors is crucial for supporting smarter healthcare transitions.
Large language models (LLMs) are increasingly used to provide medical information, yet their performance in rare pediatric cancers remains largely unexplored. This study aimed to compare the clinical accuracy, comprehensiveness, and communication quality of five widely used LLMs in answering frequently asked questions about Ewing sarcoma. Twelve representative questions covering diagnosis, treatment, prognosis, and psychosocial support were presented to ChatGPT (GPT-5.2), Claude Sonnet 4.5, Gemini 3, DeepSeek V3.2, and Grok 4. Two orthopedic oncology specialists independently evaluated each response using a 4-point Likert scale assessing clinical accuracy, completeness, clarity, and relevance. Qualitative assessments of empathy and communication quality were also performed. Statistical analyses included the Friedman, Wilcoxon signed-rank, Kruskal-Wallis, and Mann-Whitney U tests. Significant differences were observed among the five LLMs (p < 0.001). ChatGPT achieved the highest overall performance, followed by Claude and DeepSeek. DeepSeek demonstrated the greatest technical accuracy but lower communication quality, whereas ChatGPT provided the best balance between factual correctness and patient-friendly communication. Gemini and Grok produced more superficial responses with lower overall scores. Current LLMs can support patient and family education in Ewing sarcoma but should not replace specialist consultation. Although ChatGPT and Claude demonstrated the most reliable overall performance, variability among models remains substantial. Further validation and disease-specific optimization are required before routine implementation in clinical practice.
Adult type 1 diabetes mellitus (T1DM) involves complex diagnosis, treatment, and long-term self-management, creating a need for accurate and accessible health education. Large language models (LLMs) are increasingly used for medical information seeking, yet their accuracy and consistency in adult T1DM-related queries remain insufficiently evaluated. A guideline-based comparative evaluation assessed DeepSeek-V3.2 and ChatGPT-5.0 using 22 English-language prompts derived from the 2021 ADA/EASD consensus report, covering basic knowledge, diagnosis and differential diagnosis, treatment, and complications. The prompts were submitted to both models twice, two weeks apart. Responses were independently evaluated by two blinded endocrinology specialists using a predefined four-point scoring rubric, with disagreements adjudicated by a third senior endocrinologist. Short-term consistency was assessed by expert judgment and TF-IDF cosine similarity. Inter-rater agreement was good (Cohen's κ = 0.71). Expert-judged consistency was 95.45% (21/22) for both models; TF-IDF cosine similarity was 0.52 ± 0.09 for DeepSeek and 0.54 ± 0.09 for ChatGPT. Overall accuracy scores were 3.59 ± 0.59 and 3.77 ± 0.43, respectively, with no statistically significant difference (p = 0.102). Comprehensive ratings accounted for 63.64% and 77.27%, respectively, and mixed correct and incorrect or outdated information accounted for 4.55% and 0.00%. Both models may support adult T1DM-related health education, but outputs should be interpreted as supplementary educational material under professional guidance rather than as diagnostic or therapeutic advice.
The use of artificial intelligence (AI) in orthodontic practice is increasing rapidly; however, there is a notable lack of research evaluating the accuracy of large language models (LLMs) in educating patients about orthodontic retainers and related guidelines. This study utilized a cross-sectional, repeated‑measures comparative evaluation design after receiving exemption from the institutional ethics committee. A set of 110 questions related to orthodontic retainers was compiled from previous articles addressing concerns about retainers and approved by a panel of three orthodontists. These questions were submitted to large language models (LLMs), including ChatGPT, Copilot, DeepSeek and Google Gemini. The responses were then reviewed by six independent orthodontists, who rated them using a modified five-point Likert scale. The overall accuracy revealed that 68.6% of responses scored 4, while 15.3% achieved a perfect score of 5. Among the LLMs, Gemini ranked first with 96.8%, closely followed by ChatGPT at 95.6%, indicating comparable high‑level performance between these models, while DeepSeek (76.9%) and Copilot (66.2%) demonstrated comparatively lower accuracy. Gemini produced a higher proportion of perfect scores, whereas ChatGPT consistently achieved strong ratings. The mean ratings across six raters demonstrated strong reliability (ICC = 0.81), reflecting expert agreement. Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions. However, specialist oversight remains essential to ensure clinical applicability. Future research using larger and more diverse datasets is needed to assess broader educational and communication‑related outcomes.
BackgroundPosttraumatic stress disorder (PTSD) is common yet frequently underdiagnosed, in part due to barriers to systematic screening and the reliance on self-report instruments. Large language models (LLMs) have shown promise in extracting clinically relevant information from unstructured language, but their ability to infer item-level PTSD symptom severity from clinical interviews remains unclear.MethodsUsing the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WoZ), we analyzed 100 semi-structured clinical interview transcripts paired with item-level PTSD Checklist-Civilian Version (PCL-C) scores. Six LLMs (DeepSeek 3.1, Claude Sonnet 4, LLaMA 4 Scout, GPT-4o, GPT-5, and Gemini 2.5 Flash) used zero-shot prompting to predict all 17 PCL-C items. Performance was assessed for binary symptom endorsement (≥3 vs. < 3), 5-point Likert prediction, and DSM-IV symptom-cluster analyses using accuracy, F1 score, and Matthews correlation coefficient (MCC).ResultsFor binary prediction, Claude 4 achieved the highest mean accuracy (0.705; 95% CI, 0.681-0.728), followed by DeepSeek 3.1(0.699; 95% CI, 0.675-0.724) and Gemini 2.5 (0.698; 95% CI, 0.677-0.718). For Likert prediction, DeepSeek 3.1 performed best (accuracy = 0.438; 95% CI, 0.401-0.475), only modestly above the majority-class baseline (0.399; 95% CI, 0.355-0.443). Performance varied by symptom domain, with re-experiencing and hyperarousal symptoms generally predicted more accurately than avoidance/numbing symptoms. Across models, predicted item-level symptom patterns showed a meaningful alignment with observed PCL-C responses despite reduced accuracy in fine-grained severity estimation.ConclusionZero-shot LLMs' performance was insufficient for clinical application in predicting PTSD symptoms from semi-structured interview transcripts. While models showed some ability to capture overall symptom patterns, performance varied across domains and remained limited for fine-grained severity estimation. Given these constraints and the non-trauma-specific nature of the dataset, findings should be interpreted as preliminary, with only modest differences observed between models.Plain Language Summary TitleCan Artificial Intelligence Identify PTSD Symptoms from Conversations? A Study Using Clinical Interview TranscriptsPlain Language SummaryPost-traumatic stress disorder (PTSD) is a common mental health condition, but it is often missed in clinical settings. Screening usually relies on questionnaires that patients must complete themselves, which may not always happen due to time, stigma, or discomfort discussing trauma. Researchers are exploring whether artificial intelligence (AI) could help identify PTSD symptoms from conversations instead.In this study, we tested several advanced AI systems, known as large language models, to see if they could estimate PTSD symptoms based on written transcripts of clinical interviews. These interviews were not specifically designed to assess trauma, which makes the task more challenging but closer to real-world situations. We compared the AI predictions to participants' own questionnaire responses about their symptoms.We found that the AI models were somewhat able to recognize general patterns of PTSD symptoms, especially more visible ones like sleep problems or distressing dreams. However, they struggled with more internal or less obvious symptoms, such as avoidance or emotional numbness. Overall, their accuracy was moderate and not reliable enough for clinical use, particularly when trying to estimate how severe symptoms were.Importantly, differences between the AI models were small, and none performed well enough to replace existing screening methods. These findings suggest that while AI may have future potential as a supportive tool, it is not yet ready to be used for diagnosing or screening PTSD on its own.Further research using better data, improved methods, and real clinical settings is needed before this approach could be considered for practical use. Le trouble de stress post-traumatique (TSPT) est fréquent, mais il est souvent sous-diagnostiqué, en partie en raison d'obstacles au dépistage systématique et de l'utilisation d'instruments d'autodéclaration. Les grands modèles de langage (LLM) se sont révélés prometteurs pour ce qui est d'extraire des renseignements cliniquement pertinents d'un langage non structuré, mais leur capacité à déduire la gravité des symptômes de TSPT selon des éléments à partir d'entrevues cliniques demeure incertaine. À l'aide du corpus d'entrevues Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WoZ), nous avons analysé 100 transcriptions d'entrevues cliniques semi-structurées appariées au score pour les éléments de la liste de vérification pour le TSPT - Version civile (PCL-C). Six GML (DeepSeek 3.1, Claude Sonnet 4, LLaMA 4 Scout, GPT-4o, GPT-5 et Gemini 2.5 Flash) ont utilisé des requêtes sans exemple pour prédire les 17 éléments de la PCL-C. Le rendement a été évalué en ce qui concerne l'approbation binaire des symptômes (≥ 3 p/r à < 3), la prédiction de l'échelle Likert à 5 points et les analyses du groupe des symptômes dans le DSM-IV selon l'exactitude, le F-score et le coefficient de corrélation de Matthews (MCC). Pour la prédiction binaire, Claude 4 a obtenu l'exactitude moyenne la plus élevée (0,705; IC à 95 % : de 0,681 à 0,728), suivi de DeepSeek 3.1 (0,699; IC à 95 % : de 0,675 à 0,724) et de Gemini 2.5 (0,698; IC à 95 % : de 0,677 à 0,718). Pour la prédiction de l'échelle de Likert, DeepSeek 3.1 a obtenu les meilleurs résultats (exactitude = 0,438; IC à 95 % : de 0,401 à 0,475), légèrement supérieurs aux valeurs initiales de la classe majoritaire (0,399; IC à 95 % : de 0,355 à 0,443). Le rendement variait en fonction du domaine de symptômes, les symptômes de reviviscence et d'hypervigilance étant généralement prédits avec plus de précision que les symptômes d'évitement ou d'émoussement. Dans l'ensemble des modèles, les schémas des symptômes prédits au niveau des éléments ont montré une harmonisation significative avec les réponses observées au questionnaire PCL-C, malgré une exactitude réduite dans l'estimation fine de la gravité. Le rendement des LLM sans exemple était insuffisant pour une application clinique dans la prédiction des symptômes de TSPT à partir de transcriptions d'entrevues semi-structurées. Bien que les modèles aient montré une certaine capacité à saisir les tendances globales des symptômes, le rendement variait d'un domaine à l'autre et demeurait limité pour l'estimation fine de la gravité. Compte tenu de ces contraintes et de la nature non traumatique de l'ensemble de données, les résultats doivent être considérés comme préliminaires, avec de légères différences observées entre les modèles.
The potential for large language models (LLMs) to standardize clinical decision-making in ARF, particularly in optimizing high-flow nasal cannula (HFNC) therapy application, is gaining interest due to the condition's complex etiology. This study aimed to evaluate the clinical utility of artificial intelligence (AI) models in aligning with the European Respiratory Society (ERS) clinical practice guidelines for the use of HFNC in acute respiratory failure (ARF). Eight advanced LLMs were assessed across eight representative clinical scenarios related to HFNC use in ARF. Each model's responses were independently scored by three independent reviewers across four domains: accuracy, overconclusiveness, supplementary value, and completeness, using a 5-point Likert scale. Inter-rater reliability was measured by Fleiss' Kappa. Readability metrics, including Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL), were also analyzed. No differences were detected in accuracy or overconclusiveness (both P>0.99). Significant differences emerged in supplementary value among models (P<0.0001). DeepSeek-V3.1 ranked highest, significantly outperforming Claude and ChatGPT models (P<0.05). For completeness, ChatGPT-4o scored lower than both DeepSeek-V3.1 (P=0.006) and DeepSeek-R1 (P=0.01). Readability analyses revealed that DeepSeek-V3.1 achieved the highest FRE score, while DeepSeek-R1 had the lowest FKGL, reflecting better textual accessibility. Inter-rater reliability was substantial (κ=0.781). Evaluations of LLMs for HFNC guideline interpretation show that while all are accurate, DeepSeek-V3.1 and R1 excel in completeness, supplemental detail, and readability, marking them as more suitable for ARF clinical decision support. Their integration into critical workflows still demands expert supervision.
Widespread antiretroviral therapy has greatly extended the life expectancy of people living with HIV (PLWH), making cardiovascular disease (CVD) one of their primary comorbidities. Nevertheless, significant cross-specialty knowledge gaps persist in routine clinical practice. Siloed disciplinary expertise results in low clinical adherence to guideline-recommended risk management interventions, highlighting an urgent demand for integrated, evidence-based tools that break down interdisciplinary barriers. Large language models (LLMs) have demonstrated robust medical knowledge retrieval and reasoning capacity in recent years, yet no systematic evaluation has determined whether these models can bridge such knowledge gaps and facilitate multidisciplinary collaborative CVD management for PLWH. This study compared the performance of four mainstream AI models (Deepseek-V3, Deepseek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for PLWH. Based on four authoritative domestic and international guidelines on HIV and CVD care, a structured 25-question assessment battery was developed via two rounds of Delphi expert consultation, with standard reference answers and an evaluation framework finalized through expert consensus. Responses of the four LLMs were generated with standardized prompts, while 12 clinicians answered identical questions in one-on-one structured interviews, with all verbal replies transcribed verbatim. Six multidisciplinary experts independently rated all responses across four dimensions: accuracy, completeness, readability and reliability, using a 4-point ordinal scale ranging from 1 (poor) to 4 (excellent). Cumulative link mixed models (CLMMs) were applied to analyze intergroup differences. All AI models achieved statistically significantly higher scores than clinicians across all evaluation dimensions (p < 0.01). The AI group had mean scores of 3.44-3.68 (median = 4, CV: 0.145-0.178). Restricted by individual factors including specialty background, knowledge reserve, clinical experience, clinicians obtained lower mean scores of 1.78-2.05 (median = 2, CV: 0.428-0.473) with markedly greater score dispersion. Among all AI models, Deepseek-R1 delivered the optimal performance and showed statistically significant advantages over ChatGPT-4o, ChatGPT-o4-mini and Deepseek-V3 (all p < 0.01). Specialty-stratified CLMM analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (OR = 0.92, 95% CI: 0.84-1.01, p = 0.094). Dimension-specific CLMMs combined with Wilcoxon rank-sum tests confirmed that cardiologists only earned significantly higher scores in the accuracy dimension (OR = 0.81, 95% CI: 0.67-0.97, p = 0.0261). Domain-specific performance divergence was observed: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists achieved higher scores on drug adverse effect evaluation (2.23 vs 1.65). This structured Q&A study on CVD management for PLWH found that LLMs outperformed human clinicians on all assessment metrics, with Deepseek-R1 attaining a distinctly superior composite score. The findings support the promising potential of Deepseek-R1 as a cross-disciplinary decision-support tool: it integrates multi-domain complex clinical knowledge, which may help address cross-specialty knowledge barriers, could improve the completeness and precision of clinical information output, and may enhance communication and decision-making efficiency for patients with complicated multimorbidity. To maximize clinical benefits, AI systems should be integrated into multidisciplinary care workflows alongside targeted clinical training to optimize the management of complex comorbidities among PLWH.
Large language models (LLMs) and large reasoning models (LRMs) have shown excellent performance on medical benchmarks, although evaluations concerning real-world medical workflows are still lacking. Lung cancer care is particularly dependent on the multidisciplinary team (MDT) integration of radiology, pathology, staging, and treatment planning, making it a high-bar setting for evaluating LRMs. This study aimed to compare the quality of recommendations generated by 2 LRMs (GPT-5-Thinking and Deepseek-v3-r1) with each other and with human MDT decisions in real-world lung cancer cases, as well as to assess whether MDT awareness of AI comparison influences the quality of MDT decisions. This was a single-center real-world comparative study of 100 consecutive lung cancer MDT cases (50 retrograde and 50 anterograde) from the University Hospital of Split, Croatia. For each case, deidentified structured reports (containing all necessary patient or case data, while excluding MDT conclusions) were submitted once to GPT-5-Thinking and Deepseek-v3-r1 to generate recommendations for radiologic diagnostics, pathologic diagnostics, oncologic therapy, and overall usefulness. Two independent lung oncologists graded MDT decisions and model outputs on 1-5 Likert scales. An average recommendation score (avg_rec) was calculated as the mean of radiology, pathology, and therapy scores. Analyses used Wilcoxon tests for paired model comparisons, Mann-Whitney tests for between-phase comparisons, and Spearman correlations (2-sided α=.05). Ratings showed ceiling effects. In the retrograde phase (N=50), the mean (95% CI) GPT-5-Thinking scores were higher than Deepseek-v3-r1 scores for radiologic diagnostics (4.89, 4.78-4.99 vs 4.76, 4.62-4.89; P<.001), oncologic therapy (4.82, 4.69-4.94 vs 4.18, 3.82-4.54; P<.001), and usefulness (4.82, 4.69-4.94 vs 4.18, 3.84-4.53; P<.001); pathologic diagnostics were similar (4.88, 4.78-4.97 vs 4.73, 4.57-4.90; P=.15). In the anterograde phase (n=50), the mean (95% CI) GPT-5-Thinking scores remained higher for radiology (4.94, 4.85-5.03 vs 4.64, 4.47-4.81; P<.001) and pathology (4.96, 4.90-5.02 vs 4.78, 4.65-4.91; P=.008), with smaller differences for therapy (4.46, 4.18-4.74 vs 4.20, 3.86-4.54; P=.24) and usefulness (4.50, 4.24-4.76 vs 4.16, 3.83-4.49; P=.12). The mean (95% CI) GPT-5-Thinking avg_rec exceeded MDT grade in both phases (retrograde: 4.90, 4.84-4.95 vs 4.14, 3.96-4.33; P<.001; anterograde: 4.79, 4.69-4.89 vs 4.34, 4.16-4.52; P<.001); Deepseek-v3-r1 exceeded MDT in the retrograde phase (4.56, 4.40-4.72 vs 4.14, 3.96-4.33; P<.001) but not the anterograde phase (4.54, 4.41-4.67 vs 4.34, 4.16-4.52; P=.15). MDT grades did not differ between phases (P=.13). In 100 real-world lung cancer MDT cases, both LRMs produced high-quality recommendations, with GPT-5-Thinking consistently outperforming Deepseek-v3-r1 and exceeding expert-graded MDT decision quality in both phases. MDT decision quality was unchanged by awareness of AI benchmarking. LRMs can thus generate recommendations comparable to or exceeding expert MDT decisions, though the single-center design and ceiling effects limit generalizability. Whether integrating such tools into MDT workflows improves clinical decisions warrants prospective study.
Large language models (LLMs) are increasingly used in health care, with emerging applications in clinical decision support and nursing education. However, evidence on their performance in nursing contexts, particularly in oncology nursing, remains limited. Given the complexity and high-risk nature of oncology care, it is important to evaluate the performance and clinical relevance of LLM-generated responses in oncology nursing contexts. This study aimed to compare the performance of LLMs in oncology nursing decision support tasks using standardized examination questions and case-based clinical scenarios and explore LLMs' potential applicability and current limitations in oncology nursing practice. A total of 33 case-based questions derived from 10 oncology nursing clinical scenarios in a nationally used training manual, along with standardized examination-oriented questions from a commercially published preparation book for the Chinese Nursing (Intermediate) Qualification Examination, were used to evaluate the performance of 5 LLMs (DeepSeek, Qwen, Spark-Desk, WiseDiag, and ChatGPT). All models generated responses using a standardized prompt. Two oncology nurses with more than 5 years of clinical experience independently rated the case-based responses using 3 evaluation dimensions: correctness, clarity, and conciseness. Interrater reliability was assessed using the quadratic weighted Cohen κ, intraclass correlation coefficient, and Spearman rank correlation coefficient. Differences among models were analyzed using the Kruskal-Wallis test with the Dunn post hoc test. In addition, examination performance was evaluated based on total score, accuracy rate, and completion efficiency. Interrater reliability analyses indicated moderate agreement between evaluators. The median correctness, clarity, and conciseness scores were as follows: 11.50 (IQR 10.50-12.00) for DeepSeek, 11.00 (IQR 10.50-12.00) for Qwen, 10.50 (IQR 9.50-11.50) for Spark-Desk, 10.00 (IQR 9.50-11.50) for WiseDiag, and 10.00 (IQR 9.00-11.50) for ChatGPT. The Kruskal-Wallis test indicated statistically significant differences among models (H=11.416; P<.05), with post hoc analysis showing a significant difference only between DeepSeek and ChatGPT (P<.05). In examination-based tasks, all models achieved passing performance, with accuracy rates ranging from 77% (77/100) to 93% (93/100). In terms of response completion, DeepSeek and ChatGPT completed all tasks in a single interaction, whereas other models required multiple interactions due to output interruptions. LLMs showed relatively strong performance on structured knowledge and examination-based tasks but remained limited in complex oncology nursing scenarios requiring individualized assessment and dynamic clinical judgment. Their potential use may be most relevant to information retrieval, knowledge organization, and patient education. Because the correctness, clarity, and conciseness rubric showed only moderate interrater reliability, the case-based comparisons should be interpreted as preliminary signals rather than definitive evidence of between-model differences. LLM outputs should therefore be used as supportive information and interpreted alongside professional clinical judgment.
Colorectal cancer (CRC) screening relies on structured risk assessment and guideline-concordant communication, which remain challenging to implement consistently in real-world practice. Digital tools based on large language models (LLMs) may support such workflows, but their feasibility and safety in structured screening contexts have not been well evaluated. This study aimed to develop and evaluate a guideline-concordant chatbot framework for structured CRC screening. A multistage feasibility study was conducted. In phase 1, baseline performance of contemporary LLMs was assessed using 14 standardized CRC screening questions and validated expert-rated instruments. In phase 2, structured prompt versions were iteratively optimized based on screening guidelines and expert feedback and tested using simulated user scenarios. In phase 3, the optimized chatbot was evaluated in 50 screening-eligible adults to assess feasibility, safety, and guideline concordance. In phase 1, both LLMs demonstrated satisfactory performance in responding to standardized CRC screening questions, with DISCERN instrument for AI-generated content scores of 12.02 (SD 0.30) for GPT-4o and 13.36 (SD 0.26) for DeepSeek-V3, Global Quality Score scores of 3.96 (SD 0.10) for GPT-4o and 4.39 (SD 0.08) for DeepSeek-V3, Natural Language Assessment Tool for AI scores of 21.41 (SD 0.34) for GPT-4o and 22.73 (SD 0.27) for DeepSeek-V3, and Patient Education Materials Assessment Tool adapted for AI outputs scores of 0.900 (SD 0.016) for GPT-4o and 0.906 (SD 0.015) for DeepSeek-V3. In phase 2, iterative prompt optimization significantly improved all 6 expert-rated dialogue evaluation dimensions (all P values <.001). In phase 3, the optimized chatbot successfully collected complete CRC screening risk information and generated guideline-concordant screening recommendations for all participants (N=50). No unsafe or inappropriate outputs were identified. This study demonstrated the feasibility, preliminary safety, and guideline-concordant performance of a structured chatbot framework for CRC screening communication under the study conditions. Beyond patient education, the proposed framework may support key components of the CRC screening workflow, including risk information collection, risk stratification, and generation of guideline-concordant screening recommendations. Further prospective studies are needed to evaluate the framework's impact on patient-centered outcomes, screening uptake, and clinical effectiveness.
Robot-assisted radical cystectomy (RARC) is a complex procedure that requires patients to understand surgical indications, urinary diversion, perioperative treatment, complications, recovery, and long-term functional outcomes. Although artificial intelligence (AI) chatbots are increasingly used to obtain medical information, their suitability for RARC patient education remains unclear. We conducted a cross-sectional comparative evaluation of four contemporary AI chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro. A set of 20 core patient-education questions on RARC was developed by three senior urologic experts. Chatbot responses were assessed using DISCERN, the Ensuring Quality Information for Patients tool, the Global Quality Scale, and JAMA benchmark criteria. Readability was evaluated using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, and SMOG. Reliability scores differed significantly across models for DISCERN, EQIP, and GQS, while JAMA benchmark criteria were summarized descriptively as transparency signals. DeepSeek-V4 achieved the highest mean scores for DISCERN, EQIP, and GQS, while ChatGPT-5 and DeepSeek-V4 showed the strongest JAMA benchmark performance. Gemini 3.5 Pro generally had the lowest reliability and transparency scores. Readability also varied across models. DeepSeek-V4 produced the most readable responses overall, whereas Gemini 3.5 Pro generated the most complex text. However, all models exceeded the recommended sixth-grade reading level, and FRES scores remained below the recommended threshold. Contemporary AI chatbots generated responses with variable presentation quality, transparency, and readability for common RARC patient-education questions. Because factual accuracy was not directly assessed, these tools should not be interpreted as validated sources of clinical guidance and should not replace individualized counseling by urologists.
Large language models (LLMs) are increasingly used in medical education and academic writing. However, concerns remain regarding reference hallucination, citation, and the reliability of LLM-generated content. This study aimed to evaluate the performance of ChatGPT 5.2, Gemini 3 Pro, and DeepSeek V3.2 in generating anatomy-related responses by assessing bibliographic reference accuracy, citation content consistency, and the readability of LLM-generated content. A total of 120 open-ended anatomy questions covering six anatomical categories (neuroanatomy, musculoskeletal, respiratory and circulatory, gastrointestinal, urogenital and endocrine, head and neck) were submitted to each model. Individual citation components, including author names, article titles, journal names, publication details, and PMIDs, were verified against indexed sources. Citation content consistency was evaluated using a three-point Likert scale. Readability was assessed using the Flesch Reading Ease score, Flesch-Kincaid Grade Level, Coleman-Liau, and Simple Measure of Gobbledygook indices. A total of 1800 references were analyzed. ChatGPT 5.2 demonstrated the lowest hallucination rate (23.2%), whereas Gemini 3 Pro and DeepSeek V3.2 exhibited substantially higher hallucination rates (45.8% and 47.5%, respectively). DeepSeek V3.2 achieved the highest accuracy for several individual bibliographic components, including author names, article titles, volumes, issues, pages, and journal names. PMID accuracy remained limited across all models, ranging from 25.1% to 57.6%. Citation content consistency differed significantly among the models (p < 0.001), with ChatGPT 5.2 demonstrating the highest proportion of fully supported citations (67.2%), compared with Gemini 3 Pro (42.5%) and DeepSeek V3.2 (41.0%). Citation accuracy differed significantly across most anatomical subcategories, with the greatest intermodel discrepancy observed in head and neck anatomy. Readability analyses indicated that the generated responses generally required college-level reading proficiency. Although LLMs can generate plausible anatomy-related responses, substantial limitations remain regarding reference accuracy, hallucination, and citation reliability. Human verification remains essential before incorporating LLM-generated references into academic or educational materials.