Esophageal cancer remains a significant global health issue. ChatGPT-4o and DeepSeek V3 can provide the public with health-related knowledge about esophageal cancer. This study aimed to evaluate the accuracy of DeepSeek V3 and ChatGPT-4o in responding to health knowledge questions related to esophageal cancer. Fifty-two questions related to esophageal cancer were classified into themes of basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis and patient frequently asked questions (FAQs). These questions were entered into DeepSeek V3 and ChatGPT-4o to obtain responses, and 2 experienced gastroenterologists independently evaluated the accuracy and temporal stability of each response. Overall, the scores of DeepSeek V3 and ChatGPT-4o on all questions were 4 (3-4), and there was no statistically significant difference between the 2 groups. The final scores of DeepSeek V3 in basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis, and FAQs were 4 (3-4), 4 (3-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively, while the scores of ChatGPT-4o were 4 (3-4), 3 (2-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively. For temporal stability across 2 independent test runs, DeepSeek V3 presented inconsistent responses on 2 questions, and ChatGPT-4o on 1 question; no statistically significant differences were found in overall and subgroup scores between the 2 runs for both models (all P > .05). ChatGPT-4o and DeepSeek V3 showed favorable accuracy and comprehensive responses to most of the 52 esophageal cancer-related questions in this study, but our findings do not confirm their general reliability for esophageal cancer health information in routine clinical or public use.
Background Large language models (LLMs) are increasingly being explored for medical applications, including clinical decision support and oncology education. However, their performance in radiation oncology remains insufficiently characterized. Methods This study evaluated and compared the performance of ChatGPT (GPT-5.2; OpenAI, San Francisco, USA) and Gemini (Ultra/Pro; Google, Mountain View, USA) in radiation oncology. A benchmark consisting of 70 multiple-choice questions covering clinical oncology, radiation physics, and radiobiology was used to assess general knowledge. In addition, 25 clinically relevant open-ended questions were independently evaluated by three radiation oncologists using five-point Likert scales for correctness and usefulness. A mixed-effects model was applied to analyze performance. Results ChatGPT achieved an accuracy of 94.3%, while Gemini achieved 97.1% in the multiple-choice assessment. For open-ended questions, both models received similarly high ratings, with mean correctness scores of 4.71 and 4.67 and mean usefulness scores of 4.63 and 4.64 for ChatGPT and Gemini, respectively. Mixed-effects analysis demonstrated a significant effect of question type on both correctness and usefulness, whereas no significant differences between models were observed. Although most responses were rated as good or very good, limitations became apparent in more complex clinical scenarios requiring prioritization and individualized decision-making. Minor discrepancies between benchmark answers and current clinical evidence were also identified. Conclusions Both ChatGPT and Gemini demonstrated high performance on benchmark-based radiation oncology assessments. While the generated responses were generally accurate, limitations remained in complex clinical scenarios requiring nuanced clinical judgment. Further studies are needed to determine whether such performance translates into meaningful clinical utility in real-world radiation oncology practice.
Gastroesophageal reflux disease can progress to reflux esophagitis and Barrett's esophagus (BE), making accurate endoscopic diagnosis important. Artificial intelligence tools like ChatGPT-5 may assist image interpretation, though data on newer large language models remains limited. This study evaluated ChatGPT-5 for BE detection and LA esophagitis severity classification. Endoscopic images from the HyperKvasir dataset were analyzed, including BE, esophagitis A, esophagitis B-D, and normal Z-line images. Four standardized prompts were assessed: (1) BE versus normal, (2) esophagitis versus normal, (3) LA-grade severity (A vs. B-D), and (4) BE versus severe esophagitis. ChatGPT-5 was evaluated in auto mode. Two investigators analyzed 640 unique images, yielding 1280 evaluations. Sensitivity, specificity, positive/negative predictive values (PPV/NPV), F1 scores, and accuracy were calculated. In binary tasks, sensitivity was highest for severe esophagitis (B-D) (0.774). Binary accuracy was similar across tasks (~0.64), with the highest for severe esophagitis (0.655). In three-class analyses, severe esophagitis performed best (sensitivity 0.506, specificity 0.761, accuracy 0.438), with performance improving alongside disease severity. PPV trended higher for severe esophagitis compared with normal mucosa (p ~ 0.03), though significance was not retained after multiple-comparison correction. In the BE-severe esophagitis-normal comparison, severe esophagitis achieved the highest sensitivity (0.590), specificity (0.790), and accuracy (0.521). PPV trended higher for severe esophagitis versus normal mucosa (p ~ 0.05), though significance was not maintained after multiplicity adjustment. NPVs exceeded PPVs across all paradigms. ChatGPT-5 demonstrated moderate performance for esophageal image interpretation, performing best for severe esophagitis and worst for mild esophagitis/BE. Binary prompting outperformed multiclass formats, and the model functioned better as a rule-out tool.
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
To compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren-Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability. In this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0-1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures. Agreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1-2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485). Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.
Artificial intelligence (AI) is increasingly integrated into scientific publishing workflows, yet no study has formally evaluated the ability of large language models (LLMs) to reproduce human editorial desk-review (R0) decisions in a general orthopaedic surgery journal. We investigated whether three commercially available LLMs could accurately replicate the editorial decisions of the Editorial Board of Orthopaedics & Traumatology: Surgery & Research (OTSR). The study addressed four questions: (1) Is the concordance between LLM and human R0 decisions satisfactory for editorial use? (2) Do LLMs exhibit a severity bias? (3) Do LLMs generate decision letters of acceptable quality, and do they reproduce the specific critiques of human reviewers? (4) Does prompt complexity influence LLM decision-making? LLMs used without task-specific fine-tuning or prior exposure to the study corpus would demonstrate at least moderate concordance (κ ≥ 0.40) with human editorial decisions. A corpus of 32 manuscripts randomly selected from submissions to OTSR between 2025 and 2026 (n = 10 outright rejected at R0: 3 out of scope, 3 plagiarism/dual submission, 4 direct desk rejection; n = 11 accepted for peer review; n = 11 rejected after full peer review) was anonymised and independently evaluated, without task-specific fine-tuning or prior exposure to the study corpus, by ChatGPT (GPT-5.5, OpenAI), Gemini (3.1, Google), and Claude (Sonnet 4.6, Anthropic) using a structured prompt incorporating the OTSR guidelines. The primary outcome was assessed using Cohen's kappa between LLM and human binary decisions. Secondary outcomes included accuracy, inter-LLM agreement, domain-specific scoring, ARCADIA quality scoring of 115 eligible decision letters by two blinded raters with ICC, human-performed content concordance analysis, sensitivity analysis (structured vs. minimal prompt), and test-retest reproducibility at 24 hours. Overall accuracy (i.e. the decision was similar for LLM and editorial decision) was 59.4% for ChatGPT (19/32) and Claude (19/32), and 62.5% for Gemini (20/32). Cohen's kappa was near-zero for ChatGPT (κ = -0.05) and Gemini (κ = 0.00), and low for Claude (κ = 0.15). All LLMs showed systematic over-rejection of accepted manuscripts. No LLM identified plagiarism or simultaneous dual submission as a rejection motive. Test-retest concordance was 84.4 - 90.6% across models. ARCADIA quality scoring (n = 115 scorable letters, inter-rater ICC = 0.86, 95% CI 0.80-0.90) showed Claude achieved the highest scores (4.46 ± 0.32 /5), significantly above the human OTSR letters (4.01 ± 0.44, p < 0.001), ChatGPT (3.81 ± 0.49, p = 0.002), and Gemini (3.28 ± 0.49, p < 0.001). LLMs reproduced 30 - 41% of human reviewer-specific critiques, with Claude achieving the highest match (40.5%) without hallucinations. Switching to a minimal prompt markedly increased acceptance rates for Gemini (87.5%) and Claude (75%), while ChatGPT remained largely insensitive to prompt simplification (15.6%). The principal finding of this study is the dissociation between formal review quality and true editorial reliability. Although modern LLMs generated persuasive and methodologically structured decision letters, they failed to achieve meaningful concordance with real editorial decisions and displayed stable architecture-specific biases that were highly sensitive to prompt design. These results indicate that current LLMs reproduce the surface features of peer review more successfully than its underlying scientific and contextual reasoning. Consequently, LLMs may represent valuable supervised assistants for editorial workflows, but not reliable autonomous substitutes for human editorial expertise in orthopaedic scientific publishing. IV; Observational pilot study, concordance analysis.
Large language models (LLMs) are increasingly used in health care, with emerging applications in clinical decision support and nursing education. However, evidence on their performance in nursing contexts, particularly in oncology nursing, remains limited. Given the complexity and high-risk nature of oncology care, it is important to evaluate the performance and clinical relevance of LLM-generated responses in oncology nursing contexts. This study aimed to compare the performance of LLMs in oncology nursing decision support tasks using standardized examination questions and case-based clinical scenarios and explore LLMs' potential applicability and current limitations in oncology nursing practice. A total of 33 case-based questions derived from 10 oncology nursing clinical scenarios in a nationally used training manual, along with standardized examination-oriented questions from a commercially published preparation book for the Chinese Nursing (Intermediate) Qualification Examination, were used to evaluate the performance of 5 LLMs (DeepSeek, Qwen, Spark-Desk, WiseDiag, and ChatGPT). All models generated responses using a standardized prompt. Two oncology nurses with more than 5 years of clinical experience independently rated the case-based responses using 3 evaluation dimensions: correctness, clarity, and conciseness. Interrater reliability was assessed using the quadratic weighted Cohen κ, intraclass correlation coefficient, and Spearman rank correlation coefficient. Differences among models were analyzed using the Kruskal-Wallis test with the Dunn post hoc test. In addition, examination performance was evaluated based on total score, accuracy rate, and completion efficiency. Interrater reliability analyses indicated moderate agreement between evaluators. The median correctness, clarity, and conciseness scores were as follows: 11.50 (IQR 10.50-12.00) for DeepSeek, 11.00 (IQR 10.50-12.00) for Qwen, 10.50 (IQR 9.50-11.50) for Spark-Desk, 10.00 (IQR 9.50-11.50) for WiseDiag, and 10.00 (IQR 9.00-11.50) for ChatGPT. The Kruskal-Wallis test indicated statistically significant differences among models (H=11.416; P<.05), with post hoc analysis showing a significant difference only between DeepSeek and ChatGPT (P<.05). In examination-based tasks, all models achieved passing performance, with accuracy rates ranging from 77% (77/100) to 93% (93/100). In terms of response completion, DeepSeek and ChatGPT completed all tasks in a single interaction, whereas other models required multiple interactions due to output interruptions. LLMs showed relatively strong performance on structured knowledge and examination-based tasks but remained limited in complex oncology nursing scenarios requiring individualized assessment and dynamic clinical judgment. Their potential use may be most relevant to information retrieval, knowledge organization, and patient education. Because the correctness, clarity, and conciseness rubric showed only moderate interrater reliability, the case-based comparisons should be interpreted as preliminary signals rather than definitive evidence of between-model differences. LLM outputs should therefore be used as supportive information and interpreted alongside professional clinical judgment.
To evaluate the readability, quality, and misinformation of patient education materials generated by large language models, including ChatGPT-4o (OpenAI), Gemini 1.5 Pro (Google), and Copilot Pro (Microsoft), compared with American Society of Retina Specialists (ASRS) brochures for retinal diseases. A cross-sectional comparative analysis was performed by generating patient education materials on 3 retinal conditions: retinal detachment, diabetic retinopathy, and age-related macular degeneration. Materials were created using a general prompt (prompt A) and a prompt specifying a sixth-grade readability level (prompt B). Readability was evaluated using 6 validated metrics. Quality was assessed through DISCERN and the Patient Education Materials Assessment Tool. Misinformation was graded using a 5-point Likert scale. Assessments were performed independently by 2 masked retina specialists. Average readability of Gemini (11.65; P = .005) and Copilot (11.23; P = .003) materials was significantly better than that of ASRS materials (14.17), whereas ChatGPT showed no significant difference (12.85; P = .06). ChatGPT's average readability was significantly lower compared with Gemini (12.85 vs 11.65; P = .01) and Copilot (12.85 vs 11.23; P < .001). Prompt B significantly improved readability across all large language models relative to ASRS but still exceeded the sixth-grade readability level. DISCERN scores were comparable across groups. ASRS materials had an understandability score of 74.37%, which was significantly lower than ChatGPT (94.44%; P = .02) and Gemini 1.5 (95.83%, P = .02) scores. No significant differences were observed for actionability or misinformation. Readability showed no significant correlation with quality or misinformation (P > .05). Large language models, when appropriately prompted, can generate retina-related patient education material with superior readability compared with existing ASRS brochures, while maintaining comparable quality and accuracy. Large language models represent a promising approach for addressing literacy barriers, though expert oversight remains essential.
BackgroundArtificial intelligence (AI) increasingly supports clinical decision making. ChatGPT-5 (OpenAI, San Francisco, CA) and OpenEvidence (OpenEvidence Inc, Miami, FL) are regularly used by physicians, yet the limits of their reliability in decision-making remain poorly defined. This study evaluated both platforms against the National Comprehensive Cancer Network® (NCCN) Guidelines for Rectal Cancer to define where AI tools perform well and where they fall short.MethodsThe NCCN Guidelines for Rectal Cancer (Version 2.2025) were reviewed. Three questions were generated per decision-making page and classified into workup/diagnosis, treatment, and surveillance domains, yielding 138 clinical scenarios. Both platforms were queried. Responses were scored independently by two physicians on a 5-point Likert scale (5 = completely correct; 1 = absolutely incorrect). Two performance thresholds were defined: Correctness (≥3) and Accuracy (≥4). Proportions were compared with Fisher's exact test and score distributions with the Mann-Whitney U test.ResultsBoth platforms demonstrated high overall guideline concordance. ChatGPT-5 achieved Correctness in 136 (98.6%) and Accuracy in 116 (84.1%), compared to 128 (92.8%) and 112 (81.2%) for OpenEvidence, respectively. Overall Correctness favored ChatGPT-5 (P = 0.035), while Accuracy showed no significant difference (P = 0.634). Mean Likert scores were 4.57 vs 4.37 (P = 0.089). Both platforms achieved >90% Correctness across all domains.ConclusionBoth platforms demonstrate reliable identification of the broad direction of care but exhibit important limitations in completeness, nuance, and the handling of preference-sensitive decisions. These findings define the appropriate role of AI-assisted decision support as an adjunct to, rather than a substitute for, multidisciplinary review in rectal cancer management.
Accurate assessment of pulpal status is essential for achieving successful endodontic outcomes. However, direct evaluation remains inherently challenging because the pulp is surrounded by calcified tissue, necessitating reliance on clinical and radiographic examinations for diagnostic and prognostic decision-making. These procedures demand substantial clinical expertise and time, and less-experienced clinicians often face challenges that may lead to errors in diagnosis and treatment planning. Recent advancements in large language models (LLMs) offer promising opportunities to enhance clinical reasoning by facilitating the integration of evidence and supporting methodical diagnostic decision-making. This study aimed to evaluate the clinical applicability of LLMs by comparing their text-based clinical screening performance and the clinical validity of their treatment plan responses with those of human evaluators. Between January 2011 and December 2022, 100 clinical cases involving primary endodontic disease were randomly selected from the clinical records of outpatients who visited the Department of Conservative Dentistry or Advanced General Dentistry (AGD) at Yonsei University Dental Hospital. Four prompt types, combining 2 variables (language and role), were used as input for 4 LLMs. Both LLMs and human evaluators (AGD specialists, AGD residents, endodontic residents, and senior dental students) assessed the cases using text-based clinical records. Radiographic images were not directly provided. Screening performance was evaluated using a 0-to-2-point concordance scale, and treatment plan validity and relevance were assessed using a 5-point Likert scale. Among the 4 LLMs evaluated, ChatGPT achieved the highest mean concordance score on Korean-doctor prompts (mean 0.98, SD 0.82). However, this score did not reach the partially correct criterion of 1 on the 0 to 2-point scale. Clova X recorded the lowest mean score on English-patient prompts (mean 0.23, SD 0.63). Across both diagnostic categories, AGD specialists demonstrated the highest diagnostic accuracy (pulpal: 0.70; periapical: 0.65), with higher sensitivity but lower specificity than those exhibited by the other groups. ChatGPT also showed favorable performance among the LLMs, with accuracies of 0.65 (95% CI 0.55-0.74) for pulpal disease and 0.57 (95% CI 0.47-0.69) for periapical disease, which were comparable to those of AGD and endodontic residents. Under image-free clinical record review conditions, ChatGPT 4.0 showed relatively higher and more consistent performance in symptom screening and treatment planning compared to the other LLMs evaluated. However, its highest mean score of 0.98 (SD 0.82) did not reach the partially correct criterion of 1 on the 0 to 2-point scale. Hallucinations generated by LLMs and experience-dependent interpretation biases among human evaluators remain key challenges that require attention. Therefore, continuous clinical supervision and comprehensive user training are necessary for the safe and effective clinical application.
The capabilities of general-purpose large language models (LLMs) on specialized medical examinations have not been systematically compared. To evaluate the performance differences among three LLMs-DeepSeek-R1, ChatGPT-o1, and Gemini-2.0-in answering questions from China's Standardized Training Examination for Resident Physicians in Radiology, and to assess the impact of different questioning conditions (with/without answer choices, and the introduction of doubt) on model accuracy. A total of 131 questions were analyzed at one time. The LLMs were subjected to two tasks (questions with/without answer choices). Each task included three conditions: no doubt, weak doubt, and strong doubt, with the latter two presented as follow-up questions after the models' initial responses. Subjective evaluation was conducted using the 5-point Likert scale. In both tasks, Gemini-2.0 achieved the highest accuracy (0.763-0.809) and (0.595-0.679). In Task 2, all LLMs' accuracy was lower than in Task 1, but only DeepSeek-R1 showed statistical significance (p = 0.014). The introduction of doubt (weak or strong) did not significantly increase accuracy in Task 1 for any model. In Task 2, only DeepSeek-R1 showed a slight increase in accuracy under doubt conditions. All LLMs performed significantly better on single-choice questions than on multiple-choice questions (p < 0.05) and showed superior performance in case-based questions vs. knowledge-based questions. The subjective evaluation scores were 4.46 for DeepSeek-R1, 3.92 for ChatGPT-o1, and 3.83 for Gemini-2.0. It was noted that ChatGPT-o1 negotiated not to answer all questions at once and had missing responses. LLMs exhibit variable proficiency in tackling radiology resident examination questions, with Gemini-2.0 showing the highest overall accuracy. However, repeated self-examination of LLMs through the introduction of doubt does not consistently or significantly enhance their performance on radiology-related questions.
Childhood accidents are among the leading causes of injury during early childhood. This study aimed to evaluate and compare the accuracy, clarity, and comprehensiveness of pediatric first aid information generated by LLMs. A cross-sectional comparative evaluation design was employed. Twenty standardized pediatric first aid questions were developed based on international guidelines and expert consensus. Responses generated by ChatGPT, Claude, Gemini, and Copilot were independently evaluated by a pediatric nurse and a physician using a 5-point Likert scale. Inter-rater reliability was assessed using Cohen's kappa coefficient, and differences among models were analyzed using one-way analysis of variance. Moderate inter-rater agreement was observed across all evaluation domains. Statistically significant differences were identified among the four LLMs. Claude demonstrated the highest overall performance across all evaluation domains. Gemini demonstrated relatively high accuracy but lower clarity and comprehensiveness scores. Copilot performed well in clarity but showed limited depth of clinical content. ChatGPT received the lowest scores across all assessed domains. The findings reveal considerable variability in the quality of pediatric first aid information generated by LLMs. While certain models may serve as supportive educational tools, none should be considered a substitute for professional medical assessment or emergency care.
General-purpose large language model (LLM)-based systems are increasingly accessible to clinicians and are being explored for applications in clinical microbiology and infectious diseases (ID). However, rapid adoption has outpaced the development of specialty-specific guidance, raising concerns related to safety, reliability, accountability, and antimicrobial stewardship. The project aims to develop consensus-based statements, endorsed by Study Group for Artificial Intelligence and Digitalisation of the European Society of Clinical Microbiology and Infectious Diseases (ESGAID), that describe principles, opportunities, and limitations in the interactions of clinical microbiologists and ID specialists with general-purpose LLM-based systems. Secondary objectives are to quantify expert agreement and identify areas of uncertainty and disagreement. The project follows a structured expert consensus design using the RAND/UCLA Appropriateness Method. A multidisciplinary panel of 15 experts will be selected through ESGAID using predefined criteria to ensure balanced expertise across clinical microbiology, infectious diseases, ethics, legal aspects, and patient safety. Ten draft statements, each supported by a structured literature review, will be developed by project coordinators. Statements will be evaluated through iterative rounds of anonymous rating on a 1-9 scale, combined with moderated remote discussions. Median scores will classify statements as supported, uncertain, or unsupported. Consensus will be defined as ⩾70% agreement with <15% disagreement during final anonymous voting. This protocol provides a transparent and reproducible framework to generate interim, specialty-specific statements on principles, opportunities, and limitations in the interactions of clinical microbiologists and ID specialists with general-purpose LLM-based systems. By combining structured evidence review with expert judgment, the resulting statements aim to delineate guidance on principles for interacting with these systems, highlight the nature of both existing risks and excessive skepticism, and identify research priorities in a rapidly evolving technological and regulatory landscape. Interactions with AI systems in microbiology and infectious diseases: an expert consensus project. Artificial intelligence systems such as ChatGPT, Gemini, or Claude are increasingly accessed by clinical microbiologists and infectious diseases specialists to explore medical questions or summarize information. However, clear understanding and deep knowledge about key principles, opportunities, limitations, and safeguards in the interaction with these tools may still be uncommon. For example, these systems can sometimes produce convincing but incorrect answers or lead users to rely too heavily on automated outputs. This project brings together experts to develop agreed statements on the principles guiding how clinicians can interact with these systems responsibly.
To compare the quality of scientific review articles generated by two artificial intelligence systems, ChatGPT and Gemini, with those written by human authors in the field of otolaryngology. Two otolaryngology topics, chronic rhinosinusitis and infantile subglottic hemangioma, were selected. For each topic, four AI-generated reviews (GPT-4.0 and Gemini 2.0; narrative and PRISMA-style) and one human-authored peer-reviewed review were included, yielding a total of 10 manuscripts (8 AI-generated, 2 human-authored). A blinded panel of seven board-certified otolaryngologists evaluated all manuscripts using a 5-point Likert scale across seven domains: scientific accuracy, depth of content, citation quality, structure and organization, readability and tone, critical insight, and overall scientific quality. Group comparisons were performed using linear mixed-effects models with random intercepts for reviewer and manuscript. Interrater reliability was assessed using Shrout-Fleiss intraclass correlation coefficients (ICC). Manual verification of AI-generated references was conducted to assess citation accuracy and fabrication. Human-authored manuscripts received the highest ratings across all domains (overall quality 4.50 ± 0.76). GPT-4.0 demonstrated moderate performance (2.71 ± 1.46 overall), while Gemini 2.0 scored lowest (2.14 ± 1.01). Mixed-effects modeling demonstrated significant group differences across all domains (p ≤ 0.008). Citation quality showed one of the largest between-group differences and strong reliability [ICC(2,1) = 0.68; ICC(2,7) = 0.94]. Manual verification of 123 AI-generated references revealed high citation accuracy for GPT-4.0 (89.6% fully accurate; 0% fabricated) compared with Gemini 2.0 (71.7% fully accurate; 23.9% fabricated). Reviewers misclassified 50% of GPT-4.0 manuscripts as human-authored, correctly identified 93% of human-authored manuscripts, and classified 86% of Gemini manuscripts as AI-generated. GPT generated fluent, stylistically strong reviews but remained significantly inferior to human-authored manuscripts in analytical depth and citation integrity. Gemini 2.0 underperformed across all domains and demonstrated a substantial rate of fabricated citations. As large language models become integrated into academic workflows, transparent disclosure, structured fact-checking, and human oversight remain essential to safeguard scientific reliability.
To use Google Trends to assess whether interest in rotator cuff repair (RCR) has increased over the last 15 years and to identify the most frequently searched topics related to RCR. A retrospective longitudinal study was conducted on public interest in RCR using Google Trends. While using a focus group comprised of 2 sports-medicine fellowship-trained orthopaedic surgeons, 1 orthopaedic surgery resident, 3 medical students, and supplementing with ChatGPT, the 50 most commonly searched topics perioperatively regarding RCR were identified. The mean relative search volume (RSV) was obtained and compared for all 50 topics. Regarding RSV, which extends from 0 to 100, 0 represents no public interest, whereas 100 represents maximal public interest. Inclusion criteria for search topics included Google search terms concerning RCR which had a mean RSV available between January 2010 and December 2024. Analysis of variance tests were used to compare means. Statistical significance was set at a P value less than .05. Between January 2010 and December 2024, the mean RSV for "rotator cuff surgery" significantly increased (2010: 51; 2024: 90; P < .001), representing growing patient interest in rotator cuff surgery. Among the 50 perioperative topics, "recovery" and "pain" had significantly higher mean RSVs than other topics, whereas "sling," "therapy," "cost," "how long does rotator cuff surgery take," "arthroscopic surgery," "sleep," "rehab," and "healing" were also among the top ten topics with the highest mean RSV (79; 51; 20; 18; 14; 13; 9; 8; 7; 4; P < .001). Google searches for RCR have increased significantly during the past 15 years. Pain and recovery after RCR were the most frequently searched topics. This study investigates the public interest in RCR as indicated by Google search data. The level of interest has increased during the past 15 years, and the topics identified in this study should be included in patient education materials.
To use Google Trends to study whether public interest in meniscus surgery has increased during the last 15 years and after the COVID-19 pandemic, and to determine what specific topics regarding meniscus surgery the public is interested in. A longitudinal observational study was conducted between January 2010 and December 2024 on public interest in meniscus surgery utilizing Google Trends. Through establishing a focus group comprising two attending orthopaedic surgeons (sports-medicine fellowship trained), one orthopaedic surgery resident, and three medical students, and using ChatGPT, the 50 most commonly searched questions regarding meniscus surgery were identified. The mean relative search volume (RSV) was identified and compared for 50 search terms. RSV extends from 0 to 100, with 0 representing no public interest, and 100 representing maximal public interest. Analysis of variance tests compared means. Statistical significance was set at a P value less than .05. Between 2010 and 2024, the mean RSV for "meniscus surgery" increased significantly (2010: 40; 2024: 89; P < .001), after a transient decline during the COVID-19 pandemic. Among 50 topics regarding meniscus surgery, "recovery" and "pain" had the highest mean RSVs, while other topics including "arthroscopic meniscus surgery", "cost", "how long does meniscus surgery take", "therapy", "brace", "rehab", "swelling", and "arthritis" were also commonly searched (60 vs 25 vs 15 vs 9 vs 9 vs 6 vs 6 vs 5 vs 4 vs 2; P < .001). Patient interest in meniscus surgery has increased significantly during the past 15 years, and rebounded after the COVID-19 pandemic. Patients have interest in many topics regarding meniscus surgery, including recovery, pain, surgical length, cost, postoperative therapy, braces, and post-traumatic arthritis. This study showed that public interest in meniscus surgery has increased significantly in the last 15 years, and the various topics that are important to patients, as identified in this study, should be included in patient education materials.
Orthopedic-related rare diseases are difficult to diagnose because of their low prevalence, heterogeneous phenotypes, and fragmented knowledge. Large language models (LLMs) can serve as dynamic knowledge-support tools, but their diagnostic performance and effect on physicians' decision-making remain unclear. This study aims to compare the diagnostic performance of advanced LLMs for orthopedic-related rare diseases and to evaluate the effect of a 2-stage LLM-assisted diagnostic workflow on physicians' diagnostic accuracy and subjective acceptance. We selected 40 orthopedic-related rare diseases from the Chinese Rare Disease Catalog. A total of 4 general-purpose LLMs each generated 1 primary diagnosis and 5 differential diagnoses per case. Diagnostic accuracy, defined as a correct primary diagnosis, was compared using the Cochran Q test and pairwise McNemar tests with Bonferroni correction. A representative LLM was integrated into a 2-stage workflow involving 27 intermediate and 15 senior orthopedic physicians. Physicians first diagnosed all cases independently and then rediagnosed the same cases after reviewing nonauthoritative LLM suggestions. Physician diagnostic data were primarily analyzed using mixed-effects logistic regression at the individual-diagnosis level. Case-level group accuracy was additionally assessed using >50% and ≥2/3 accurate-physician thresholds. After both rounds, physicians completed an 8-item Likert-scale questionnaire assessing subjective acceptance and workflow perceptions. Claude Sonnet 4.5, ChatGPT-5.0, and Gemini 2.5 Pro each achieved 90% (36/40) primary-diagnosis accuracy, whereas DeepSeek-V3.2 achieved 67.5% (27/40; Cochran Q P<.001). Before LLM assistance, mean physician-level accuracy was 42.22% for intermediate physicians and 58.67% for senior physicians; after assistance, it increased to 68.80% and 83.33%, respectively. In the primary mixed-effects logistic regression analysis, physician seniority group and LLM assistance stage were significantly associated with diagnostic correctness (both P<.001), whereas the group-by-stage interaction was not significant (P=.10). Secondary case-level analyses using the >50% threshold showed improvement from 40% (16/40) to 67.5% (27/40) for intermediate physicians and from 57.5% (23/40) to 82.5% (33/40) for senior physicians, with similar findings using the ≥2/3 threshold. Cases accurately diagnosed by all 3 agents increased from 16 to 27. The questionnaire showed high internal consistency (Cronbach α=0.902) and generally positive attitudes, with no significant differences between physician groups (P=.11 to P=.78). LLMs achieved high diagnostic accuracy for orthopedic-related rare diseases. In the 2-stage LLM-assisted workflow, LLM assistance was associated with higher diagnostic correctness in both physician groups, although seniority-related differences in the magnitude of benefit require evaluation in larger studies. Senior physicians retained higher diagnostic correctness than intermediate physicians. Secondary case-level analyses suggested attenuation of group-level gaps in case-recognition patterns. Physicians reported broadly positive workflow perceptions. Given the same-day repeated-case design and potential short-term recall bias, these exploratory findings should be interpreted cautiously and warrant prospective randomized, crossover, washout-period, independent-case, or real-world evaluations of LLM-assisted diagnostic workflows in orthopedics.
This in vitro study aimed to compare and analysis degree of Manufacture accuracy and geometric discrepancy measurements of Cobalt-Chromium (C0-Cr) bar joint attachment supporting a mandibular complete overdenture fabricated directly by Direct Laser Sintering (DMLS) and indirectly fabricated by conventional casting of 3D printed resin pattern using Co-Cr alloys with the same reference design. A total of twelve bar joint attachment for two implant-retained mandibular overdentures were fabricated through additive manufacturing process indirectly by using 3D printed resin pattern and directly by using DMLS (n = six for each group). All the sample specimens of metal bars were scanned and superimposed with the reference STL reference design by using the Geomagic Control X software program. There was no significant statistical difference in manufacture accuracy and geometric discrepancy in the bars fabricated by DMLS compared to indirectly 3D-printed conventional Co-Cr bars in mandibular implant retained overdenture. The Casted group has a mean of 0.1128 ± 0.006, while the DMLS group has a mean of 0.1206 ± 0.01. There was no significant difference between the groups, Regarding Internal Discrepancy, the Casted group shows a mean of 0.1659 ± 0.118, and the DMLS group has a mean of 0.1678 ± 0.035. The research indicates no significant difference between the groups for this variable. No significant difference in Manufacture accuracy and geometric discrepancy of direct laser sintering versus cast 3D printed resin pattern of Co-Cr bars in mandibular implant retained complete overdenture.
Retinal image quality significantly affects the performance of diagnostic artificial intelligence (AI) models and is typically improved with pupil dilation in clinical settings. However, in real-world settings where dilation is not feasible, suboptimal image quality remains a challenge for AI deployment. In this study, we fine-tuned a CofeNet model to enhance the quality of undilated retinal images. We performed pixel-level alignment and fine-tuned a generative CofeNet model using 313 paired and spatially aligned undilated and dilated retinal images. Model performance was evaluated on an internal test (120 image pairs) and two external tests. Image similarity was assessed using peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). Image gradability was evaluated for undilated and model-enhanced images by three graders on a three-point scale (gradable, force gradable, ungradable), with consensus-based grading. Inter-rater agreement among the three graders was evaluated using Fleiss' kappa (k). In the internal test set, PSNR and SSIM improved significantly from 23.2 and 0.823 (undilated) to 28.7 and 0.923 (model-enhanced) when compared to ground truth dilated images (both p < 0.001). The gradability of the images improved from 41% (49 of 120 images) to 76% after enhancement, although the inter-rater agreement decreased (overall Fleiss' κ from 0.397 to 0.234). Similar trends were observed in the external tests. The enhanced images demonstrated improved quality and greater structural similarity to dilated images, suggesting the potential of the CofeNet model as an alternative approach to enhance image quality. Further validation is necessary to determine its clinical utility.