Modern large language models (LLMs) like ChatGPT (based on the GPT-4 architecture) and DeepSeek offer unprecedented capabilities for generating scientific text. However, their performance in replicating structured, high-quality scientific writing, especially compared to human-authored abstracts, remains insufficiently evaluated. To compare the abstract quality produced by human authors, ChatGPT/GPT4, and DeepSeek across: six evaluation criteria Clarity, Coherence, Conciseness, Accuracy, IMRaD Structure, and Language Quality, using blinded expert ratings and non-parametric statistical methods, specifically the Kruskal-Wallis test followed by pairwise Wilcoxon rank-sum tests with false discovery rate correction. We selected 23 medical and healthrelated research topics, each yielding three abstracts (human, ChatGPT, DeepSeek), for a total of 69 abstracts. Three raters scored each abstract. Kruskal-Wallis tests assessed group differences; Cliff's Delta (δ) was calculated as a nonparametric effect size for each comparison, suitable for ordinal data. Across criteria, ChatGPT and DeepSeek significantly outperformed human authors in Clarity, Coherence, IMRaD Structure, and Language Quality. In contrast, Conciseness and Accuracy showed negligible effect sizes (|δ| <0.10), suggesting parity across all three sources. ChatGPT and DeepSeek achieved significantly higher scores in clarity, coherence, structure, and language quality, while showing comparable performance in conciseness and accuracy. These findings complement recent evaluations showing competitive medical and reasoning performance of DeepSeek models compared to proprietary LLMs. While shortform abstracts, expert oversight, and domain expertise remain critical, the results suggest that LLMs-particularly GPT4 and DeepSeek-can serve as effective tools in drafting scientific abstracts.
To evaluate and compare the performance of four general-purpose large language models (LLMs) (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) in answering specialised clinical questions related to total knee arthroplasty (TKA) derived from the World Expert Meeting in Arthroplasty (WEMA). This is a cross-sectional comparative study. Twenty questions on TKA supported by moderate-strong level of evidence were randomly selected from the WEMA. Three orthopaedic surgeons independently performed a blinded assessment of all LLM-generated responses. An adapted version of the QUEST rating system, a comprehensive framework designed for the objective human assessment of LLM performance across healthcare-related subdomains, was used. Furthermore, the same three evaluators subjectively selected the best-performing LLM response for each question. The four LLMs presented statistically significant differences in overall performance based on the QUEST framework (score range 1-5): Gemini 2.5; 4.77 ± 0.06, Claude 4; 4.72 ± 0.07, ChatGPT-5; 4.70 ± 0.08 and GROK 4; 4.62 ± 0.09 (p < 0.001). Gemini 2.5 achieved the highest scores in the Accuracy (4.58 ± 0.70), Comprehensiveness (4.87 ± 0.34) and Trust (4.52 ± 0.62) dimensions. However, Claude 4 obtained the highest score for the Currency (4.23 ± 0.67) dimension. When assessors subjectively selected the superior answer for each question, Claude 4 was chosen most frequently, in 46.7% of cases, followed by ChatGPT-5 in 25.4%, Gemini 2.5 in 22.9% and GROK 4 in 7.5% of cases. Four general-purpose LLMs (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) demonstrated good overall performance when addressing specialised clinical questions related to TKA. No single model consistently outperformed the others across all evaluated domains. Level V.
Gastroesophageal reflux disease can progress to reflux esophagitis and Barrett's esophagus (BE), making accurate endoscopic diagnosis important. Artificial intelligence tools like ChatGPT-5 may assist image interpretation, though data on newer large language models remains limited. This study evaluated ChatGPT-5 for BE detection and LA esophagitis severity classification. Endoscopic images from the HyperKvasir dataset were analyzed, including BE, esophagitis A, esophagitis B-D, and normal Z-line images. Four standardized prompts were assessed: (1) BE versus normal, (2) esophagitis versus normal, (3) LA-grade severity (A vs. B-D), and (4) BE versus severe esophagitis. ChatGPT-5 was evaluated in auto mode. Two investigators analyzed 640 unique images, yielding 1280 evaluations. Sensitivity, specificity, positive/negative predictive values (PPV/NPV), F1 scores, and accuracy were calculated. In binary tasks, sensitivity was highest for severe esophagitis (B-D) (0.774). Binary accuracy was similar across tasks (~0.64), with the highest for severe esophagitis (0.655). In three-class analyses, severe esophagitis performed best (sensitivity 0.506, specificity 0.761, accuracy 0.438), with performance improving alongside disease severity. PPV trended higher for severe esophagitis compared with normal mucosa (p ~ 0.03), though significance was not retained after multiple-comparison correction. In the BE-severe esophagitis-normal comparison, severe esophagitis achieved the highest sensitivity (0.590), specificity (0.790), and accuracy (0.521). PPV trended higher for severe esophagitis versus normal mucosa (p ~ 0.05), though significance was not maintained after multiplicity adjustment. NPVs exceeded PPVs across all paradigms. ChatGPT-5 demonstrated moderate performance for esophageal image interpretation, performing best for severe esophagitis and worst for mild esophagitis/BE. Binary prompting outperformed multiclass formats, and the model functioned better as a rule-out tool.
To investigate the effectiveness of a case-based learning (CBL) model integrated with ChatGPT in ophthalmology clinical teaching and to compare its educational outcomes with those of the traditional multimedia lecture-based approach. A total of 98 fifth-year clinical medicine students from Shanghai Jiao Tong University School of Medicine (Class of 2021) were randomly assigned to an experimental group (CBL + ChatGPT, n = 49) or a control group (traditional lecture-based teaching, n = 49). The teaching intervention lasted for 8 weeks, with one 90-min session per week. The experimental group adopted a standardized human-AI interaction protocol, including a structured questioning framework, no fewer than three rounds of dialog, and real-time teacher supervision and correction. The control group received conventional multimedia lectures. Outcomes included theoretical knowledge examination scores, case analysis scores, recognition rates of the teaching model across seven dimensions, and overall teaching satisfaction. Subgroup analyses were further conducted according to students' baseline academic performance (high-, medium-, and low-foundation groups). The experimental group achieved significantly higher theoretical examination scores than the control group (86.4 ± 5.2 vs. 78.9 ± 6.1, P < 0.01), as well as higher case analysis scores (84.7 ± 6.3 vs. 74.2 ± 7.5, P < 0.01). Recognition rates across all seven evaluation dimensions, including interactivity, immediate feedback, and clinical relevance, were significantly higher in the experimental group (all P < 0.05). The overall satisfaction rate was 89.8% in the experimental group and 63.3% in the control group (P < 0.01). Subgroup analysis demonstrated that students with weaker academic foundations benefited most from the intervention, showing the greatest improvement in scores (+12.3 points, P < 0.01). These students also reported significantly higher recognition of personalized learning experiences than students with stronger academic backgrounds (95.2% vs. 81.2%, P < 0.05). The CBL teaching model integrated with ChatGPT significantly improves both objective learning outcomes and subjective satisfaction in ophthalmology clinical education, particularly among students with weaker academic foundations. This teaching approach demonstrates substantial potential for broader implementation in medical education.
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
AI is increasingly used for medical education and assessment, yet effectiveness on specialized exams remains underexplored. We aim to compare multiple AI models with national human averages on SBCOC entrance exams (2021-2023), evaluate answer accuracy and citation reliability, and contrast models. Five models-ChatGPT-4 (standard and literature-trained), ChatGPT-o1-pro, Gemini, and Meta Llama 3.1-answered official SBCOC exams (50 items/year) individually using a unified prompt and identical images. Each response included one choice (A-D) and a cited source. Outcomes were accuracy and source reliability (peer-reviewed articles/textbooks versus websites). Statistics used chi-square tests and one-way ANOVA with Tukey post-hoc (p<0.05). Across 150 questions, ChatGPT-o1-pro led (66%, 62%, 68%) and exceeded human national averages each year (p<0.05). Gemini (40%, 28%, 38%) and Meta Llama 3.1 (32%, 44%, 50%) underperformed relative to humans, while both ChatGPT-4 versions hovered near the ≥50% pass threshold without significant differences. Regarding sources, ChatGPT-o1-pro and Llama predominantly cited articles or books, ChatGPT-4 alternated between literature and websites, and Gemini did not specify references. ChatGPT-o1-pro outperformed the national human average and mostly used credible sources; ChatGPT-4 matched humans, while Gemini and Meta Llama 3.1 lagged. Level of evidence IV; Case series. A IA é cada vez mais usada na educação e avaliação médicas, mas sua efetividade em exames especializados permanece pouco explorada. O objetivo foi comparar modelos de IA às médias humanas nos exames de ingresso da SBCOC (2021–2023), avaliar acurácia e confiabilidade das citações e contrastar modelos. Cinco modelos—ChatGPT-4 (padrão e treinado com literatura), ChatGPT-o1-pro, Gemini e Meta Llama 3.1—responderam às provas da SBCOC (50 itens/ano) individualmente, com prompt unificado e imagens idênticas. Cada resposta continha uma alternativa (A–D) e uma fonte. acurácia e confiabilidade das fontes (artigos/livros vs. sites). Estatística: qui-quadrado e ANOVA de uma via com pós-teste de Tukey (p<0,05). Em 150 questões, o ChatGPT-o1-pro liderou (66%, 62%, 68%) e superou médias humanas em todos os anos (p<0,05). Gemini (40%, 28%, 38%) e Meta Llama 3.1 (32%, 44%, 50%) abaixo dos humanos; ambas as versões do ChatGPT-4 próximas ao limiar de aprovação ≥50%, sem diferenças significativas. Quanto às fontes, ChatGPT-o1-pro e Llama citaram predominantemente artigos/livros; o ChatGPT-4 alternou entre literatura e sites; Gemini não especificou referências. O ChatGPT-o1-pro superou a média humana e usou majoritariamente fontes confiáveis; o ChatGPT-4 igualou os humanos; Gemini e Meta Llama 3.1 ficaram aquém. Nível de evidência IV; série de casos.
As large language models (LLMs) are increasingly used to interpret medical concerns, rigorous evaluation of their performance on clinically relevant tasks is essential. However, the new state-of-the-art models from OpenAI and Google as of February 2026, have not been evaluated for their accuracy and consistency in interpreting ECGs independent of clinical context. We aim to compare ChatGPT (GPT-5.2 Thinking) and Gemini (Gemini 3 Pro) on electrical axis and heart rhythm identification to assess current clinical usability and identify areas for improvement. ECGs were obtained from the Lobachevsky University Electrocardiography Databases on PhysioNet. The LLM responses were evaluated for first-shot accuracy and consistency across three different trials. First-shot accuracy was further split into the various types of rhythms and axes to examine systematic trends in model outputs. Both models demonstrated comparable overall first-shot accuracies for electric axis and rhythm classification. Both models struggled with less common rhythm categories, including multifocal rhythms and tachycardias. The Macro F1 analysis indicated low overall classification performance for both ChatGPT and Gemini in terms of axis and rhythm. ChatGPT achieved a higher Macro F1 point estimate for axis classification compared with Gemini, though both models struggled with a Macro F1 of 45.2% and 32.4% respectively. The Macro F1 scores of both ChatGPT and Gemini suggest that they are not reliable for independent clinical diagnoses in cardiology, with difficulty shown in interpreting ECGs for rhythm and axis. This investigation aims to guide continued improvement of these LLM models for physician assistance.
Large language models (LLMs) are increasingly used in health care, with emerging applications in clinical decision support and nursing education. However, evidence on their performance in nursing contexts, particularly in oncology nursing, remains limited. Given the complexity and high-risk nature of oncology care, it is important to evaluate the performance and clinical relevance of LLM-generated responses in oncology nursing contexts. This study aimed to compare the performance of LLMs in oncology nursing decision support tasks using standardized examination questions and case-based clinical scenarios and explore LLMs' potential applicability and current limitations in oncology nursing practice. A total of 33 case-based questions derived from 10 oncology nursing clinical scenarios in a nationally used training manual, along with standardized examination-oriented questions from a commercially published preparation book for the Chinese Nursing (Intermediate) Qualification Examination, were used to evaluate the performance of 5 LLMs (DeepSeek, Qwen, Spark-Desk, WiseDiag, and ChatGPT). All models generated responses using a standardized prompt. Two oncology nurses with more than 5 years of clinical experience independently rated the case-based responses using 3 evaluation dimensions: correctness, clarity, and conciseness. Interrater reliability was assessed using the quadratic weighted Cohen κ, intraclass correlation coefficient, and Spearman rank correlation coefficient. Differences among models were analyzed using the Kruskal-Wallis test with the Dunn post hoc test. In addition, examination performance was evaluated based on total score, accuracy rate, and completion efficiency. Interrater reliability analyses indicated moderate agreement between evaluators. The median correctness, clarity, and conciseness scores were as follows: 11.50 (IQR 10.50-12.00) for DeepSeek, 11.00 (IQR 10.50-12.00) for Qwen, 10.50 (IQR 9.50-11.50) for Spark-Desk, 10.00 (IQR 9.50-11.50) for WiseDiag, and 10.00 (IQR 9.00-11.50) for ChatGPT. The Kruskal-Wallis test indicated statistically significant differences among models (H=11.416; P<.05), with post hoc analysis showing a significant difference only between DeepSeek and ChatGPT (P<.05). In examination-based tasks, all models achieved passing performance, with accuracy rates ranging from 77% (77/100) to 93% (93/100). In terms of response completion, DeepSeek and ChatGPT completed all tasks in a single interaction, whereas other models required multiple interactions due to output interruptions. LLMs showed relatively strong performance on structured knowledge and examination-based tasks but remained limited in complex oncology nursing scenarios requiring individualized assessment and dynamic clinical judgment. Their potential use may be most relevant to information retrieval, knowledge organization, and patient education. Because the correctness, clarity, and conciseness rubric showed only moderate interrater reliability, the case-based comparisons should be interpreted as preliminary signals rather than definitive evidence of between-model differences. LLM outputs should therefore be used as supportive information and interpreted alongside professional clinical judgment.
Accurate assessment of pulpal status is essential for achieving successful endodontic outcomes. However, direct evaluation remains inherently challenging because the pulp is surrounded by calcified tissue, necessitating reliance on clinical and radiographic examinations for diagnostic and prognostic decision-making. These procedures demand substantial clinical expertise and time, and less-experienced clinicians often face challenges that may lead to errors in diagnosis and treatment planning. Recent advancements in large language models (LLMs) offer promising opportunities to enhance clinical reasoning by facilitating the integration of evidence and supporting methodical diagnostic decision-making. This study aimed to evaluate the clinical applicability of LLMs by comparing their text-based clinical screening performance and the clinical validity of their treatment plan responses with those of human evaluators. Between January 2011 and December 2022, 100 clinical cases involving primary endodontic disease were randomly selected from the clinical records of outpatients who visited the Department of Conservative Dentistry or Advanced General Dentistry (AGD) at Yonsei University Dental Hospital. Four prompt types, combining 2 variables (language and role), were used as input for 4 LLMs. Both LLMs and human evaluators (AGD specialists, AGD residents, endodontic residents, and senior dental students) assessed the cases using text-based clinical records. Radiographic images were not directly provided. Screening performance was evaluated using a 0-to-2-point concordance scale, and treatment plan validity and relevance were assessed using a 5-point Likert scale. Among the 4 LLMs evaluated, ChatGPT achieved the highest mean concordance score on Korean-doctor prompts (mean 0.98, SD 0.82). However, this score did not reach the partially correct criterion of 1 on the 0 to 2-point scale. Clova X recorded the lowest mean score on English-patient prompts (mean 0.23, SD 0.63). Across both diagnostic categories, AGD specialists demonstrated the highest diagnostic accuracy (pulpal: 0.70; periapical: 0.65), with higher sensitivity but lower specificity than those exhibited by the other groups. ChatGPT also showed favorable performance among the LLMs, with accuracies of 0.65 (95% CI 0.55-0.74) for pulpal disease and 0.57 (95% CI 0.47-0.69) for periapical disease, which were comparable to those of AGD and endodontic residents. Under image-free clinical record review conditions, ChatGPT 4.0 showed relatively higher and more consistent performance in symptom screening and treatment planning compared to the other LLMs evaluated. However, its highest mean score of 0.98 (SD 0.82) did not reach the partially correct criterion of 1 on the 0 to 2-point scale. Hallucinations generated by LLMs and experience-dependent interpretation biases among human evaluators remain key challenges that require attention. Therefore, continuous clinical supervision and comprehensive user training are necessary for the safe and effective clinical application.
The heterogeneity of health data systems remains a major barrier to interoperability and data integration, and a challenge for clinical decision support and AI-driven research. The coexistence of multiple Common Data Models (CDMs) complicates integration, requiring complex manual or rule-based schema mapping, which is labor-intensive, error-prone, and difficult to scale. We evaluated three leading Large Language Models (LLMs): ChatGPT, Gemini, and DeepSeek, for their ability to match data elements from four widely used CDMs (PCORnet, OMOP, Sentinel, and i2b2) to FHIR R4 via web-based chat interfaces. Reference mappings from the HL7 Common Data Models Harmonization Implementation Guide were used as the benchmark. Standardized prompts were submitted across 36 total runs (3 LLMs × 4 schemas × 3 repetitions), and accuracy was measured using F1-scores with 95% confidence intervals, alongside strict (AND) and relaxed (OR) match criteria. DeepSeek obtained the highest numerical mean F1-scores (0.78-0.99) with strong consistency (AND Match = 0.87). Gemini showed good coverage (OR Match = 0.96) but greater variability (AND Match = 0.64), while ChatGPT achieved moderate accuracy with lower consistency (AND Match = 0.51). A two-way ANOVA indicated no statistically significant differences by LLM type (p = 0.138), CDM schema (p = 0.194), or their interaction (p = 0.451). Performance varied by domain, with Demographics and Condition mapped most accurately, while Observation showed the highest error rates (44% of total errors). Out of 2,583 total matchings, 314 errors were observed and categorized into four failure patterns: semantic confusion (42%), structural misalignment (28%), ambiguous terminology (18%), and terminology variance (12%). This study demonstrates the feasibility of using LLMs to automate health data integration as a complementary approach to traditional ontology-based methods, particularly for well-structured, context-rich CDMs. However, safe deployment requires human-in-the-loop validation, governance, prompt optimization, metadata integration, and validation to ensure clinical and research reliability.
Persistent gaps in faculty capacity hinder the integration of genomics into nursing curricula, highlighting the need for strengthened faculty development. The NIH-funded R25 "Translation and Integration of Genomics is Essential to Doctoral Nursing" (TIGER) program is designed to strengthen genomic literacy among nursing faculty. Grounded in adult learning theory, TIGER incorporates SMART (Specific, Measurable, Achievable, Relevant, and Time-bound) goal development as a pedagogical strategy. Midway through the program, participants engage with faculty and peers to review progress, share implementation approaches, and address barriers. Participants (N = 62) developed SMART goals spanning genomic education, research, and practice. ChatGPT-assisted analysis of TIGER 2022 to 2024 goals identified key themes, including strengthening personal genomic knowledge, conducting curriculum mapping, integrating genomics into courses, and advancing genomics-focused scholarship. Adult learning theory fosters relevant, empowering learning, while SMART goals offer a structured, actionable pathway for nursing faculty to achieve sustainable genomic integration in curricula.
Despite the goal of utilization management (UM) to facilitate high-value cancer care, concerns about care denials persist. Evaluating justifications for denied claims can clarify the quality of UM decisions. To examine second-level appeals for denied anticancer medications in Medicare Part D to assess the quality of UM decisions. This cross-sectional study used a large language model (ChatGPT 5 Mini, OpenAI) to analyze texts of second-level appeals decisions on denied Part D medications, rendered by a Medicare-contracted reviewer. Data were analyzed from March 13, 2025, to May 12, 2026. From decision summaries, the reasons for first-level denials (as cited by Part D plans) and the outcomes of second-level appeals (favorable vs unfavorable) were examined. The reasons for unfavorable reviews (clerical vs nonclerical) were also examined. We included 4952 second-level appeals from 2020 to 2024 by beneficiaries with hematological (n = 1155 [23.3%]), prostate (n = 876 [17.7%]), digestive (n = 650 [13.1%]), breast (n = 613 [12.%]), and lung, bronchus, and respiratory (n = 430 [8.7%]) cancers. Most appeals were for targeted therapy (n = 2889 [58.9%]) or hormonal therapy (n = 959 [19.4%]). The most common reasons for initial denial or appeals were "non-medically acceptable" indications (ie, off-label use; 3410 [68.9%]), formulary exceptions (545 [11.0%]), and failure to meet preapproval coverage criteria (mostly on-label prescribing; 463 [9.3%]). Overall, 19.5% (95% CI, 18.4%-20.6%) received favorable reviews, and 55.2% (95% CI, 53.8%-56.6%) and 25.3% (95% CI, 24.1%-26.5%) received unfavorable reviews due to clerical and nonclerical issues, respectively. Favorable reviews were more prevalent among appeals for drugs initially failing to meet preapproval coverage criteria (61.1% [95% CI, 56.7%-65.6%]), but less so among appeals for off-label therapies (16.5% [95% CI, 15.2%-17.8%]). More than 87.9% (95% CI, 86.1%-89.6%) of successfully appealed off-label therapies were for medically acceptable indications supported by Medicare-approved clinical compendia and/or peer-reviewed literature. In this cross-sectional study of second-level appeals for denied anticancer medications in Medicare Part D, a considerable number of successfully appealed claims (especially for therapies prescribed on label) were found, which suggests potentially inappropriate UM decisions by Part D plans. However, the appeal outcomes varied substantially by the nature of the denials. Streamlining the preapproval process for on-label therapies and establishing accessible, transparent criteria for off-label prescribing may optimize UM processes for anticancer medications.
As educational digitalization continues to deepen and generative artificial intelligence rapidly enters higher education teaching contexts, faculty professional development is undergoing a profound transformation from tool use to pedagogical redesign and from individual adaptation to organizational change. However, current faculty development practices still tend to substitute standardized training for differentiated support, making it difficult to address the diverse needs of teachers at different career stages in digital competence, curricular innovation, disciplinary integration, and academic leadership. To enhance the transparency of the framework-building process, this study was informed by explicit review questions, database-based literature retrieval, and thematic synthesis. Building on scholarship on teacher digital competence, teacher professional development theory, and the practice of digital transformation in higher education, with Heilongjiang University of Chinese Medicine serving as the institutional context for framework formulation, this paper offers a systematic account of the stage-specific developmental tasks of university faculty in the digital era and advances a multi-level, developmental framework. The framework is organizationally supported by coordination across three institutional levels-university, school, and department-structured around four developmental stages: entry and adaptation, stable development, deep integration, and academic leadership, and implemented through five pathways: blended professional learning, cross-boundary linkage, community-based mutual learning, diagnostic mentoring, and evidence-informed practice. The study contends that faculty development should move beyond a narrow focus on digital tool training and instead emphasize the systematic development of capacities related to pedagogical improvement, disciplinary reconfiguration, digital ethics, and organizational leadership. This framework helps enhance the stage-specific relevance, practical explanatory power, and organizational operability of faculty development, while also providing a theoretical reference for designing faculty development systems in digitally transformed universities.
This review examines how pharmaceutical-inspired discovery logic can help revitalize insecticide innovation by integrating target-based design, chemoinformatics, and artificial intelligence into a more structured discovery pipeline. It outlines the innovation deficit in insecticide discovery and explains why pharmaceutical concepts such as validated target selection, hit-to-lead progression, multi-parameter optimization, and Design-Make-Test-Analyze cycles provide a useful conceptual model for insect-control discovery. The review evaluates established and emerging insect molecular targets, including classical neurophysiological targets and underexploited insect-selective pathways, with attention to structural tractability, ortholog-based selectivity, and resistance relevance. It further synthesizes the roles of chemoinformatics and artificial intelligence in molecular representation, virtual screening, activity prediction, structure-based design, active learning, generative design, and safety-aware optimization. Particular attention is given to the opportunities and limitations of these approaches in the context of sparse insect-specific datasets, uneven assay standardization, applicability-domain constraints, and the persistent gap between computational promise and field-usable products. The review also considers repurposing strategies, scaffold innovation, selectivity and pollinator safety, environmental sustainability, resistance-informed design, and the translational barriers that limit movement from in silico leads to deployable insecticides. Overall, the evidence suggests that the strongest future for insecticide discovery lies not in artificial intelligence alone, but in a connected discovery ecosystem that links target biology, structural insight, chemistry, predictive modeling, validation practice, and sustainability-oriented design.
暂无摘要(点击查看详情)
Generative artificial intelligence is increasingly embedded in learning as a source of explanation, feedback, and interactive support. However, it remains unclear how such support shapes learners' cognitive processing and their subsequent willingness to continue learning. From an educational psychology perspective, this study examines how information quality, perceived ease of use, and perceived interactivity are associated with sustained learning intention through intrinsic and extraneous cognitive load, and whether prior cultural knowledge moderates the associations between cognitive load and sustained learning intention. This cross-sectional survey study was conducted in The Art of Life: Mawangdui Han Culture Immersive Digital Exhibition as a high-complexity digital cultural learning context. Survey data were collected from 572 learners who reported a recent and complete experience of using generative AI tools for understanding or further learning content related to Mawangdui Han culture. Partial least squares structural equation modeling was used to test net effects, statistical indirect associations, and moderation, and fuzzy-set qualitative comparative analysis was used to identify configurations associated with high and low sustained learning intention. Information quality and perceived ease of use were significantly negatively associated with both intrinsic and extraneous cognitive load, whereas perceived interactivity was significantly positively associated with both forms of load. Intrinsic and extraneous cognitive load were both negatively associated with sustained learning intention. Prior cultural knowledge weakened the negative association between extraneous cognitive load and sustained learning intention but strengthened the negative association between intrinsic cognitive load and sustained learning intention. The fsQCA results further revealed multiple asymmetric configurations associated with high and low sustained learning intention. These findings suggest that the value of generative AI-supported learning in this context lies not in maximizing interaction, but in providing cognitively manageable support. In this study, sustained learning intention refers to a self-reported intention measure rather than observed long-term learning behavior.
As Academic Medicine marks its centennial, academic medical centers (AMCs) face a defining inflection point driven by the convergence of generative artificial intelligence (AI), machine learning, computational biology, digital health platforms, autonomous systems, and increasingly continuous streams of clinical and behavioral data. Together, these technologies are transforming how health professionals are educated, how biomedical knowledge is generated, how care is delivered, and how health outcomes are measured. In this Commentary, the authors argue that these changes require more than technological adoption; they demand a redefinition of the mission and responsibilities of academic medicine. AMCs must move beyond their traditional roles in education, research, and clinical care to become leaders in responsible innovation, ethical governance, workforce transformation, public trust, community engagement, health equity, and global stewardship. The authors examine implications for individualized education, AI-enabled clinical care, accelerated scientific discovery, and the democratization of health knowledge across diverse settings. The authors contend that AMCs must deliberately redesign curricula, research ecosystems, care models, and institutional incentives to ensure that technological advances strengthen rather than diminish human judgment, compassion, and equity. The next century of academic medicine will be defined not by whether AMCs adopt emerging technologies, but by whether they shape their development and use in ways that preserve the human relationships and societal responsibilities at the heart of medicine.
To systematically evaluate the capability of the Large Language Model, Generative Pre-Trained Transformer (GPT-4o), to deliver Motivational Interviewing. Large Language Models are a type of Artificial Intelligence trained to understand and generate human language. OpenAI's GPT-4o was prompted to conduct Motivational Interviews with simulated patient actors. The interview transcripts were analysed using the Motivational Interviewing Treatment Integrity code, which includes thresholds for assessing competency in Motivational Interviewing. Conversation sequences were furthermore examined to understand GPT-4o's patterned responses across Motivational Interviewing behaviours. A total of 36 Motivational Interviewing transcripts were generated covering a range of chronic conditions and health behaviours. GPT-4o performed above 'good' competency thresholds for behaviour counts (Percent Complex Reflection and Reflection to Question ratio). Total Motivational Interviewing-Adherent behaviour counts were very high with very few Total Non-Motivational Interviewing-Adherent counts. GPT-4o did not perform as well on global competency scores, performing below 'fair' competency thresholds for both relational and technical skills. A distinct pattern was observed across the interviews where GPT-4o gradually shifted away from Motivational Interviewing towards attempts to persuade participants to change, which is inconsistent with the principles and processes of Motivational Interviewing. GPT-4o has some capacity to mimic Motivational Interviewing behaviours but requires further augmentation to perform it with proficiency. Using Artificial Intelligence has the potential to increase accessibility and reduce the costs of delivering Motivational Interviewing. Further research is needed to explore effective augmentation techniques before Large Language Models such as GPT-4o can be used to deliver Motivational Interviewing.
Many patients now turn to generative AI for health advice before seeing a doctor. However, there is little evidence about how well it works in primary care. Research focused on general practice is limited, and real-world harms are often missed because there is no standard way to record AI-related incidents. The gap between what AI provides and what general practice needs is not just technical; it is structural and requires active oversight. AI can recognise patterns, summarise documents, and organise data, but it does not have the judgement, long-term patient knowledge, or moral reasoning that define good general practice. Evidence shows that careful use of AI can help, but careless use can cause real harm. To safely use AI, we must protect what makes primary care effective: knowing patients over time, reasoning through diagnoses across visits, and building healing relationships that combine clinical skill with whole-person care. No training set or workforce change can replace this. Evaluation should look beyond accuracy, considering continuity of care, the risk of losing clinical skills, patient experience, and fairness. General practitioners (GPs) must take an active role in guiding how AI is used, rather than simply adding it to current practice.