Large language models (LLMs) have shown promise in medical education, but most are trained primarily on Western medical data and lack sufficient exposure to Traditional Chinese Medicine (TCM). Despite growing interest in LLMs for TCM education, systematic evaluations of their performance remain limited. However, it remains unclear at which cognitive levels-from basic recall to clinical reasoning-these models can reliably support learners, and whether their multimodal capabilities meet the demands of authentic TCM pedagogical tasks. This study aimed to evaluate LLMs' performance on authentic Chinese medicine knowledge tasks, including both text-based and image-based assessments, to provide evidence for their potential application in TCM education. Seven LLMs were evaluated across four task types: textual objective questions, textual subjective syndrome differentiation tasks, image-based objective herb identification, and image-based subjective tongue diagnosis. For objective tasks, performance was measured by accuracy. For subjective tasks, performance was evaluated using natural language processing metrics and blinded expert ratings. Statistical analyses included Cochran's Q test, Friedman rank-sum test, and Bonferroni-corrected post-hoc comparisons. Inter-rater reliability was verified using the intraclass correlation coefficient. Significant tiered performance differences emerged across 7 models. For textual objective questions, Doubao-1.5-Pro achieved the highest accuracy (91.23%), with performance declining markedly from lower-order memorization to higher-order clinical reasoning tasks. For textual subjective tasks, SparkDesk-v3.5 (94.11%), Doubao-1.5-Pro (93.03%), and Gemini-3-Preview (92.99%) formed the top tier. A notable subjective-objective performance inversion was identified: SparkDesk-v3.5 ranked lowest in objective accuracy but highest in subjective scoring. Our exploratory image-based results suggest that no model exceeded 60% accuracy in herbal identification (N = 55), while tongue diagnosis performance was comparatively higher but should also be interpreted cautiously (N = 223). Qualitative analysis using a four-category error typology revealed visual misidentification, cross-modal association failures, terminological-level errors, and reasoning-level inconsistencies. Top-tier large language models demonstrate potential as virtual teaching assistants for foundational Traditional Chinese Medicine knowledge navigation. However, their deployment in complex syndrome differentiation and multimodal teaching poses substantial risks due to model errors and insufficient clinical reasoning. Image-based findings should be considered preliminary and require validation in larger, more balanced multimodal datasets. Clearly defined applicability boundaries, rigorous manual review procedures, and continuous enhancement of domain-specific multimodal datasets are essential before integration into educational practice.
使用 AI 将内容摘要翻译为中文,便于快速阅读
使用 AI 分析这篇文章的核心发现、关键要点和深度见解
由 DeepSeek AI 提供分析 · 首次使用需配置 API Key
PubMed · 2026-01-01
PubMed · 2026-01-01
PubMed · 2026-01-01
PubMed · 2026-01-01