This study aims to compare the performance of different large language models (LLMs) in pre-exam learning for the Chinese dental licensing examination, with a focus on evaluating their differences in answering, explanation, and teaching effectiveness, to provide a reference for the application of LLMs in dental education. Three evaluation scenarios were designed: selecting correct answers, providing answer explanations, and adversarial testing. DeepSeek-R1, Qwen 2.5-MAX, Doubao 1.5 Pro, Xinghuo Spark-X1, ERNIE 4.0 Turbo, GPT-4o, and Huaxi Zhilian were selected for comparative testing. Evaluation metrics included accuracy, net accuracy, and pedagogical effectiveness. In the scenario of selecting correct answers, all LLMs exceeded the passing threshold, with Huaxi Zhilian achieving the highest accuracy (84%). In the answer explanation scenario, Huaxi Zhilian demonstra-ted the highest net accuracy (92%), followed by DeepSeek-R1 (89%), among models. Regarding pedagogical effectiveness, Huaxi Zhilian ranked highest in relevance, practicality, and clarity, whereas GPT-4o led in conciseness, among the investigated LLMs. In adversarial testing, Huaxi Zhilian and DeepSeek-R1 exhibited the smallest declines in accuracy and net accuracy, respectively, among the tested models. In pre-exam learning for the Chinese dental licensing examination, knowledge-enhanced LLMs specifically optimized for dentistry (e.g., Huaxi Zhilian) outperform reasoning LLMs pretrained on general Chinese corpora (e.g., DeepSeek-R1) and those primarily trained on English corpora (e.g., GPT-4o). However, the performance of all models declines under adversarial conditions. Future research should focus on addressing identified weaknesses to enhance the utility of LLMs in dental education further. 目的: 比较不同的大语言模型(LLM)在口腔执业医师考前学习中的应用表现,重点评估其在答题、解析和教学效果方面的差异,为LLM在口腔医学教育中的应用提供参考。方法: 设计3个评估场景,包括选择正确答案、答案解析和对抗性测试。选择DeepSeek-R1、通义千问2.5-MAX、豆包1.5 Pro、星火Spark-X1、文心一言4.0 Turbo、GPT-4o和华西口腔智联大模型进行对比测试,评估的指标为准确率、净正确率和教学效果评分。结果: 在选择正确答案场景方面,所有LLM均达到了及格线,华西口腔智联大模型的准确率最高(84%)。在答案解析场景方面,华西口腔智联大模型的解析净正确率最高(92%),其次为DeepSeek-R1(89%)。在教学效果评分中,华西口腔智联大模型在相关性、实用性和清晰度方面得分最高,GPT-4o在简洁性方面分数最高。在对抗性测试场景下,华西口腔智联大模型和DeepSeek-R1分别在准确率和净正确率方面下降幅度最小。结论: 在口腔执业医师考前学习应用场景中,针对口腔医学进行知识增强的LLM(如华西口腔智联大模型)表现优于以中文语料库为核心预训练的推理LLM(如DeepSeek-R1),优于以英文语料库为核心预训练的LLM(如GPT-4o)。在对抗性测试中,所有模型的表现均有所下降。未来相关模型的研究可针对评估中表现不足之处进行优化,提升其在口腔执业医师考前学习中的应用效果。.
使用 AI 将内容摘要翻译为中文,便于快速阅读
使用 AI 分析这篇文章的核心发现、关键要点和深度见解
由 DeepSeek AI 提供分析 · 首次使用需配置 API Key
arXiv · 2024-05-28
arXiv · 2025-11-18
arXiv · 2022-05-17