Patients increasingly consult large language models (LLMs) before specialist review and may do so in widely differing emotional registers. We examined whether re-casting an unruptured intracranial aneurysm (UIA) case from a third-person vignette into a first-person patient account, with anxiety about treatment or non-treatment, alters the recommendations LLMs return. Sixty-seven UIA cases previously discussed at our multidisciplinary team (MDT) were presented to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking under four conditions: third-person vignette (3P), neutral first-person (1PN), and first-person accounts anxious about treatment (1PAT) or non-treatment (1PANT). Each case was submitted five times (4,020 prompts). Majority recommendations were benchmarked against MDT consensus (Cohen's κ, McNemar), tested for directional asymmetry (Bowker), classified as progressive or regressive, and linguistically analysed. Concordance with MDT was fair-to-moderate (71.6-80.0%; κ 0.34-0.51). ChatGPT over-treated under all conditions (p = 0.001-0.013) and Gemini over-treated only at 3P baseline (p = 0.001), abolished under first-person framing. Claude showed no directional preference. First-person framing drove significant clipping-to-coiling migration in Gemini and ChatGPT (p = 0.001 and 0.015). Claude alone hedged under treatment-directed anxiety (p = 0.002). Anxiety-induced shifts were predominantly regressive. Gemini under rupture-anxiety combined the highest empathy-opener rate (85%), highest confidence-marker density and lowest hedging density of any cell, a sycophantic phenotype invisible to categorical analysis. Patient persona and affective framing measurably and model-specifically alter LLM recommendations for UIAs, with anxiety-induced shifts moving models away from rather than toward expert consensus. Clinicians should anticipate AI-shaped expectations varying with a patient's emotional state and counsel against framing-induced advice. This study was retrospectively registered on the Open Science Framework (https://doi.org/10.17605/OSF.IO/HCU6D).
This study was aimed to compare the efficacy of three most popular large language models (LLMs)-Claude Opus 4.6, ChatGPT Thinking 5.4 and DeepSeek v3.2 in answering frequently asked questions (FAQs) about scoliosis. 20 scoliosis related questions (four categories, five questions in each category) were submitted to each LLM. A panel of 9 experts (two spine surgeons, two pediatric orthopedic surgeons and five physical therapists, all blinded to the LLMs and responses) rated independently each response generated by LLMs on a 6 points Likert scale (1 as strongly disagree to 6 as strongly agree). 540 total ratings were collected. Intergroup comparisons were conducted by Kruskal Wallis test and Mann Whitney U pairwise tests. Paired question level analysis was achieved by Friedman test and Wilcoxon signed rank comparisons. Claude's score was 5.53 ± 0.76 much higher than both ChatGPT (4.84 ± 0.86, p < 0.001) and DeepSeek (4.86 ± 0.84, p < 0.001), but no difference was found between ChatGPT and DeepSeek (p = 0.749). Claude performed on top for 19 of 20 questions (95%) and was favored by 7 of 9 reviewers. Consistency of Claude was also highest [CV = 13.8% vs. 17.9% (ChatGPT) and 17.2% (DeepSeek)]. Although all three LLMs achieved favorable overall ratings (>4.8/6), Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency. Within the scope of the present evaluation, Claude demonstrated the strongest overall performance among the three LLMs tested.
Careful evaluation of research methodology is fundamental to scientific progress but represents a significant burden on human experts. The complexity of functional MRI (fMRI) methods makes transparent reporting, as suggested by OHBM COBIDAS guidelines, particularly critical. Large Language Models (LLMs) present a potential solution for rapid, scalable methodological assessment. We evaluated three state-of-the-art LLMs (Gemini 2.5 Pro, Claude 4 Sonnet, ChatGPT-o3-pro) against human expert ratings. Fifty fMRI articles (taken from 2016 to 2025) were independently evaluated by ten human experts and three LLMs using an 82-item COBIDAS based rubric. Human raters demonstrated excellent inter-rater reliability (ICC = 0.801), while LLMs showed poor internal agreement (ICC = 0.254). When comparing total scores across papers, Gemini showed strong positive correlation with human consensus (r = 0.693, p < 0.0001), Claude showed moderate positive correlation (r = 0.394, p = 0.004), while ChatGPT showed negative correlation (r = -0.172, p = 0.233). Gemini maintained high reliability when added to human raters (combined ICC = 0.811), achieving 85.3 % exact agreement and 98.8 % within-1-point agreement. Domain-specific analysis revealed Gemini's consistently high agreement across all six COBIDAS sections (experimental design: 0.915, statistical modeling: 0.880), while ChatGPT and Claude showed weaker, more variable performance. Obvious differences emerged in determining non-applicable items: humans marked 40.5 % as not applicable versus 32.3 % for Gemini, 9.2 % for ChatGPT and 21.1 % for Claude. ChatGPT exhibited extreme score volatility, with papers ranging from 0 to 121 points compared to humans' 44.2-77.7 range. LLM scoring required 1-7 min versus 30-35 min for humans. This proof-of-concept study demonstrates that LLM-assisted methodological evaluation is feasible for complex neuroimaging research and could likely be applied to other research fields.
To evaluate the performance of four artificial intelligence (AI) systems (ChatGPT 4o, Claude 3.7, Gemini 2.0, and Grok 2) in analysing nasal deformities. The artificial intelligence chatbots were compared to experts in terms of their capacity to analyse nasal deformities. A quantitative analysis compared AI-generated MIRA scores with expert MIRA scores using error measures, Bland-Altman analysis, concordance metrics, and intraclass correlation coefficients to evaluate agreement and systematic bias. A qualitative evaluation was conducted using a 5-point Likert scale to characterise the major nasal type (tension nose, saddle nose, deviated nose, etc.). Fifty adult patients seeking rhinoplasty were evaluated by the chatbots and two experts based on standardised photographs. The evaluations by the two surgeons demonstrated very strong concordance (ICC = 0.997) for nasal analysis using the MIRA scale. Only Claude 3.7 and the experts had comparable total MIRA score evaluations (p > 0.05). Detailed analysis of MIRA sub-scores showed a significant difference between chatbots and experts across all models (p < 0.05), including Claude. Grok 2 (p < 0.001) demonstrated the poorest performance. The qualitative description of the nose by ChatGPT 4o achieved the best results, with an accuracy rate reaching 70%. No model achieved significant performance on MIRA sub-scores in the quantitative analysis of nasal deformity. The qualitative assessment shows that ChatGPT4o could assist, under supervision, with rhinoplasty assessments to analyse major nose types. However, it was effective in only two-thirds of cases. To date, AI tools are not reliable for analysing nasal deformities. This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .
The current study aimed to quantify the diagnostic accuracy of commonly utilized chatbots including Gemini, Copilot, Claude, and specialized architectures like Manus in the detection and differential diagnosis of various jaw lesions, while concurrently evaluating the clinical safety and fidelity of the information they provide. Cone beam computed tomography (CBCT) dataset from 97 patients presented with jaw lesions were collected and anonymized. Panoramic 2D views were reconstructed from Digital Imaging and Communication in Medicine (DICOM) of all cases using Bluesky Plan software and provided to 4 chatbots (Gemini 2.5 Pro, Copilot, Claude and Manus). Moreover, the DICOM data was provided to Manus followed by prompting. The reports generated were evaluated for accuracy, relevance and feasibility. Statistically significant differences were detected between the chatbots in all measured parameters. In all evaluated parameters Manus CBCT showed the most accurate results (95% of lesions were detected and correctly diagnosed,). Gemini 2.5 pro ranked second where 80% of lesions were detected and 56% were correctly diagnosed. Manus Pan showed less accurate results. The least accurate results were detected in Copilot and Claude. Significant discrepancies exist among artificial intelligence (AI) chatbots regarding their diagnostic accuracy in reporting jaw lesions. Notably, the integration of raw 3-dimensional CBCT data substantially optimizes chatbot performance in lesion detection and diagnosis, as demonstrated by Manus architecture.
Artificial intelligence (AI) chatbots or large language models (LLMs) are adept at generating language, but their increasing use in the healthcare field, including endodontics, raises concerns about their accuracy. The potential of LLMs to assist clinicians in their decision-making processes regarding vital pulp therapy (VPT) is worth exploring. This study aims to evaluate and compare the responses provided by OpenAI GPT-5.1 Instant, DeepSeek-R1, Claude, Google Gemini, Comet, and Perplexity to clinically relevant questions related to VPT according to the guidelines set by the American Association of Endodontists, European Society of Endodontics, and Indian Endodontic Society. Twenty-three open-ended questions covering various aspects of VPT were developed and presented to OpenAI GPT-5.1 Instant, DeepSeek-R1, Claude, Google Gemini, Comet, and Perplexity. Two experienced endodontists, who were blinded to the different chatbots, evaluated the answers on a 3-point Likert scale. To assess the reproducibility of these answers, the same questions were presented again after 1 month and subsequently saved in a separate Microsoft Word file. The findings were recorded in an Microsoft Excel Sheet, and then statistical analysis was performed. All the LLMs were able to answering all the questions on VPT with almost similar reproducibility across two different intervals. Most tested LLMs, regardless of whether they are free or subscription-based, demonstrated high accuracy and reproducibility when evaluated on guidelines-based questions related to VPT.
Accurate assessment of distal radius fracture stability is essential for appropriate triage and timely referral to hand specialists. The LaFontaine criteria provide a structured radiographic framework for predicting instability but are not routinely reported. Recent advances in multimodal large language models (LLMs) capable of direct image interpretation have generated interest in their potential role as adjunctive diagnostic tools. However, their real-world performance in structured radiographic assessment remains unclear. A cross-sectional diagnostic accuracy and agreement study was performed using 20 distal radius fracture radiographs. Five hand surgeons independently assessed each case for the 5 LaFontaine criteria, with majority agreement serving as the reference standard. Two publicly available multimodal LLMs (ChatGPT and Claude) were evaluated using a standardized, single-prompt approach designed to approximate real-world use. Both models were provided identical radiographs and clinical prompts then asked to classify each criterion and overall fracture stability. Agreement and diagnostic performance were calculated relative to surgeon consensus. Hand surgeons demonstrated high consistency in identifying LaFontaine criteria. Agreement between LLMs and clinician consensus varied across individual features, with several criteria showing limited agreement. Both models achieved similar overall accuracy for fracture stability classification (0.75; 95% CI, 0.56-0.94). ChatGPT demonstrated moderate agreement with surgeon consensus, while Claude showed fair agreement. Agreement between clinician-determined instability and independent operative recommendations was moderate. Multimodal LLMs demonstrated variable and generally limited agreement with clinician consensus in classifying fracture stability. Performance across several criteria approached chance levels, and specificity was limited. These findings should be interpreted as exploratory and hypothesis-generating; current models are not yet reliable for clinical decision-making and require substantial validation before potential use as adjunctive triage tools. Diagnostic Level III.
Enrollment in phase I oncology trials remains low largely because potentially eligible patients are not identified and evaluated quickly enough. Current clinical trial matching systems can identify candidate patients from the electronic health record, but cases with missing or uncertain eligibility data are often routed for offline manual review. This delay impedes clarification and prolongs the final eligibility determination. This study evaluated TrialTriage, a semiautonomous system built on the n8n platform and designed to resolve eligibility ambiguity during prescreening for phase I oncology trials. When eligibility information is missing or uncertain, TrialTriage emails the investigator, captures the reply, and reruns classification within the same workflow. TrialTriage combined large language model-based variable extraction from free-text clinical narratives and investigator email replies with a deterministic rule engine applying a prespecified 7-criterion protocol. Each case was classified as eligible, not eligible, or ambiguous. Ambiguous cases triggered a structured email query to the investigator, followed by reclassification after a reply. Two requests were sent at 24-hour intervals; after 48 hours without a reply, the case was referred for manual review. The system was tested on 90 synthetic patient cases generated independently by Claude Sonnet 4.6, Gemini 3.1, and Grok 4, with 30 cases per model and balanced distributions of eligible, not eligible, and ambiguous cases. Answer keys were reviewed for accuracy before system execution. Five independent reviewers classified the Claude dataset using a uniform survey form. TrialTriage's classifications were 100% concordant with the author-confirmed ground truth in all 90 synthetic cases (95% CI 96.0%-100.0%). All ambiguous cases were correctly escalated to investigator query. The mean processing time was 2.3 (SD 0.5) minutes per 30-case dataset (range 1.8-2.8 min, approximately 3.5-5.5 s per case). The 5 reviewers achieved a mean accuracy of 96.7% (SD 3.3%), with a Fleiss κ of 0.910, and required a mean of 9.8 (SD 4.8) minutes to review 30 cases. In a subset test of 6 first-pass ambiguous cases, 4 of 6 were reclassified definitively after investigator response, while 2 remained ambiguous because the replies lacked actionable information. TrialTriage demonstrates the feasibility of a semiautonomous prescreening workflow in which ambiguous cases trigger an immediate investigator email query and are reclassified after reply capture with new information within the same system. The main contribution is the integration of email ambiguity resolution into the workflow rather than immediate deferral to offline manual review. Because the evaluation used synthetic cases and label definitions aligned with the same protocol rules used to design the rule engine, these findings should be interpreted as proof of concept and implementation fidelity rather than evidence of real-world clinical performance. Prospective validation using data from real-world electronic health records would be a plausible next step.
Validated measures of pain catastrophizing primarily assess catastrophizing as a stable trait. However, emerging evidence suggests catastrophizing fluctuates with context, highlighting a need for ecologically valid methods to capture it. This study evaluated large language models (LLMs) as implicit markers of catastrophizing from free-text responses from ninety-one adults with chronic pain receiving long-term opioid therapy (57.3% Female; mean age = 60.5 years). Patients completed baseline measures, including the trait pain catastrophizing scale (PCS), followed by a 10-minute writing task after random assignment to a negative, positive, or neutral pain-coping condition. State affect and pain were assessed before and after writing tasks and again after a cold pressor task (4°C; ≤ 2 minutes). A state PCS followed the cold pressor task. Free-text responses were analyzed using four LLMs (Claude Opus 4; GPT Mini 4o; Llama 4 Maverick; and Gemini 2.5 Pro). ANOVA-based results supported discriminant validity, as all four LLM-derived pain catastrophizing scores differentiated negative from positive and neutral pain-coping conditions. Convergent validity was model-dependent; only Gemini-derived scores correlated with state catastrophizing (r = .22) and pain unpleasantness (r = .23). Divergent validity was mixed. LLM-derived scores were unrelated to pain intensity, but Gemini and Claude-derived scores showed small correlations with trait PCS (r's = .21; 28, respectively). All LLM-derived scores also correlated with negative affect (r's = .29-.41), comparable in magnitude to state PCS, suggesting limited specificity. These findings provide preliminary evidence that certain LLMs may serve as implicit markers of state pain catastrophizing, but further study is needed.
Large language models (LLMs) have revolutionized biomedical research, yet they remain prone to hallucinations and struggle with the precise, multi-hop reasoning required for biomedical analysis. To bridge this gap between generative capability of AI model and factual rigor, this article introduces GeneGenie, a model-agnostic, multi-agent framework built upon a directed acyclic graph architecture. Unlike static prompting strategies, GeneGenie implements a deterministic five-node pipeline that orchestrates query planning, intelligent retrieval-augmented generation across curated databases (GenCC, HGNC, and UniProt), and the dynamic execution of bioinformatics tools, including NCBI E-Utilities and local BLAST+. We evaluated the system using the updated 16-module GeneTuring benchmark, comprising 1600 question-answer pairs. The experimental design compared six state-of-the-art models-including GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Pro-operating in a standalone "Direct Mode" versus the agentic "Graph Mode." The results demonstrate that the graph-based architecture consistently outperforms single-model baselines across all metrics. Notably, among the six selected LLM models we explored, Gemini 2.5 Pro achieved the highest performance, correctly answering 1158 questions (72.375% accuracy), compared with the best baseline score of only 15.8%. Furthermore, our evaluation utilized an "LLM-as-Judge" semantic assessment, revealing that the agentic approach significantly enhances not only lexical accuracy but also the completeness and factual grounding of responses. While limitations remain in named entity recognition for protein-coding genes, GeneGenie establishes a robust, reproducible paradigm for future biomedical AI systems, proving that tool-augmented orchestration is superior to reliance on raw model scale alone.
Large language models (LLMs) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear. We systematically reviewed LLM performance across clinical tasks in inflammatory arthritis. We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating LLM performance on clinical tasks in inflammatory arthritis. Two reviewers (Y.A., A.G.) screened 113 records. Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4). Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions. Over 20 distinct LLMs were evaluated, including ChatGPT-3.5 to ChatGPT-4o, Gemini 2.0, DeepSeek-R1/V3, Claude, and Perplexity; ChatGPT/GPT variants were the most frequently tested models (16 of 18 studies), so the current evidence base is predominantly GPT/ChatGPT-based. Findings spanned patient education (n=11), guideline adherence (n=6), clinical reasoning (n=3), and other applications (n=1). All readability assessments exceeded recommended thresholds. Guideline concordance ranged from 48% to 96%. Accuracy was lower for case-based clinical scenarios (4.24/6) than FAQ and guideline-based questions (5.32-5.36/6; p=0.044). When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0). LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions. None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows. Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.
Background Clinical histories accompanying imaging orders guide protocol selection and diagnostic focus. However, they are often incomplete, potentially compromising diagnostic accuracy and workflow efficiency. Purpose To evaluate whether large language models (LLMs) can improve the clinical utility of provided imaging indications by leveraging clinical notes. Materials and Methods This retrospective study curated a dataset from deidentified electronic health records at the University of California San Francisco (January 2012 to August 2024), consisting of radiology reports with paired referring clinician-provided and radiologist-curated indications linked to clinical notes. The dataset was stratified across five body systems and five pathophysiologic categories to derive LLM selection and reader study internal test sets. For the reader study, 20 radiologists with 2-25 years of experience compared indications from the referring clinician, radiologist, and best-performing LLMs. Readers scored comprehensiveness, factuality, and conciseness and ranked indications for usefulness in protocoling, usefulness in interpretation, and overall ranking. Models and clinicians were compared using cumulative link mixed models with Tukey-adjusted post hoc comparisons. Results From 28 313 patients (mean age, 59 years ± 20.6 [SD]; 14 912 women), 250 examinations from 247 patients were sampled for the reader study. After nine exclusions, 241 examinations were analyzed, yielding 482 reader-examination evaluations. Indications from the best-performing proprietary (Claude 3.5 Sonnet; Anthropic) and open-source (Qwen 2.5-7B Instruct; Alibaba) LLM were rated as more comprehensive (Likert rating of 5: 37.14% and 28.42%, respectively; both P < .001) and factual (68.05% and 59.75%; both P < .001) than referring clinician indications. The proprietary LLM ranked most useful in protocoling (rank 1: 40.87%; all P < .001), useful in interpretation (44.61%; all P < .001), and overall ranking (44.19%, all P < .001). Comprehensiveness (65.77% of ratings; both P < .001) most strongly influenced overall rankings. Conclusion LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by Yilmaz and Cardoza-Ochoa in this issue.
To validate DT-RAG, a curated retrieval-augmented generation system for dental traumatology decision support, against eight commercial large language models. A knowledge base of 250 curated text units ("chunks") from five authoritative sources (IADT 2020, ESE 2021, Krastl 2021, AAE 2013, Cochrane) was built and deployed with Gemini 2.5 Flash as base model. In Study 1, 99 binary clinical questions were submitted in three runs to DT-RAG and eight LLMs; modal accuracy was compared by McNemar exact tests with Holm correction. In Study 2, seven blinded specialists scored DT-RAG against three frontier LLMs on ten clinical scenarios using a 92-point rubric; differences were estimated by linear mixed-effects regression. DT-RAG achieved 96.0% modal accuracy (95% CI 90.1-98.4), significantly exceeding every commercial LLM (best comparator GPT-5.5 87.9%; paired difference +8.1 pp, 95% CI +2.9 to +14.9; p_Holm = 0.013). The curated knowledge base elevated the base model from 49.5% to 98.0% valid rationale rate, eliminating confabulations in this evaluation (0 vs 21). In Study 2, DT-RAG achieved the highest mean score (82.1/92; 89.3%), significantly exceeding Claude Opus 4.5 (72.6; paired difference +9.6 points, exact Wilcoxon p = 0.016), Gemini 2.5 Pro (60.3) and GPT-4.1 (50.4); all seven evaluators ranked DT-RAG first (Kendall's W = 0.97). DT-RAG, a curated retrieval-augmented configuration, outperformed frontier-tier general-purpose LLMs on this dental traumatology benchmark, with no confabulated rationales observed among the responses assessed. This approach may be applicable to other well-defined clinical domains with authoritative guidelines. Curated retrieval augmentation produced accurate, source-traceable and reproducible responses in dental traumatology under benchmark and simulated-scenario conditions, making every error auditable against its source. Clinical safety requires prospective evaluation.
暂无摘要(点击查看详情)
Injection-related bacterial infections represent a major but under-recognised health issue among people who inject drugs (PWID). Harm reduction interventions (HRIs) like needle and syringe programmes (NSP) and opioid agonist treatment (OAT) could mitigate their burden. This review aimed to identify and synthesise evidence on the effectiveness of HRIs in preventing bacterial infections among PWID. Systematic review with meta-analysis of studies with more than 40 participants from Medline, Embase, Cochrane Library and Web of Science, published between 1990 and 2023 in English or French. We included interventional and observational studies that reported a quantitative effect measure for an HRI conducted in community, harm reduction, healthcare and outreach settings, targeting bacterial infections among PWID, defined as individuals who had injected drugs at least once within the previous year. Twelve studies met the inclusion criteria, for a total of n = 11 611 participants. The primary outcome was the impact of HRI on the prevalence or incidence of bacterial infections, measured as a risk difference, relative risk, number needed to treat, relative risk reduction, odds ratio, incidence rate ratio, hazard ratio or preventable fraction among the unexposed. Risk of bias was assessed using the Newcastle-Ottawa Scale for observational studies and the Cochrane RoB 2 tool for randomised trials. Meta-analysis was performed when at least 3 comparable estimates were available. Overall, the available evidence was sparse and heterogeneous, with substantial variability in study design, intervention definitions and outcome measurement across the 12 included studies. Sterile injecting equipment provision was found protective in 2/6 studies (n = 1938), 1/6 (n = 5209) found increased risk and 3/6 (n = 1323) reported no statistically significant association. OAT was protective in 2 studies (n = 2934) when comparing current PWID or those who had never used OAT to past PWID. Only 1 study (n = 1876) evaluated a combination of these interventions, showing a statistically significant reduction in skin and soft tissue infections. Among hygiene interventions, 1 of 2 studies (n = 59) reported a statistically significant protective effect, and the same for drug consumption rooms (n = 665). Overall, 8/10 studies assessed were judged to be at high risk of bias. A random-effect meta-analysis of crude odds-ratios (ORs) associated with NSPs yielded a pooled OR of 1.25 (95% confidence interval = 1.07-1.47). Evidence on the effectiveness of harm reduction interventions in preventing bacterial infections among people who inject drugs is limited and inconsistent, as most studies are observational, focus on skin and soft tissue infections and present substantial methodological limitations.
Protein kinase CK2 is the subject of numerous studies in medicinal chemistry due to its involvement in the development of several diseases, primarily cancers. Its overexpression in tumor cells is related to key processes such as tumor immune evasion and cell proliferation. The scientific approach of this study aims to investigate the thermal shift assay (TSA) as a pre-screening tool and to complement it with a co-crystallization approach in post-screening. Therefore, the synthesis of seven small-molecule CK2 inhibitors derived from indeno[1,2-b]indoles was supplemented by 18 related derivatives from our in-house compound library. The 25 molecules belong to four sub-scaffolds, namely 4b,9b-dihydroxy-4b,5,6,7,8,9b-hexahydroindeno[1,2-b]indole-9,10-dione (D-0), 5,6,7,8-tetrahydroindeno[1,2-b]indole-9,10-dione (D-1), 9-hydroxy-5H-indeno[1,2-b]indol-10-one (D-2), and 5H-indeno[1,2-b]indole-6,9,10-trione (D-3). The most active CK2 inhibitors identified by capillary electrophoresis (CE)-based assay belong to the D-1 sub-scaffold. In the TSA, these compounds also generate significant shifts of the melting temperature (Tm) of CK2, indicating a clear correlation between the results of the CE-based assay and those of the TSA. The contribution of co-crystallization in post-screening also demonstrated the effectiveness of D-1 sub-scaffold compared with D-0 sub-scaffold.
Large language models (LLMs) are increasingly used in higher education, but multi-country evidence on dental students' use, verification, and integrity practices is limited. To compare senior dental students' LLM use, perceived time and academic impact, reliability judgements, verification practices, and integrity safeguards across five countries. An anonymous cross-sectional online survey was administered to final-year dental students in the United Arab Emirates (UAE), Jordan, Malaysia, Oman, and Brazil. Measures included tools used, frequency and motivations, learning activities, perceived time and academic impact, verification frequency and strategies, guideline awareness, and integrity safeguards. Analyses used Kruskal-Wallis and chi-square tests with Benjamini-Hochberg adjustment, effect sizes, Spearman correlations, and ordinal logistic models. In total, 454 students participated (UAE 160, Jordan 101, Malaysia 75, Oman 62, Brazil 56; mean age 22.9; 74.9% female). ChatGPT predominated (95.9%), followed by Gemini, formerly Bard (18.0%), DeepSeek (16.4%), and Claude (7.4%). Tool diversity varied across country-based cohorts, with Oman showing greater multi-tool uptake. Use was frequent (several times/week 39.2%, daily 28.6%). Key motivations were saving time (73.0%), clarifying concepts (56.9%), and summarising (54.1%). Common activities included understanding complex concepts (75.3%), summarising lecture notes (70.0%), exam preparation (61.5%), and assignment research (53.2%); exam-time assistance was reported by 25.6%. Verification was 'always' 20.0% and 'often' 34.1%, varying across country-based cohorts, with Oman verifying less frequently than other cohorts. Guideline awareness was 40.3% overall (UAE 61.3% vs Brazil 8.3%). Integrity safeguards commonly involved paraphrasing (69.6%), citations (39.2%), and plagiarism checks (38.0%); disclaimers were uncommon (9.2%). LLM-use frequency correlated with broader academic use (ρ = 0.289) but not with integrity concern (OR = 0.963). LLM use is widespread and heterogeneous across settings, including non-trivial higher-stakes use. Dental programmes should implement explicit training in verification, evidence traceability, and disclosure, supported by clear, enforceable guidance and assessment designs aligned with real-world LLM practices.
Understanding interfacial heat transfer between polymers and water is crucial for the design of biomaterials, drug delivery platforms, and nano-fluidic systems. In this study, we employed all-atom molecular dynamics (MD) simulations to quantify the interfacial thermal conductance between an infinitely diluted polyethylene glycol (PEG) 36-mer chain and explicit water over the temperature range of 280-350 K. To compare the conformational behavior of the PEG chain, we examined its radius of gyration and observed a temperature-dependent chain collapse consistent with previous coarse-grained models. By employing a transient non-equilibrium MD approach, we imposed temperature difference across the interface and analyzed the energy relaxation behavior to compute heat transfer across the polymer-water interfaces. Moreover, we investigate the impact of polymer conformation on heat transfer by considering modulations of the Lennard-Jones polymer/solvent interactions, different from the original PEG-water interactions. Our results demonstrate that both temperature and Lennard-Jones interfacial interaction strength influence interfacial thermal conductance, with temperature playing the dominant role. Structural factors such as chain conformation and interfacial area were found to mediate the effect of interfacial interaction. Additional analysis of the vibrational density of states and the mean square displacement reveal that vibrational coupling has minimal impact on thermal conductance across interfaces, whereas increased water thermal motion enhances energy transfer. These findings highlight the structural and dynamical origins of interfacial thermal conductance and provide atomistic insights into the tuning of interfacial heat transport in molecular systems through temperature and solvent interactions.
Hallux valgus is the most prevalent forefoot condition and is associated with substantial pain, functional impairment and reduced health-related quality of life. Despite established clinical effectiveness, since 2021 an increasing number of Integrated Care Boards in the UK have classified surgical correction as a procedure of limited clinical benefit, citing a perceived absence of population-level cost-effectiveness data. National-scale evidence is required to inform commissioning decisions and ensure equitable access to care. A cost-utility analysis was performed from the perspective of the UK National Health Service (NHS) using British Orthopaedic Foot and Ankle Society (BOFAS) Registry data for adults undergoing primary hallux valgus correction by osteotomy (open or minimally invasive surgery, MIS). Fusion procedures were excluded. EuroQol-5 Dimension five-level (EQ-5D-5L) utility scores at baseline and 12 months were used to estimate quality-adjusted life year (QALY) gains. A six-state Markov model simulated lifetime costs and outcomes over 40 annual cycles from the UK NHS perspective, with costs and benefits discounted at 3.5% per annum. Incremental cost-effectiveness ratios (ICERs) were calculated against conservative management and deterministic sensitivity analysis was performed across procedural cost, utility gain and benefit duration. A pre-specified subgroup analysis compared open and MIS techniques. From 1111 registry pathways, 306 patients had complete EQ-5D-5L datasets for cost-utility modelling, comprising 129 open and 177 MIS procedures. EQ-5D-5L utility improved from 0.69 (95% CI 0.65-0.72) at baseline to 0.84 (95% CI 0.81-0.88) at 12 months in the open group, and from 0.69 (95% CI 0.66-0.72) to 0.82 (95% CI 0.79-0.84) in the MIS group (both p < 0.001). The base-case lifetime Markov model produced an ICER of £ 8737 per QALY for open correction and £ 11,969 per QALY for MIS correction, both well below the NICE willingness-to-pay threshold. In sensitivity analysis using incremental costs against conservative management, the ICER ranged from cost-saving (-£374 per QALY) to £ 3219 per QALY across all tested scenarios. Open correction was the dominant strategy in the pre-specified subgroup analysis, primarily driven by lower implant costs and higher removal rates in current literature. Hallux valgus correction surgery is highly cost-effective from the UK NHS perspective, with cost per QALY values substantially below those reported for total hip and total knee arthroplasty. The current restriction of access in some UK regions is not supported by national health-economic evidence. III (economic and decision analysis based on prospective registry data).
Timely and highly accurate diagnoses by physicians play a crucial role in improving the quality and effectiveness of patient treatment outcomes. Currently, the use of artificial intelligence capabilities in this area has garnered the attention of many health science researchers. Therefore, the main goal of this preliminary study was to compare the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in definitive and differential diagnoses using standardized clinical vignettes. This descriptive comparative study evaluated the diagnostic accuracy of 10 emergency medicine physicians and 4 large language models (LLMs)-ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1-using 10 standardized clinical vignettes. All LLMs were accessed via official web interfaces. Clinical vignettes were developed from emergency department presentations, validated by an expert panel, and presented as text-only inputs to all evaluators. Diagnostic accuracy was assessed using a standardized scoring protocol: definitive diagnoses required an exact match with expert-derived reference standards, while differential diagnoses required ≥ 3 matches. Overall diagnostic accuracy was the primary outcome. Data were analyzed using Pearson's Chi-square test, McNemar's test, and Generalized Estimating Equation (GEE) logistic regression with Bonferroni correction (SPSS version 28). A total of 280 diagnostic evaluations (10 clinical cases assessed by 14 evaluators across 2 diagnosis types) were analyzed. Overall diagnostic accuracy was 61.79%. AI models demonstrated significantly higher overall accuracy (73.75%) compared to emergency medicine physicians (57.00%, p = 0.014). Across all evaluators, definitive diagnoses were more accurate than differential diagnoses (70.71% vs. 52.86%). Generalized Estimating Equation (GEE) analysis revealed a significant interaction between evaluator group and diagnosis type (p = 0.036). Specifically, physicians experienced a significant decline in accuracy when providing differential diagnoses compared to definitive diagnoses (45.0% vs. 69.0%, p = 0.003; remained significant after Bonferroni correction). In contrast, AI models maintained consistently high accuracy across both diagnosis types, with no significant difference between definitive (75.0%) and differential (72.5%) diagnoses (p = 1.000). Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks. While human physicians struggled significantly with differential diagnoses, AI models maintained high and stable performance regardless of the diagnostic complexity. These findings indicate that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool, particularly in complex scenarios requiring differential diagnostic reasoning. Due to study limitations, such as the small number of clinical scenarios and assessors, these findings should be interpreted with significant caution.