Triaxial compression testing is a widely used method for investigating the elastic and inelastic responses of geomaterials. However, in situ characterization of the process leading to failure at both the specimen and grain scales during such experiments, particularly under elevated confining pressures and reactive environments relevant to geological reservoirs and geotechnology, has remained scarce. This limitation has hindered the advancement of a mechanistic understanding of deformation and failure in geomaterials. To address this gap, we developed a novel high-pressure triaxial apparatus, termed MISTRAL. MISTRAL is a miniaturized triaxial device capable of applying confining pressures relevant to reservoir conditions (up to 100 MPa) while allowing deviatoric loading up to 400 MPa. MISTRAL allows the injection of both reactive fluids (pH 5-9) and carbon dioxide (gas and supercritical) under a pressure of up to 100 MPa. The device is uniquely designed to accommodate in situ x-ray imaging through laboratory sources and operando permeability quantification for the real-time investigation of hydraulic, chemical, and deformation reactions at both the sample and grain scales. In this work, we introduce the design and capabilities of MISTRAL, along with results from its deployment in laboratory environments. The executive drawings of MISTRAL are fully provided to enable reproducibility and deployment. An experiment on a standard sandstone material subjected to a confining pressure of 100 MPa demonstrates MISTRAL's capacity to resolve micromechanical responses under reservoir conditions. The findings underscore the utility of the instrument in advancing our understanding of deformation processes in geomaterials, with implications for both natural and engineered systems.
Mistletoe extract is a widespread complementary therapy mainly used for quality-of-life improvement in cancer patients. Advanced pancreatic cancer is associated with poor quality of life and better therapies for symptomatic relief are highly needed. MISTRAL aimed to assess the impact of mistletoe extract on quality of life, body weight, observed costs and blood biomarkers in patients with advanced pancreatic cancer. MISTRAL was an investigator-initiated, phase III, randomized, double-blind, placebo-controlled, parallel-group, superiority, multicenter, clinical trial with a nested biomarker study. Registration EudraCT 2014-004552-64, NCT02948309. At 9 oncology centers, 290 participants were randomized to standard treatment (palliative chemotherapy or best supportive care) plus subcutaneous mistletoe extract or placebo. Main inclusion criteria were advanced pancreatic cancer, performance status 0-2, main exclusion criteria neuroendocrine pancreatic tumor. EORTC-QLQ-C30, EORTC-QLQ-PAN26, body weight, cost parameters and biomarkers were assessed from baseline up until 9 months. No statistically significant differences for quality of life and weight were evident between treatment arms. Parameters for observed costs for supportive and inpatient care (days at hospital, parenteral nutrition infusions, nutritional supplement drinks, number of visits of palliative home care teams, symptom-relieving medication) were similar in both arms. Thus, calculation of costs was not performed. No effect on explored biomarkers (differential blood count, lymphocyte subpopulations, C-reactive protein, albumin and Ca19-9) was found except for a statistically significant increase of eosinophils in the mistletoe arm without association to clinical effect. Since no benefit was observed, there is no clinical reason to recommend mistletoe extract in patients with advanced pancreatic cancer.
To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes. Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar's odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139. This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions). Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.
Diabetes management increasingly relies on telehealth platforms in which patients generate structured and unstructured data. This unstructured data, in the form of free-text notes often contain additional information beyond the structured data. Extracting this information can enhance patient profiles and optimize treatment. In particular, the extraction of physical activity information from these notes is considered important. This study evaluates rule-based/regex algorithms and a locally deployed Mistral LLM for physical activity information extraction and data augmentation, with their performances benchmarked against a state-of-the-art GPT-4.1. Data from 943 patients collected over 12 years in the DiabMemory system, supplemented by 100 synthetic notes, were analyzed. Patients' privacy was preserved by applying a free text pseudonymization algorithm to all notes and by using locally deployed approaches, thereby avoiding third-party cloud services. Three tasks were conducted: (1) extraction of physical activity (PA) data from free-text notes using regex and a locally deployed Mistral LLM, (2) integration of extracted data with structured activity records using a rule-based approach and the local Mistral LLM, and (3) benchmarking local approaches against GPT-4.1 based on the synthetic notes. Both local methods achieved strong performance in task 1, with minimum F1-scores of 0.84. In task 2, rule-based augmentation (F1 = 0.73) surpassed the Mistral LLM (F1 = 0.37). Task 3 showed GPT-4.1 outperforming the local LLM but not consistently surpassing regex. The rule-based algorithms also required substantially less computation time than either LLM. The regex algorithm achieved superior accuracy and efficiency but required extensive dataset-specific development, while prompt engineering for the LLM required less knowledge and the development time for regex exceeded that of LLM prompt engineering. Findings of this work generally align with prior studies but are limited by the rather small test set and use of synthetic data. Local NLP approaches can enhance structured PA data in diabetes telehealth. Rule-based algorithms remain a strong option where computational resources are limited, though future work should validate these findings on larger and more diverse datasets.
Large language models (LLMs) are increasingly explored as decision-support tools in medical imaging. However, their ability to align with country-specific guidelines, which often diverge, remains uncertain. We set out to evaluate the geographic neutrality of three state-of-the-art LLMs-GPT-o3, Mistral Large, and DeepSeek R1-and a biomedical LLM (MedGemma 1.5 4B), when applied to neuroradiology scenarios with conflicting U.S. and non-U.S. Vignettes derived from contradictory international guidelines were presented to each model under two conditions: an implicit setting, where no guideline was specified and vignettes were provided in English and French; and an explicit setting, where prompts directed models to follow a named guideline. Performance was reviewed against the target guideline, and mitigation strategies were tested. Thirty clinical vignettes presenting conflicting guidelines were evaluated by GPT-o3, Mistral Large, and DeepSeek R1. In the implicit setting, all models favored U.S. guidelines, with GPT-o3, Mistral, and DeepSeek aligning with them in 27 of 30 scenarios (90.0%; 95% CI, 74.4-96.5). In the explicit setting, adherence declined sharply for non-U.S. recommendations for all models. Providing the complete guideline text was the most effective mitigation strategy, restoring accuracies above 90% across all models. Across languages and model origins, LLMs exhibited a systematic bias toward U.S. neuroradiology guidelines, even when explicitly instructed otherwise. This U.S.-centrism likely reflects training data imbalances and raises concerns for safe global deployment. Strategies for local contextualization, such as guideline integration at deployment, are necessary to ensure context-appropriate clinical decision support. Question Do large language models display geographical neutrality in neuroradiology decision support? Findings Even models developed in France and China systematically preferred United States guidelines, aligning with them in most implicit scenarios while failing to follow explicit guidelines from other sources. Clinical relevance This systematic United States-centric bias poses clinical and legal risks for global deployment. Safe implementation requires specific localization strategies, such as providing full guideline texts, to ensure recommendations align with local practice standards.
This study presents a hybrid ontology-based framework for clinical concept extraction from narrative EHR discharge summaries using large language models (LLMs) and standardized biomedical terminologies. The framework integrates multiple NLP components in sequence: SparkNLP for chunk detection and named entity recognition (NER), SentenceBERT embeddings for semantic similarity candidate generation, zero-shot inference with LLaMA3-8B and Mistral-7B for concept selection, and UMLS REST API normalization to CUIs and SNOMED CT terms. This coordinated integration of linguistic, semantic, and ontological modules forms a flexible architecture rather than a single-model comparison. We applied the framework to ten MIMIC-III discharge summaries spanning Chief Complaint, Brief Hospital Course, and History of Present Illness sections. Clinicians labeled extracted concepts as correct, partial, incorrect, missing, or spurious to assess model performance. LLaMA3-8B achieved the highest F1 score (0.77) and lowest false positive rate (3.04%), outperforming both Mistral-7B and cTAKES. While cTAKES demonstrated high precision, it had low recall and a significantly higher FPR (29.95%), indicating frequent misclassification. Mistral-7B offered faster processing for shorter notes, while LLaMA3-8B delivered higher accuracy for more detailed sections. LLMs outperformed traditional rule-based systems by more effectively handling context, modifiers, abbreviations, and multi-word expressions. Prompt refinement and semantic similarity embedding enhanced extraction quality. SparkNLP supported chunking but introduced errors related to spacing and abbreviation handling. We presented a flexible, context-aware framework for clinical concept extraction using LLMs, offering key advantages over rule-based tools. Future work should incorporate full ontology mapping, integrate assertion detection, and validate performance across diverse clinical datasets and domain-adapted LLMs.
Clinical note documentation is a vital yet time-intensive task in health care. While advancements in natural language processing have transformed many domains, generating accurate summaries of doctor-patient conversations remains underexplored due to the limited availability of open-source datasets. Large language models (LLMs), with their training on vast datasets, present a promising solution to this challenge. Precision in clinical summarization is crucial, as it directly impacts patient care and safety. This study aimed to evaluate the effectiveness of parameter-efficient, fine-tuned, decoder-only LLMs for clinical note generation from doctor-patient conversations. We focus on assessing medical accuracy, robustness, and the feasibility of parameter-efficient fine-tuning (PEFT) approaches under practical resource constraints. We used the Medical Training Summarization Dialog dataset containing 1700 doctor-patient conversations paired with clinical notes. Several decoder-only LLMs, including Mistral, Meditron, and Llama, were fine-tuned using PEFT techniques to reduce computational and memory overhead. Evaluation was performed using standard automatic metrics, including the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers score, to assess content overlap and semantic similarity between generated and reference clinical notes. In addition, an expert physician assessed the LLM-generated notes for medical accuracy, completeness, concision, relevance, and clinical coherence and readability. Model performance was evaluated using the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers scores, demonstrating that Meditron-7B and Llama3-8B achieved state-of-the-art results among open-source, parameter-efficient, fine-tuned models, with Mistral-7B also performing competitively. The findings indicate that decoder-only LLMs, particularly Llama variants, outperform traditional models. Moreover, fine-tuning with higher quantization has the potential to further enhance performance. Human expert evaluation further indicated that Llama3-8B and Mistral-7B produced clinically coherent and accurate summaries, with Meditron-7B and Llama3-3B also performing reliably across evaluation criteria. The findings suggest that higher quantization during fine-tuning may improve efficiency without substantially compromising performance. This study underscores the potential of the PEFT of decoder-only LLMs to transform clinical workflows by streamlining medical documentation, thereby enabling health care professionals to dedicate more time to patient care. These models offer a scalable and resource-efficient alternative to traditional architectures and have the potential to streamline clinical documentation workflows.
Retrospective identification of acute nonarteritic anterior ischemic optic neuropathy (NAION) cases is critical for research on risk factors. However, reliance on International Classification of Diseases (ICD) 10th edition coding for case identification has limited accuracy, and manual review of longitudinal electronic health records is time-intensive. The purpose of this study is to evaluate automated methods for retrospective identification of acute NAION cases using large language models (LLMs) that preserve patient privacy. Retrospective cross-sectional study. 165 patients with ≥1 ICD-10 code for ischemic optic neuropathy (H47.01∗) in the electronic health record at an academic medical center. Five locally deployed LLM models (Mistral Small 3.1, Magistral Small, Gemma3, MedGemma, GPT-OSS 20B) were used to implement 4 approaches for acute NAION diagnostic classification using unstructured ophthalmology records (basic prompting, retrieval-augmented generation [RAG], 2-step agentic workflow, and 3-step agentic workflow). Ten percent of subjects were used for prompt refinement. Large language model/approach diagnostic classifications were compared against expert neuro-ophthalmologist diagnoses based on chart review. Positive predictive value (PPV) of LLM approaches for acute NAION case identification with expert chart review diagnosis serving as gold standard. Secondary outcomes included negative predictive value, sensitivity, specificity, accuracy, F1 score, and distribution of LLM/approach classifications. 7/17 prompt refinement subjects and 58/148 testing subjects had acute NAION by expert chart review corresponding to PPV of 0.39 for ≥1 ICD code. Large language model approaches accurately identified 20 ± 12 (mean, standard deviation) acute NAION cases in the test set with PPV of 0.78 ± 0.16 and accuracy of 0.69 ± 0.06. The Mistral model using a 3-step agentic approach had the best-balanced performance (39 cases identified, 0.85 PPV, 0.82 accuracy, 0.75 F1 score). Privacy-preserving agentic LLM approaches can achieve high PPV for acute NAION case identification using unstructured ophthalmology longitudinal electronic health records. These results exceed the performance of using structured ICD codes to identify cases, offering a scalable, efficient method for case identification in retrospective research while maintaining patient confidentiality and local data control. This method has application for enhancing research efficiency and accuracy for studies on NAION risk factors, with potential applicability to other conditions requiring complex diagnostic review. Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Emergency department (ED) overcrowding causes diagnostic challenges, prolonged wait times, and impairs appropriate triage, often due to human error and fatigue. Large language models can assist ED staff in triage, improving patient care by mitigating these problems. We designed an evaluation method (Skyer benchmark) to assess fifteen large language models, including DeepSeek-R1 (70B, 7B), ChatGPT versions (4, 4.5-preview), Gemini iterations (1.5-pro, 2.0-Pro-experimental, 2.5_03-25, 2.5_05-06), Mistral-7B, Llama-3.3, Gemma models (2-27b-it, 3-12b-it, 3-27b-it), Qwen-2.5, and Phi-4-14B. We assessed the performance of models in ED triage by using 55 realistic clinical pediatric scenarios. The main objective was to evaluate models' triage performance using a weighting system that accounted for the impacts of over-triage and under-triage, in addition to simple accuracy. A secondary objective was to assess models' consistency, by repeating tests across scenarios three times. Our findings indicate that only two models, ChatGPT-4.5-preview and Gemini-2.5_05-06, demonstrated superior and reliable triage performance. ChatGPT-4.5-preview (77% accuracy, mean weight 377.5 out of 550) and Gemini-2.5_05-06 (74% accuracy, mean weight 365/550) significantly outperformed human-triage-experts accuracy (64% accuracy, mean weight 253.5/550) and other models. This difference was statistically significant (p-value < 0.05), with an extremely large effect size (Cohen's D = 2.18 and 1.98). Furthermore, they demonstrated sufficient reliability due to acceptable consistency in their triage performance (85% and 82%). Our work establishes Skyer as an evaluation approach of large language models. Skyer selected the best-performing models, which demonstrated significantly higher triage performance than the human experts and showed consistent results. These large language models can streamline the delivery of quality healthcare in overcrowded EDs. Despite promising outcomes, we identified limitations prohibiting these large language models as replacement of human experts. Instead, we demonstrate their potential for a substantial role in assisting the staff in overcrowded EDs. RéSUMé: OBJECTIFS: Le surpeuplement des services d’urgence (SU) entraîne des difficultés diagnostiques, des temps d’attente prolongés et un triage inapproprié, souvent en raison d’erreurs humaines et de la fatigue. Les grands modèles de langage peuvent aider le personnel du service d’urgence dans le triage, améliorant ainsi les soins aux patients en atténuant ces problèmes. MéTHODES: Nous avons conçu une méthode d’évaluation (benchmark de Skyer) pour évaluer quinze grands modèles de langages, y compris DeepSeek-R1 (70B, 7B), les versions ChatGPT (4, 4.5-preview), les itérations Gemini (1.5-pro, 2.0-Pro-experimental, 2.5_03-25, 2.5_05-06), Mistral-7B, Llama-3.3, les modèles Gemma (2-27b-it, 3-12b-it, 3-27b-it), Qwen-2,5 et Phi-4-14B. Nous avons évalué la performance des modèles dans le triage ED en utilisant 55 scénarios pédiatriques cliniques réalistes. L’objectif principal était d’évaluer la performance de triage des modèles à l’aide d’un système de pondération qui tenait compte des impacts du triage excessif et du triage insuffisant, en plus de la simple précision. Un objectif secondaire était d’évaluer la cohérence des modèles, en répétant trois fois les tests entre les scénarios. RéSULTATS: Nos résultats indiquent que seuls deux modèles, ChatGPT-4.5-preview et Gemini-2.5_05-06, ont démontré des performances de triage supérieures et fiables. ChatGPT-4.5-preview (précision de 77%, poids moyen 377,5 sur 550) et Gemini-2.5_05-06 (précision de 74%, poids moyen 365/550) ont significativement surpassé la précision des experts en triage humain (précision de 64%, poids moyen 253,5/550) ainsi que d’autres modèles. Cette différence était statistiquement significative (valeur p > 0,05), avec une taille d’effet extrêmement grande (D de Cohen = 2,18 et 1,98). De plus, ils ont démontré une fiabilité suffisante en raison d’une cohérence acceptable dans leur performance de triage (85% et 82%). CONCLUSION: Notre travail établit Skyer comme une approche d’évaluation des grands modèles linguistiques. Skyer a sélectionné les modèles les plus performants, avec une performance de triage significativement supérieure à celle des experts humains. Ces grands modèles de langage peuvent rationaliser la prestation de soins de santé de qualité dans les services d’urgence surpeuplés. Malgré des résultats prometteurs, nous avons identifié des limites interdisant ces grands modèles de langage en remplacement des experts humains. Au lieu de cela, nous démontrons leur potentiel pour un rôle substantiel dans l’assistance au personnel dans les services d’urgence surpeuplés.
This study aimed to generate medical qualification exam questions and their corresponding answers from real-world electronic health records (EHRs) with large language models (LLMs), and to compare their output to that of human medical experts. Utilizing a multicenter bidirectional anonymized database China Elderly Comorbidity Medical Database (CECMed), a total of 8 LLMs: ERNIE 4, ChatGLM 4, Doubao, Hunyuan, Spark 4, Qwen, Llama 3, and Mistral were tasked with generating open-ended questions and answers based on a subset of sampled admission reports. LLMs generated the medical question and answer through few-shot prompting. An independent expert panel scored the AI-generated outputs based on multiple criteria, including coherence, sufficiency of key information, information correctness, factual consistency, evidence of statement, and professionalism, using 5-point Likert scales. For question generation, ERNIE 4 achieved the highest cumulative score (16.47). Human experts surpassed LLMs in sufficiency of key information (3.67) but lagged in information correctness (3.63 vs. LLMs' 4.03-4.57). The information correctness of ERNIE was significantly higher than the human's [0.93 (0.62, 1.24), p < 0.01]. For answer generation, humans led overall (14.49), while Doubao outperformed the other LLMs in coherence (3.57), factual consistency (3.60), and professionalism (3.53). The coherence of human's was significantly better than that of 8 LLMs, especially outperformed Llama [0.8 (0.37, 1.23), p < 0.01] and Mistral [0.87 (0.45, 1.28), p < 0.01]. Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming. This study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians. Although current LLMs performed dissatisfactorily in some aspects, medical students and interns may find LLMs a useful auxiliary tool to support their learning. https://clinicaltrials.gov/study/NCT06316544, identifier: NCT06316544.
This study develops an open-source large language model-based chatbot tailored for Korean health consultations. The chatbot was implemented using the retrieval-augmented generation (RAG) technique alongside metadata filtering to enhance its performance. This study aims to analyze and compare the performance of a RAG-based chatbot with other leading language models in the context of Korean health consultations. A 10.4 GB Korean medical document corpus (487,277 segments) was constructed from official websites of major Korean hospitals, public health sources, and medical textbooks. This study quantitatively compared 5 open-source large language models (Qwen3:4B, Mistral:7B, Llama-3.1:8B, Gpt-Oss:20B, and Gemma3:27B) in 3 configurations: baseline (model only), RAG-only, and RAG with metadata filtering. The RAG system used a specialized Korean embedding model (upskyy/bge-m3-korean) and an Elasticsearch store. Performance was assessed by an emergency medicine specialist using a validation set of 226 questions across 7 common diseases and scoring responses based on accuracy, safety, and helpfulness. The application of RAG alone failed to yield statistically significant performance improvements and, in some cases (Llama 3.1: 8B and Gemma 3: 27B), resulted in decreased scores. However, the combination of RAG with metadata filtering yielded statistically significant (P<.05) performance increases in most models. Notably, the average score for Mistral:7B increased from 3.79, SD 0.08, to 4.10, SD 0.10, and Gpt-Oss:20B increased from 4.43, SD 0.05, to 4.51, SD 0.04, with the latter achieving the highest safety score (4.61, SD 0.03). The Gemma3:27B model, which possessed a high baseline performance (4.42, SD 0.03), was an exception, exhibiting no significant improvement (P=.14) even with filtering. The effectiveness of RAG for specialized domains such as Korean medical consultation is highly dependent on a metadata filtering process that controls the quality of retrieved information; simple information augmentation is insufficient. Furthermore, the benefit of RAG is limited when a model's intrinsic knowledge (eg, Gemma3:27B) already meets or exceeds the quality of the external knowledge base. This finding indicates that performance enhancement strategies must account for both the retrieval mechanism's quality and the model's preexisting capabilities.
Paper-based case report forms remain common in clinical trials, often requiring manual transcription into electronic systems. Automating this process could reduce workload and errors, but standard approaches for text or mark recognition are limited by rigid templates and poor generalizability. Recent advances in visual language models (VLMs) offer a potential alternative by jointly processing images and text. In this study, we benchmarked three open-source, locally executable VLMs (Qwen 2.5, Mistral Small 3.1, and Granite 3.2 Vision) for information extraction from printed case report forms used in an Italian stroke trial. Tasks included identifying the form title and handwritten Record ID, extracting handwritten dates, and detecting marked checkboxes. Experiments were conducted on 80 smartphone-acquired images, simulating real-world data entry. Results showed that Qwen performed best in title recognition (91%) and date extraction (75%), while Mistral achieved higher accuracy for Record IDs (53%) and checkboxes (80%). Granite, despite being the fastest, often failed to follow output formatting instructions. Across models, handwritten fields remained particularly challenging. These findings highlight both the promise and current limitations of open VLMs for streamlining clinical research workflows.
Embedding systematic, structured data extraction within electronic health records (EHR) is vital for improved real-time insights into care delivery. This study evaluates the feasibility of using large language models (LLMs) to extract structured advance care planning (ACP) information from unstructured Goals of Care (GoC) clinical notes in the EHR. A sample of 100 de-identified GoC notes was manually annotated by clinicians across four ACP categories: Patient Priorities, Code Status, Decision Maker, and Documentation. Two LLMs (Mistral 24.07 and LLaMA 3.1) were prompted to extract structured outputs without domain-specific fine-tuning. Model outputs were compared to human annotations using cosine similarity of BioBERT embeddings. Mistral 24.07 achieved high semantic similarity in Code Status (0.814), Documentation (0.781), and Patient Priorities (0.770), but lower alignment in Decision Maker (0.609). LLMs can effectively extract structured ACP information, particularly in well-documented categories, suggesting potential for scalable, data-driven feedback loops that improve the provision of care. However, accuracy challenges remain, and further refinement is needed for nuanced qualitative content categories.
People who inject drugs (PWID) face a high risk for serious infections, yet International Classification of Diseases (ICD) codes fail to identify this population. Large language models (LLM) offer a promising alternative by extracting information from unstructured clinical text. This study evaluated the diagnostic performance of off-the-shelf LLMs in identifying PWID and related attributes from hospital discharge summaries. In this cross-sectional study, discharge summaries from the Infectious Diseases service at St Vincent's Public Hospital, Sydney, between 2018 and 2022 were reviewed. A single reviewer manually annotated each de-identified summary for PWID status, drugs reported, injection recency and opioid agonist therapy. Eight LLMs (Gemma3, Llama 3.3, Mistral, Phi4, hippomistral, llama3-med [8B and 70B] and OpenBioLLM) were compared using prevalence-weighted average-F1 scores. Diagnostic metrics with bootstrapped 95% confidence intervals were calculated for each annotated category. Of 859 first admissions, manual review identified 149 (17.1%) PWID. ICD codes showed low sensitivity (≤ 0.32) but high specificity (≥ 0.97) for identifying PWID. The best-performing model (Llama 3.3) achieved a prevalence-weighted average-F1 of 0.845 (0.733, 0.927). For injecting drug use, sensitivity was 0.819 (95% CI 0.753, 0.879) and specificity 0.999 (0.996, 1.00). Identification of heroin, methamphetamine, cannabis and methadone was near perfect (F1 > 0.973), while illicit prescription opioid and benzodiazepine use were identified less accurately (F1 = 0.400 and 0.606). LLMs accurately identify PWID from discharge summaries, outperforming ICD codes. Challenges remain for certain substances, underscoring the need for task-specific tuning, external validation and integration with structured data to enhance surveillance and interventions.
Large language models (LLMs) have recently been integrated into dental practice to support clinical reasoning and preventive decision-making. This study compared the performance of five advanced chatbots-ChatGPT-5, Claude 4.5 Sonnet, Gemini 2.5 Pro, LLaMA 3.1, and Mistral 7B-in providing evidence-based responses for caries risk assessment and preventive management in pediatric cases. Twenty-five validated, case-based questions were developed in accordance with internationally recognized pediatric and preventive dentistry guidelines. Responses were evaluated by six pediatric dentistry experts for accuracy, completeness, relevance, clarity, and usefulness using Likert-type scales. Response time, word count, and linguistic readability characteristics (Flesch Reading Ease Score and Flesch-Kincaid Grade Level) were additionally analyzed to compare textual complexity across chatbot-generated responses. Data normality was assessed using the Shapiro-Wilk test; parametric tests (ANOVA with Bonferroni correction) or non-parametric tests (Kruskal-Wallis with Dunn's post hoc) were applied as appropriate. Statistically significant differences were observed across all qualitative criteria, including accuracy, completeness, relevance, clarity, and usefulness (p < 0.001). ChatGPT-5 consistently ranked among the top-performing models, showing balanced and high-quality responses across domains, while Claude 4.5 Sonnet achieved the highest accuracy and completeness scores. Gemini 2.5 Pro produced the fastest responses (p < 0.001), whereas Claude 4.5 Sonnet generated the longest and most linguistically complex outputs. Readability metrics also differed significantly among models (p < 0.001), with Mistral 7B and LLaMA 3.1 showing the highest readability. All evaluated chatbots generated generally relevant responses for caries risk assessment and preventive counseling; however, substantial inter-model differences were observed in qualitative performance, linguistic complexity, and response characteristics. Occasional inconsistencies and outdated content highlight the need for cautious interpretation and further externally validated evaluation before broader clinical implementation.
Analyzing semi-spontaneous speech is a promising direction for supporting Alzheimer's disease (AD) assessment, yet progress is limited by the scarcity of annotated clinical data. Large Language Models (LLMs) offer new opportunities to generate synthetic narratives that may resemble speech patterns of both patients with AD and healthy controls during cognitive evaluation tasks such as the Cookie Theft Picture description. This study evaluates whether models including GPT, T5/Flan-T5, LLaMA, Mistral, and Qwen can generate clinically plausible picture-description narratives under two configurations: Human-to-Bot, where an LLM responds directly to real interviewer prompts, and Bot-to-Bot, where two LLMs simulate both interviewer and participant roles. Models were fine-tuned on transcripts from the DementiaBank Pitt Corpus and assessed using lexical and semantic metrics, as well as human expert ratings. Generated narratives were further used to augment training data for an AD vs. healthy control classifier based on BERT embeddings and an MLP architecture. LLMs differed substantially in their ability to reproduce clinically meaningful and semantically coherent narratives of patient-interviewer interactions. Mistral, LLaMA, and Qwen achieved the strongest automatic evaluation metrics, e.g., BERTScores above 0.90 in the Human-to-Bot condition-and produced narratives rated by human experts as fluent, plausible, and diagnostically informative. When combining real and synthetic narratives for classifier training, the highest F1-score reached 0.84, outperforming models trained on real data alone (F1 = 0.74). Synthetic data generated in Human-to-Bot settings contributed most to diagnostic improvements, whereas Bot-to-Bot interactions exhibited greater variability and reduced clinical realism. LLMs can generate high-quality synthetic narratives that enhance downstream AD classification and show promising clinical plausibility in cognitive assessment contexts. Incorporating LLM-generated data provides a scalable strategy for mitigating data scarcity in dementia research. Future work should focus on improving fully synthetic dialogue quality, expanding multilingual capabilities, and refining evaluation frameworks to better capture clinically relevant linguistic features.
Clinical retrieval-augmented generation depends on embedding models. A companion study found that non-retrieval-trained encoders underperformed retrieval-trained general-purpose embeddings and produced near-degenerate embedding geometry, but did not localize the architectural origin, separate training-domain from training-objective effects, or test whether the degradation can be corrected without retraining. This study aimed to (1) characterize layer-wise retrieval and geometric trajectories across 13 transformer configurations on clinical documents, (2) separate training-objective from training-domain effects through matched architectural comparisons, (3) reanalyze the panel under a per-query layer-wise linear mixed-effects (LME) framework, and (4) evaluate deployment-relevant post hoc geometric correction. Layer-wise embeddings were extracted from 13 transformer configurations on 3 clinical corpora (n=100 documents each: MTSamples, PMC-Patients, and Mistral-7B-Instruct-generated synthetic notes) under 2 query formats (keyword and natural-language via GPT-4o). Retrieval performance (mean reciprocal rank at cutoff 10 [MRR@10] and recall at cutoff 10 [recall@10]) and geometric properties (participation ratio, average pairwise cosine, and anisotropy) were measured at every layer. A per-query layer-wise LME model was fit independently per configuration. Corpus-only zero-phase component analysis (ZCA) whitening was evaluated as the primary deployment-relevant intervention, with 5-fold cross-validation, an epsilon sweep, a lexical-overlap audit, and a chunking sensitivity analysis. Document embeddings clustered into 3 anisotropy tiers: extreme (average pairwise cosine >0.92) for non-retrieval-trained encoders and large language models (LLMs), moderate (0.65-0.92) for general retrievers and most LLMs, and reduced (<0.65) for BioLORD-2023, instruction-tuned E5-Mistral-7B, and Nomic-embed-text-nopfx. The per-query random-slope LME identified 2 layer-depth patterns: classical degradation with depth in 3 non-retrieval-trained encoders (all P<.001), vs net improvement with depth in the remaining 10 models (all P<.001), with the strongest negative coefficients in decoder LLMs. Matched-contrast tests confirmed significant training-objective × layer-depth interactions in all 3 matched pairs (all P<.001). Corpus-only ZCA whitening produced a 2-tier pattern under 5-fold cross-validation: tier 2 non-retrieval-trained models showed positive ΔMRR@10 (+0.066 to +0.304), while tier 1 retrieval-trained models showed negative ΔMRR@10 (-0.021 to -0.051). Best Match 25-vs-embedding Spearman rank correlations spanned -0.02 to 0.37, indicating substantial nonlexical contribution to retrieval. Same-source ranking stability replicated at 4-5× corpus scale (ρ=0.952 for PMC-500, ρ=0.929 for MTSamples-400) for the BERT-scale subset. Anisotropy in transformer embeddings is widespread across architectural classes and is lower in configurations with retrieval-specific training. Corpus-only ZCA whitening is a deployment-compatible, retraining-free post hoc correction candidate that improved retrieval for non-retrieval-trained models on this controlled benchmark but requires target-corpus validation before clinical deployment. The matched-comparison evidence supports training objective rather than training domain as the stronger explanatory axis, though residual confounding is not eliminated. The principal contribution is mechanistic: layer-level localization of embedding degradation and the geometric basis for the 2-tier intervention response.
Pediatric asthma exacerbations are a common emergent condition treated by prehospital emergency medical services (EMS). However, retrospective identification of those patients for research, educational, and quality improvement purposes can be difficult given the unique structure of EMS' electronic health records. Computable phenotypes are algorithms used to identify patients of interest. The objective of this study was to explore the performance of large language models (LLMs) as pediatric asthma computable phenotypes for EMS data. This is a retrospective, observational study testing the performance of state-of-the-art, open-source, general-purpose LLMs (Gemma-2, Llama-3.1-8B, Llama-3.3-70B, and Mistral-0.3), and one LLM specifically designed for medical use (OpenBioLLM), as pediatric asthma exacerbation computable phenotypes for EMS data. The goal of the phenotype was a binary classification of yes or no for an EMS encounter for a pediatric asthma exacerbation. We examined EMS patient encounter data from the ESO Data Collaborative between January 1, 2018 and December 31, 2021 for patients ages 2-18 years. Two pediatric emergency medicine physicians independently reviewed and annotated 1,000 encounters to label patients as pediatric asthma exacerbations. We tested models on structured and/or unstructured data, and explored basic and chain-of-thought prompts. We measured model performance using specificity, sensitivity, positive predictive value, negative predictive value, and macro F1. After applying the inclusion-exclusion criteria, 24,283 patient encounters remained. The median age was 12 years, with a slight majority (51%) of female patients. The best performing LLM overall was Llama 3.3's 70 billion parameter model using unstructured and structured data with 10-shot chain-of-thought prompts, with a F1 score of 0.894. Unstructured data alone gave the best F1 score for all models except Llama-3.3-70B and Mistral-0.3. Chain-of-thought prompts were more likely to produce better results than basic prompts, with 4 out of the 5 models giving their best performance when prompted by a chain-of-thought prompt. In this work we demonstrate that open source LLMs can be tailored for use as accurate prehospital pediatric asthma computable phenotypes using EMS data. To our knowledge, this represents the first LLM-based computable phenotypes trained on EMS data.
Background: Reliable structured-output generation is a prerequisite for using large language models (LLMs) in automated clinical documentation workflows, but many evaluations focus on clinical quality before testing whether outputs are parseable, schema-compliant, and stable. Methods: We evaluated three locally deployed open-weight LLMs in the 7- to 8-billion-parameter range (Llama3-Med42-8B, Meta-Llama-3-8B-Instruct, and Mistral-7B-Instruct-v0.3) for structured admission-note editing. Seventy de-identified English-language admission notes (35 internal medicine and 35 surgical) were processed by each model in three independent runs under two output-control conditions: a free-text JSON prompt and a schema-enforced structured-output condition. A total of 1260 local inferences were performed in LM Studio on consumer-grade hardware. Automated proxy metrics assessed JSON/schema validity, run-to-run stability, instruction compliance, verbosity, numeric-token preservation, and uncertainty-marker change without clinician adjudication of clinical correctness. Results: Under the free-text JSON prompt, the tested Mistral-7B-Instruct-v0.3/embedded-prompt configuration had the weakest structural reliability (74.3-78.6% first-pass validity per run; 18.6-21.4% persistent parse/schema failures after retry), with at least one final failure for 17 of 70 notes. In a message-format sensitivity analysis using Meta-Llama-3-8B-Instruct, embedding system instructions in the user message increased first-attempt invalid outputs compared with separate system/user roles (55/700, 7.9% vs. 12/700, 1.7%). Under schema enforcement, all models produced 70 of 70 first-pass valid, schema-compliant outputs in every run. Documentation behavior nevertheless differed by model, including differences in verbosity and numeric-token preservation. Conclusions: Schema enforcement removed parsing failures in this sample but did not eliminate model-specific editing behavior. Proxy-based screening can identify structurally unstable model-prompt or model-format configurations before clinician review.
Colorectal cancer is a leading cause of cancer-related deaths in the United States, and colonoscopy remains the gold standard for early detection and prevention. However, many procedures are postponed due to inadequate bowel preparation, a preventable failure often caused by patients' difficulty in understanding and following written prep instructions. Prior interventions such as reminder apps and instructional videos have improved adherence only modestly, largely because they cannot answer patient-specific questions. Recent advances in large language models (LLMs) raise the possibility of developing conversational assistants that can provide interactive support to patients in procedure preparation. This study evaluated the correctness, harmfulness, and diversity of synthetic dialogues generated by leading LLMs acting as both simulated AI Coaches and patients for colonoscopy preparation. Five leading LLMs-OpenAI's o3, GPT-4.1, and GPT-5.1; Meta's Llama 3.3 70B; and Mistral's Large-2411-were used to generate 250 patient-AI Coach dialogues per model. Dialogues consisted of 3 to 7 question-answer pairs concerning diet, medications, and other prep-related topics. A multiprompt, multiquestion approach was designed to elicit diverse patient questions, and an error taxonomy was established to assess model capabilities in responding to questions. Human raters, including 3 medical experts, evaluated the generated questions for difficulty and the responses for correctness, error type, and potential harmfulness. Automatic evaluation using an LLM-as-a-judge approach complemented human evaluation. Question diversity was assessed using lexical diversity metrics (Distinct-1 and Distinct-2) and entropy. In addition, we evaluated a safety filtering mechanism in which responses judged incorrect by an automated evaluator were replaced with a deferral message instructing patients to contact their health care provider. Differences in response correctness across models were evaluated using permutation tests conducted at the dialogue level. Interrater agreement among human evaluators was assessed using the Gwet AC1 statistic. The study was conducted between May and September 2025. Automatic evaluation results closely aligned with human judgments: leading models approached but did not achieve adequate performance. Closed-weight models (GPT-5.1, GPT-4.1, and o3) outperformed open-weight models (Llama and Mistral) on correctness, with the reasoning models (GPT-5.1 and o3) performing best. This turn-level ranking was preserved under the supplementary single-prompt baseline, although dialogue-level rankings differed. All models produced harmful errors, primarily due to omissions or misinterpretations of prep instructions. The multiprompt generation strategy substantially increased the diversity of patient questions compared with a single-prompt baseline. Applying an automated safety filter reduced overall error rates but failed to eliminate harmful responses. Although LLMs demonstrate strong potential to support colonoscopy preparation, none are yet reliable enough for unsupervised deployment in patient-facing contexts. Persistent harmful errors and the limited effectiveness of simple filtering mechanisms highlight the need for improved instruction adherence, stronger safety mechanisms, and validation using real patient queries.