Mars rover missions require decision-support systems that can interpret terrain, telemetry, environmental conditions, and mission objectives under delayed communication with Earth. This study evaluates whether multi-agent orchestration improves simulated Mars rover decision support compared with a single-agent baseline. A controlled benchmark of 100 synthetic mission-inspired rover scenarios was evaluated using OpenAI GPT-4o and GPT-5.5, with five repeated runs per scenario and architecture. Model-facing scenario inputs were separated from evaluator-side labels so that expected actions and hazards were reserved for scoring only. Performance was measured using decision accuracy, exact and substring-based semantic hazard F1, hazard error counts, latency, token usage, scenario-level paired statistical comparisons, and GPT-4o specialist-agent ablations. Across the tested OpenAI configurations, the single-agent architecture showed numerical advantages in decision accuracy and hazard-label alignment, but these decision-quality differences were not consistently significant under scenario-level statistical analysis with Holm-Bonferroni adjustment. The only decision-quality metric remaining significant was GPT-5.5 exact hazard F1, although absolute values were very low. The most reliable difference was computational efficiency: the single-agent architecture required substantially lower latency and token usage than the prompt-defined multi-agent orchestration architecture. Multi-agent orchestration generated broader hazard lists, including plausible non-canonical observations, but did not reliably improve aggregate decision accuracy or hazard F1. These findings suggest that, for short-context, tool-less, static decision-support tasks where all relevant context is available in a single input, multi-agent orchestration should be treated as a cost-bearing design choice rather than an assumed improvement. The study contributes a reproducible architecture-level benchmark for evaluating when LLM-based orchestration is worth its operational cost in mission-inspired workflows.
(1) Background: Clinical trial data extraction from registries such as ClinicalTrials.gov remains labor-intensive and error-prone, often missing critical details hidden in unstructured protocol descriptions. Large Language Models (LLMs) offer potential to automate this process, yet systematic multi-model comparisons on real clinical trial data remain scarce. (2) Methods: Four LLMs (OpenAI o4-mini-high, Anthropic Claude-Sonnet-4, Google Gemini 2.5-Pro, and Meta Llama-4-Maverick) extracted brain stimulation parameters from 67 transcranial direct current stimulation (tDCS) trials in Parkinson's disease via a structured JSON schema. Pairwise inter-model agreement was quantified with Cohen's Kappa and percentage agreement across binary, categorical, and multi-component task tiers. (3) Results: Under exact-string matching, agreement was near-perfect for binary classifications (non-invasive classification: 100%; brain stimulation presence: 99.3%, κ = 0.50) and substantial for categorical extractions (primary stimulation type: 96.4%, κ = 0.70), but fell to 48.6% (κ = 0.43) for complex anatomical targets. Numeric parameters revealed model-specific strengths: o4-mini-high and Claude-Sonnet-4 achieved perfect duration agreement (r = 1.000, n = 19) while Llama-4-Maverick diverged substantially (r < 0.12). Validation against an expert gold standard (100% inter-annotator agreement on a 20-trial overlap) confirmed high extraction accuracy across all features (mean 93.7-98.9%). Crucially, the low agreement on anatomical targets proved to be an artifact of exact-string scoring: under the same semantic matching used to measure accuracy, inter-model agreement rose to 97.0%, coinciding with the 95.5% expert accuracy. Inter-model agreement therefore tracks accuracy once both are measured on a common basis. (4) Conclusions: Exact-string inter-model agreement decreases with task complexity, but this decline largely reflects interchangeable free-text wording rather than reduced accuracy. Evaluated semantically, agreement and expert accuracy are both high and closely aligned. A residual risk is not low accuracy but the rare error shared across all models, which agreement cannot detect, and which overall accuracy can itself mask when one class dominates. These findings inform hybrid human-AI systematic review pipelines in which targeted expert oversight focuses on shared-error and minority-class detection.
To systematically evaluate the capability of the Large Language Model, Generative Pre-Trained Transformer (GPT-4o), to deliver Motivational Interviewing. Large Language Models are a type of Artificial Intelligence trained to understand and generate human language. OpenAI's GPT-4o was prompted to conduct Motivational Interviews with simulated patient actors. The interview transcripts were analysed using the Motivational Interviewing Treatment Integrity code, which includes thresholds for assessing competency in Motivational Interviewing. Conversation sequences were furthermore examined to understand GPT-4o's patterned responses across Motivational Interviewing behaviours. A total of 36 Motivational Interviewing transcripts were generated covering a range of chronic conditions and health behaviours. GPT-4o performed above 'good' competency thresholds for behaviour counts (Percent Complex Reflection and Reflection to Question ratio). Total Motivational Interviewing-Adherent behaviour counts were very high with very few Total Non-Motivational Interviewing-Adherent counts. GPT-4o did not perform as well on global competency scores, performing below 'fair' competency thresholds for both relational and technical skills. A distinct pattern was observed across the interviews where GPT-4o gradually shifted away from Motivational Interviewing towards attempts to persuade participants to change, which is inconsistent with the principles and processes of Motivational Interviewing. GPT-4o has some capacity to mimic Motivational Interviewing behaviours but requires further augmentation to perform it with proficiency. Using Artificial Intelligence has the potential to increase accessibility and reduce the costs of delivering Motivational Interviewing. Further research is needed to explore effective augmentation techniques before Large Language Models such as GPT-4o can be used to deliver Motivational Interviewing.
As large language models (LLMs) are increasingly used to interpret medical concerns, rigorous evaluation of their performance on clinically relevant tasks is essential. However, the new state-of-the-art models from OpenAI and Google as of February 2026, have not been evaluated for their accuracy and consistency in interpreting ECGs independent of clinical context. We aim to compare ChatGPT (GPT-5.2 Thinking) and Gemini (Gemini 3 Pro) on electrical axis and heart rhythm identification to assess current clinical usability and identify areas for improvement. ECGs were obtained from the Lobachevsky University Electrocardiography Databases on PhysioNet. The LLM responses were evaluated for first-shot accuracy and consistency across three different trials. First-shot accuracy was further split into the various types of rhythms and axes to examine systematic trends in model outputs. Both models demonstrated comparable overall first-shot accuracies for electric axis and rhythm classification. Both models struggled with less common rhythm categories, including multifocal rhythms and tachycardias. The Macro F1 analysis indicated low overall classification performance for both ChatGPT and Gemini in terms of axis and rhythm. ChatGPT achieved a higher Macro F1 point estimate for axis classification compared with Gemini, though both models struggled with a Macro F1 of 45.2% and 32.4% respectively. The Macro F1 scores of both ChatGPT and Gemini suggest that they are not reliable for independent clinical diagnoses in cardiology, with difficulty shown in interpreting ECGs for rhythm and axis. This investigation aims to guide continued improvement of these LLM models for physician assistance.
Explore the perspectives of primary caregivers towards pediatric tissue-based research participation. Cross-sectional. Two academic pediatric gastroenterology sites in the United States, including one serving a largely rural referral population and one urban clinic population. Primary caregivers of children who underwent endoscopy between 2017-2018 at UVA or were seen in the clinic setting between 2024-2025 at Tulane and referred by their child's gastroenterologist to complete an electronic survey. Primary caregiver attitudes, motivations, and concerns toward pediatric tissue-based research were explored using descriptive-focused coding in NVivo and a large language model (LLM) processing pipeline based on OpenAI's GPT-4 for thematic, emotional, and sentiment analyses. Data were analyzed from 92 primary caregivers. Overall, respondents were amenable to having their children provide specimens for research. Primary motivations included a desire to help others or advance science, and perceived medical benefits for their child so long as specimen collection did not cause additional distress. Discomfort with participation was often linked to prior traumatic clinical experiences, concerns about additional biopsies causing unnecessary discomfort, or privacy issues. A desire to help others and potentially their own child was the strongest motivator for participation, while scheduling constraints and perceived risks to the child's health were the main barriers. At both sites, primary caregivers expressed strong willingness to participate in pediatric research. Primary concerns included perceived invasiveness of biospecimen collection and potential for additional discomfort. Limitations of the study included the unstructured nature of the data making the analysis and interpretation challenging. Strengths included two demographically diverse sites, intentional enrollment of primary caregivers of children both with and without invasive diagnostic testing, and use of LLM based analyses.
Large language models (LLM) are being rapidly integrated into healthcare, particularly to streamline time- and labor-intensive administrative processes. However, the potential for artificial intelligence (AI) systems to demonstrate bias when employed for insurance authorization remains poorly understood. As insurers increasingly adopt AI to make coverage decisions, this study examined bias in LLM-driven prior authorization in otolaryngology, using oral cavity squamous cell carcinoma (SCC) as a case study. Using OpenAI's generative transformer GPT-4o, this study assessed for LLM bias when simulating insurance coverage decisions for head and neck cancer reconstruction. A standardized clinical scenario was constructed involving patients with T2N2 oral cavity SCC, all requiring surgical resection and reconstruction. The LLM was prompted to choose between a radial forearm free flap (RFFF) and split-thickness skin graft (STSG) for reconstruction, across 19,900 simulations. Patient profiles were systematically varied by age, sex, race/ethnicity, zip code-based income level, socioeconomic status (SES), and substance use history. The LLM output showed significant disparities in approval decisions. RFFF was more frequently approved for younger, Asian, or white patients from high-income zip codes or high SES backgrounds (p < 0.0001). Older, Black, and Hispanic patients, and those from lower-income areas or with substance use histories, were less likely to receive RFFF authorization (p < 0.0001). On sensitivity analysis, inclusion of tumor-specific information markedly skewed recommendations towards RFFF across sociodemographic backgrounds. In this experimental study, the LLM's outputs exhibited significant disparities for oral cavity cancer reconstruction based on patient demographic variables in the setting of limited clinical information. Inputs to LLMs for clinical decision-making should include pertinent and detailed information to reduce the risk of bias. As insurers increasingly integrate AI for prior authorization, recognition of its biases, rigorous safeguards, and increased regulatory governance are needed to promote equitable health care.
Critical thinking (CT) is essential in orthopaedic nursing practice, where nurses must accurately assess fractures, identify complications, plan mobility strategies, and support postoperative recovery. Concept mapping (CMAP) can facilitate structured clinical reasoning, while artificial intelligence (AI) may further support conceptual organization and reflective learning processes. However, evidence regarding AI-assisted CMAP in orthopaedic nursing education remains limited. To evaluate the effectiveness of an AI-assisted CMAP program in improving CT among undergraduate students enrolled in an orthopaedic nursing course. A two-group quasi-experimental pretest-posttest design was employed. Participants were allocated to either an experimental group receiving AI-assisted CMAP using ChatGPT (OpenAI) during a structured workshop or a control group receiving standard instructional activities. CT was evaluated using a validated 12-item CMAP-based assessment tool administered before and after the intervention. Data were analyzed using mixed-design ANOVA and ANCOVA. CT scores improved significantly over time, with greater improvement observed in the experimental group. Mean CT scores in the experimental group increased from 2.40 ± 0.50 at pretest to 2.97 ± 0.72 at posttest, whereas scores in the control group increased from 2.43 ± 0.68 to 2.57 ± 0.63. A significant interaction effect between time and group was identified (F(1,58) = 9.705, p = .003, partial η2 = 0.143). Students in the experimental group also reported high satisfaction with the AI-assisted CMAP program. AI-assisted CMAP may support CT development and structured clinical reasoning in orthopaedic nursing education. Integrating AI-assisted reflective support into concept-mapping activities may provide an effective pedagogical approach for helping nursing students organize and integrate complex clinical information.
Background: Family caregivers of children with autism spectrum disorder (ASD) increasingly utilize large language models (LLMs) for health information. This study presents a systematic comparative evaluation of three widely used LLMs as localized ASD health information tools in Saudi Arabia. Methods: Twenty-four clinically validated, caregiver-oriented questions were posed to Google Gemini 1.5 Pro, OpenAI ChatGPT (GPT-4o), and DeepSeek-V3 using a standardized prompt. Three expert raters independently evaluated responses across four dimensions: scientific accuracy, PEMAT-P understandability, PEMAT-P actionability, and neurodiversity (ND)-affirming language. Readability was assessed via Flesch-Kincaid Grade Level (FKGL) and SMOG indices. Non-parametric Kruskal-Wallis tests with post hoc Mann-Whitney U comparisons and one-sample t-tests were applied. Results: Gemini achieved the highest mean accuracy (2.96/3.00), significantly outperforming DeepSeek (p = 0.003, r = -0.37). Accuracy failures across all LLMs clustered on regional epidemiological, genetic risk, and financial inquiries. ChatGPT achieved significantly higher understandability than Gemini (p < 0.001, r = 0.55), while DeepSeek achieved significantly higher actionability than Gemini (p < 0.001, r = 0.59). However, all three LLM scores fell short of the Agency for Healthcare Research and Quality (AHRQ) 80% actionability benchmark (all p < 0.001). All LLMs exceeded patient education readability benchmarks (FKGL ≤ 6, SMOG ≤ 8; all p < 0.001); ChatGPT was the most readable (FKGL = 7.86; SMOG = 9.97) and Gemini the most complex. No model differed significantly on ND-affirming language, defaulting to a mixed medical-affirming register. Conclusions: Evaluated LLMs demonstrated distinct, specialized strengths: Gemini was the most accurate, ChatGPT the most readable, and DeepSeek the most actionable. Importantly, all models failed to meet established consumer education standards for readability and actionability. LLMs require extensive plain-language adaptation and cultural customization. Clinicians must guide families on navigating LLM outputs, particularly concerning country-specific epidemiological, economic, and healthcare service queries.
Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions. We developed an automated pipeline using OpenAI generative pre-trained transformer (GPT)-4o and Anthropic Claude Sonnet-3.5 to generate neurosurgical board-style questions from Neurosurgery Publications articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian. SANS questions were more often perceived as human-made than GPT-generated (residents, P = .0002; attendings, P = .1091) and Claude-generated (residents, P = .0002; attendings, P = .0272) questions. Notably, 54% of AI-generated questions misled at least one evaluator, and 23% misled both. In quality assessments, SANS questions outperformed GPT-generated (residents, P < 10-5; attendings, P = .0001) and Claude-generated (residents and attendings, P < 10-5) questions. Particularly, 25% of AI-generated questions were rated as suitable for board examinations vs 72% of human-generated questions when measured by evaluator consensus (P < 10-7). Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.
Background Large language models (LLMs) are increasingly being explored for medical applications, including clinical decision support and oncology education. However, their performance in radiation oncology remains insufficiently characterized. Methods This study evaluated and compared the performance of ChatGPT (GPT-5.2; OpenAI, San Francisco, USA) and Gemini (Ultra/Pro; Google, Mountain View, USA) in radiation oncology. A benchmark consisting of 70 multiple-choice questions covering clinical oncology, radiation physics, and radiobiology was used to assess general knowledge. In addition, 25 clinically relevant open-ended questions were independently evaluated by three radiation oncologists using five-point Likert scales for correctness and usefulness. A mixed-effects model was applied to analyze performance. Results ChatGPT achieved an accuracy of 94.3%, while Gemini achieved 97.1% in the multiple-choice assessment. For open-ended questions, both models received similarly high ratings, with mean correctness scores of 4.71 and 4.67 and mean usefulness scores of 4.63 and 4.64 for ChatGPT and Gemini, respectively. Mixed-effects analysis demonstrated a significant effect of question type on both correctness and usefulness, whereas no significant differences between models were observed. Although most responses were rated as good or very good, limitations became apparent in more complex clinical scenarios requiring prioritization and individualized decision-making. Minor discrepancies between benchmark answers and current clinical evidence were also identified. Conclusions Both ChatGPT and Gemini demonstrated high performance on benchmark-based radiation oncology assessments. While the generated responses were generally accurate, limitations remained in complex clinical scenarios requiring nuanced clinical judgment. Further studies are needed to determine whether such performance translates into meaningful clinical utility in real-world radiation oncology practice.
Accurate, consistent and comprehensive metadata are essential for the reuse of functional genomics data deposited in repositories such as the Gene Expression Omnibus (GEO), however, achieving this often requires careful manual curation, which is time-consuming, costly and prone to errors. In this paper, we evaluate the performance of Large Language Models (LLMs), focusing on OpenAI's GPT-4o, as an assistive tool for entity-to-ontology annotation of two commonly encountered descriptors in transcriptomic experiments, mouse strains and cell lines. Using over 9 000 manually curated experiments from the Gemma database and over 5 000 associated journal articles, we assess the model's ability to identify relevant free-text entries and map them to appropriate ontology terms. Using zero-shot prompting and retrieval-augmented generation (RAG) to incorporate domain-specific ontology knowledge, GPT-4o correctly annotated 77% of mouse strain and 59% of cell line experiments, and uncovered manual curation errors in Gemma for over 200 experiments (2% of total). GPT-4o substantially outperformed non-LLM alternatives, and was statistically indistinguishable from the highest-performing 2026 frontier models. Model errors often arose from typographical mistakes or inconsistent naming in the GEO record or publication, and resembled those made by human curators. Along with annotations, our approach requests that the model output supporting context and verbatim quotes from the sources. These were typically accurate and enabled rapid curator verification. We further found that for the difficult cell line task, an ensemble of LLMs can boost precision at the cost of recall. These findings suggest that while LLMs are not ready to fully replace manual curators, they can effectively support them. A human-in-the-loop workflow, in which LLM's annotations are provided to human curators for validation, should improve the efficiency and quality of large-scale biomedical metadata curation.
To evaluate the readability, quality, and misinformation of patient education materials generated by large language models, including ChatGPT-4o (OpenAI), Gemini 1.5 Pro (Google), and Copilot Pro (Microsoft), compared with American Society of Retina Specialists (ASRS) brochures for retinal diseases. A cross-sectional comparative analysis was performed by generating patient education materials on 3 retinal conditions: retinal detachment, diabetic retinopathy, and age-related macular degeneration. Materials were created using a general prompt (prompt A) and a prompt specifying a sixth-grade readability level (prompt B). Readability was evaluated using 6 validated metrics. Quality was assessed through DISCERN and the Patient Education Materials Assessment Tool. Misinformation was graded using a 5-point Likert scale. Assessments were performed independently by 2 masked retina specialists. Average readability of Gemini (11.65; P = .005) and Copilot (11.23; P = .003) materials was significantly better than that of ASRS materials (14.17), whereas ChatGPT showed no significant difference (12.85; P = .06). ChatGPT's average readability was significantly lower compared with Gemini (12.85 vs 11.65; P = .01) and Copilot (12.85 vs 11.23; P < .001). Prompt B significantly improved readability across all large language models relative to ASRS but still exceeded the sixth-grade readability level. DISCERN scores were comparable across groups. ASRS materials had an understandability score of 74.37%, which was significantly lower than ChatGPT (94.44%; P = .02) and Gemini 1.5 (95.83%, P = .02) scores. No significant differences were observed for actionability or misinformation. Readability showed no significant correlation with quality or misinformation (P > .05). Large language models, when appropriately prompted, can generate retina-related patient education material with superior readability compared with existing ASRS brochures, while maintaining comparable quality and accuracy. Large language models represent a promising approach for addressing literacy barriers, though expert oversight remains essential.
The rapid evolution of large language models (LLMs) in natural language processing has substantially elevated their semantic understanding and logical reasoning capabilities. Such proficiencies have been leveraged in autonomous driving systems, contributing to significant improvements in system performance. Models, such as OpenAI o1 and DeepSeek-R1, leverage chain-of-thought (CoT) reasoning, an advanced cognitive method that simulates human thinking processes, demonstrating remarkable reasoning capabilities in complex tasks. By structuring complex driving scenarios within a systematic reasoning framework, this approach has emerged as a prominent research focus in autonomous driving, substantially improving the system's ability to handle challenging cases. This article investigates how CoT methods improve the reasoning abilities of autonomous driving models. Based on a comprehensive literature review, we present a systematic analysis of the motivations, methodologies, challenges, and future research directions of CoT in autonomous driving. Furthermore, we propose the insight of combining CoT with self-learning to facilitate self-evolution in driving systems. To ensure the relevance and timeliness of this study, we have compiled a dynamic repository of literature and open-source projects, diligently updated to incorporate forefront developments. The repository is publicly available at https://github.com/cuiyx1720/Awesome-CoT4AD.
Rapid evidence synthesis during emerging infectious and re-emerging disease outbreaks is critical, yet traditional systematic reviews rarely meet urgent timelines. Large language models (LLMs) may accelerate evidence synthesis by extracting data from publications. We compared an LLM-assisted data extraction system with manual extraction. We conducted a 1:1, open-label, 2-period, randomized crossover trial at the National Center for Global Health and Medicine, a national reference center for emerging infectious diseases in Japan (2025). Five experienced reviewers extracted predefined items from mpox-related articles under 2 conditions: (i) LLM-assisted extraction using OpenAI's o3 model to generate structured summaries and (ii) manual review of PDF files. The primary outcome was task completion time; secondary outcomes were extraction accuracy and adverse events. Mixed-effects models included condition as a fixed effect and participant and paper IDs as random effects. The protocol, source code, and data are available at https://github.com/SRWS-PSG/emerging_infection_24K13518_open. Five evaluators (4 physicians and 1 pharmacist; 6-10 years postgraduation) completed 20 task-level evaluations (LLM, n = 9; no LLM, n = 11). Mean completion time was 27.5 minutes with LLM assistance versus 34.5 minutes without. The LLM-assisted condition was 7.9 minutes faster on average (95% CI -1.5 to 17.3; P = .099). Extraction accuracy was 100% in both conditions, and no adverse events were reported. LLM assistance might reduce data extraction time by ∼23% (7.9 minutes per article; 95% CI -1.5 to 17.3 minutes) with no observed loss of accuracy. Although statistical uncertainty remains, LLM integration may offer practical value for rapid evidence synthesis during public health emergencies as tools and prompting strategies mature.
Artificial intelligence (AI) is increasingly integrated into scientific publishing workflows, yet no study has formally evaluated the ability of large language models (LLMs) to reproduce human editorial desk-review (R0) decisions in a general orthopaedic surgery journal. We investigated whether three commercially available LLMs could accurately replicate the editorial decisions of the Editorial Board of Orthopaedics & Traumatology: Surgery & Research (OTSR). The study addressed four questions: (1) Is the concordance between LLM and human R0 decisions satisfactory for editorial use? (2) Do LLMs exhibit a severity bias? (3) Do LLMs generate decision letters of acceptable quality, and do they reproduce the specific critiques of human reviewers? (4) Does prompt complexity influence LLM decision-making? LLMs used without task-specific fine-tuning or prior exposure to the study corpus would demonstrate at least moderate concordance (κ ≥ 0.40) with human editorial decisions. A corpus of 32 manuscripts randomly selected from submissions to OTSR between 2025 and 2026 (n = 10 outright rejected at R0: 3 out of scope, 3 plagiarism/dual submission, 4 direct desk rejection; n = 11 accepted for peer review; n = 11 rejected after full peer review) was anonymised and independently evaluated, without task-specific fine-tuning or prior exposure to the study corpus, by ChatGPT (GPT-5.5, OpenAI), Gemini (3.1, Google), and Claude (Sonnet 4.6, Anthropic) using a structured prompt incorporating the OTSR guidelines. The primary outcome was assessed using Cohen's kappa between LLM and human binary decisions. Secondary outcomes included accuracy, inter-LLM agreement, domain-specific scoring, ARCADIA quality scoring of 115 eligible decision letters by two blinded raters with ICC, human-performed content concordance analysis, sensitivity analysis (structured vs. minimal prompt), and test-retest reproducibility at 24 h. Overall accuracy (i.e. the decision was similar for LLM and editorial decision) was 59.4% for ChatGPT (19/32) and Claude (19/32), and 62.5% for Gemini (20/32). Cohen's kappa was near-zero for ChatGPT (κ = -0.05) and Gemini (κ = 0.00), and low for Claude (κ = 0.15). All LLMs showed systematic over-rejection of accepted manuscripts. No LLM identified plagiarism or simultaneous dual submission as a rejection motive. Test-retest concordance was 84.4-90.6 % across models. ARCADIA quality scoring (n = 115 scorable letters, inter-rater ICC = 0.86, 95% CI 0.80-0.90) showed Claude achieved the highest scores (4.46 ± 0.32 /5), significantly above the human OTSR letters (4.01 ± 0.44, p < 0.001), ChatGPT (3.81 ± 0.49, p = 0.002), and Gemini (3.28 ± 0.49, p < 0.001). LLMs reproduced 30-41 % of human reviewer-specific critiques, with Claude achieving the highest match (40.5%) without hallucinations. Switching to a minimal prompt markedly increased acceptance rates for Gemini (87.5%) and Claude (75%), while ChatGPT remained largely insensitive to prompt simplification (15.6%). The principal finding of this study is the dissociation between formal review quality and true editorial reliability. Although modern LLMs generated persuasive and methodologically structured decision letters, they failed to achieve meaningful concordance with real editorial decisions and displayed stable architecture-specific biases that were highly sensitive to prompt design. These results indicate that current LLMs reproduce the surface features of peer review more successfully than its underlying scientific and contextual reasoning. Consequently, LLMs may represent valuable supervised assistants for editorial workflows, but not reliable autonomous substitutes for human editorial expertise in orthopaedic scientific publishing. IV; Observational pilot study, concordance analysis.
Cardiac amyloidosis (CA) is increasingly recognized in clinical practice. Whether a guideline-based large language model can deliver clinician-level answer quality for CA remains unknown. This study aimed to develop a guideline-based custom generative pretrained transformer (GPT) (AmyloGPT) and evaluate whether its response quality matches or exceeds that of board-certified cardiologists for questions regarding CA. AmyloGPT was built in OpenAI's GPT Builder without programming, integrating the 2020 Japanese Circulation Society CA guidelines as its knowledge base. Ten nonspecialist physicians generated 71 unique clinical questions. Five board-certified cardiologist answerers drafted responses. In a prospective, blinded, comparative study, evaluators (10 nonspecialists and 3 board-certified cardiologists) assessed paired responses for preference (forced-choice) and response quality using five-point Likert scales. Compared with cardiologist answers, AmyloGPT was preferred in 81.1% (95% CI: 78.1%-83.8%) of evaluations by nonspecialist evaluators and 83.6% (95% CI: 78.6%-88.6%) of those by cardiologist evaluators (both P < 0.001). Among nonspecialists, AmyloGPT received higher median ratings for intent alignment and clinical usefulness (both P < 0.001). Among cardiologist evaluators, AmyloGPT received higher median ratings across all 5 quality dimensions: accuracy, consistency, validity, completeness, and absence of bias (all P < 0.001). A no-code, guideline-based custom GPT delivered superior response quality to that of cardiologists for CA questions. This approach allows clinicians without programming skills to build disease-specific large language models, potentially supporting equitable care where specialist access is limited. However, further studies are needed to evaluate potentially inaccurate outputs such as hallucinations.
Background/Objectives: The integration of artificial intelligence (AI) into pharmaceutical development has the potential to accelerate early-stage formulation design. In this study, large language models (ChatGPT (GPT-4o, OpenAI) and DeepSeek (DeepSeek-R1, DeepSeek AI) were evaluated as supportive tools for the design of sustained-release lornoxicam matrix tablets. Using constrained formulation prompts and a predefined excipient space, each model generated candidate formulations intended for direct compression, with the objective of producing sustained-release systems capable of mimicking the dissolution behaviour of a commercial reference product (LOROX OD 16 mg). Methods: The proposed formulations were prepared experimentally and evaluated for physicochemical properties, including weight variation, hardness, friability, and drug content, as well as in vitro dissolution performance over 24 h. Dissolution profiles were compared with the reference product using similarity (f2) and difference (f1) factors, and release behaviour was further characterized using kinetic models. Results: All formulations demonstrated sustained-release behaviour without evidence of dose dumping. One ChatGPT-generated formulation (F3C) met the regulatory criteria for dissolution similarity to the reference product (f1 = 9.66, f2 = 71.31), while the remaining formulations showed variable release behaviour with f2 values ranging from 28.61 to 49.70. However, F3C exceeded the pharmacopeial friability limit marginally (1.108%), while DeepSeek formulations F5D and F6D exceeded pharmacopeial assay acceptance limits. Kinetic modelling indicated a range of transport mechanisms from anomalous diffusion to super Case II transport depending on polymer composition. Conclusions: Although both AI systems successfully generated experimentally viable formulations, prediction accuracy analysis showed high trend-level correlations between AI-predicted and experimental dissolution profiles. However, the magnitude of quantitative error was substantial, with RMSE values exceeding 17% and MAPE values ranging from approximately 38% to 60%. These findings indicate that the models captured general release trends but did not provide reliable quantitative dissolution predictions.
Background/Objectives: Secondary hypertension (SH) requires complex diagnostic reasoning and guideline-based management, posing substantial challenges for artificial intelligence-driven clinical decision-support systems. This study aimed to comparatively evaluate the performance of three large language models (LLMs) in diagnostic reasoning, clinical management, follow-up planning, and patient-oriented communication in SH. Methods: In this cross-sectional blinded study, three LLMs-GPT-5.2 (OpenAI), Claude Sonnet 4.6 (Anthropic), and Gemini 3.0 Pro (Google)-were evaluated using 10 expert-developed clinical case vignettes representing major etiologies of SH. Model outputs were anonymized and independently assessed by three senior clinicians (two endocrinologists and one cardiologist) using a 7-point Likert scale across five domains: (1) diagnostic accuracy and hallucination control, (2) quality and comprehensiveness, (3) reliability and clinical guidance, (4) efficiency of diagnostic workup, and (5) clinical usability. Group differences were analyzed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons. Inter-rater agreement was assessed using two-way mixed-effects intraclass correlation coefficients with absolute agreement. Results: A total of 90 blinded expert evaluations were analyzed. GPT-5.2 (6.0, Q1-Q3 5.40-6.05) and Gemini 3 Pro (5.2, Q1-Q3 4.55-6.20) (H = 40.055, p < 0.001). The results indicated a clear performance hierarchy, with Claude Sonnet 4.6 receiving the highest overall scores, followed by GPT-5.2 and Gemini 3 Pro. Pairwise analyses showed higher scores for Claude Sonnet 4.6 than the other models in most domains, while efficiency of diagnostic workup showed smaller between-model differences. GPT-5.2 generally showed intermediate performance, with higher ratings than Gemini 3 Pro in reliability and clinical usability. Performance differences were most pronounced in domains requiring complex clinical reasoning, whereas efficiency of diagnostic workup scores was relatively comparable among models. Claude Sonnet 4.6 ranked first in nine of the ten clinical vignettes. Inter-rater agreement analyses demonstrated consistent ranking patterns among evaluators. Conclusions: These exploratory findings suggest heterogeneous and model-dependent performance of LLMs in secondary hypertension-related clinical tasks. A clear clinician-rated performance hierarchy was observed, with differences most apparent in domains requiring complex clinical reasoning. However, given the pilot vignette-based design and limited sample size, these results should be interpreted as hypothesis-generating and require confirmation in larger, multicenter validation studies before routine clinical implementation can be considered.
Lnk is an adaptor protein that attenuates cytokine receptor signaling in hematopoietic and immune cells, but its role in the behavior of CD8+ T cells within solid tumors is not well defined. Using a murine melanoma model, we compared tumor growth and immune infiltration in Lnk-deficient mice and in wild-type controls and evaluated CD8+ T cell trafficking toward chemokine signals produced by melanoma cells. Lnk-deficient mice developed smaller tumors containing greater numbers of intratumoral CD8+ cytotoxic T cells. Tumor chemokine profiles, including high levels of interleukin-8 family signals such as CXCL2, were comparable between groups, yet CD8+ T cells lacking Lnk showed enhanced migration toward melanoma-derived cues ex vivo and accumulated more effectively within tumors in vivo. Pharmacologic inhibition of CXCR1/2 abolished this migratory advantage. Upon stimulation with interleukin-8 family chemokines, Lnk-deficient CD8+ T cells exhibited increased activation of STAT3 and ERK signaling pathways. Moreover, antisense-mediated downregulation of Lnk in wild-type T cells resulted in enhanced accumulation within melanoma tumors following adoptive transfer into wild-type hosts. These findings identify Lnk as a negative regulator of CXCR1/2-dependent trafficking of CD8+ T cells in melanoma and suggest that modulating this intracellular checkpoint may improve T cell trafficking into solid tumors and enhance the effectiveness of adoptive cellular immunotherapies.
The approval of lifileucel in 2024 marked an important milestone in oncology as the first cellular therapy authorized for a solid tumor. This milestone stands in sharp contrast to the success of CAR-T cells in hematologic malignancies, where six products have been licensed, and highlights the central challenge that solid tumors remain largely unconquered. At the mechanistic core lies a three-stage framework describing the major barriers encountered by therapeutic T cells in solid tumors, a series of escalating barriers that any therapeutic T cell must overcome to achieve durable tumor control: (1) Access: overcoming stromal and vascular barriers that restrict T-cell infiltration into tumors, (2) Recognition: identifying malignant cells in the setting of antigen heterogeneity and immune evasion, and (3) Persistence: maintaining T-cell function within the immunosuppressive tumor microenvironment. Historically, CAR-T and TIL therapies were viewed in competition, each occupying distinct niches. The field is increasingly adopting a convergent paradigm in which both platforms address a common challenge: overcoming the biological barriers that limit durable responses in solid tumors through complementary engineering and biological strategies. We review the biological obstacles, emerging convergence strategies, and translational frameworks including biomarker-guided patient selection that define this new area. Therapeutic selection may increasingly be guided by a tumor's dominant biological barriers rather than by platform classification alone.