共找到 20 条结果
Multimodal Artificial Intelligence (AI) models-integrating diverse data such as imaging and clinical records-are advancing rapidly in healthcare, yet a significant disconnection persists between these complex predictive architectures and the explainable AI (XAI) techniques used to interpret them. We conducted a scoping review over 4 bibliographic databases to investigate the use of explainability methods in cross-modal medical AI studies. From 82 included studies, we found that the landscape remains dominated by independent feature attribution (assigning importance scores to individual modality in isolation), with the majority of studies relying on post-hoc methods (applied after a model decision is reached) that treat the model as a 'black box'. While emerging trends like visual grounding (linking textual justifications directly to specific image regions) and model reasoning show promise, a critical gap remains in explaining the underlying reasoning process. Standardised evaluation is missing in the majority of studies relying solely on qualitative measures. Only a minority of studies achieve good reproducibility with public codebase. We provide suggestions for the field to transition from individual and post-hoc XAIs toward intrinsically explainable designs where the reasoning logic is built directly into the model architecture to ensure that AI outputs align with human-centric clinical workflows and applications.
Artificial intelligence (AI) is increasingly explored across deep brain stimulation (DBS) for movement disorders, yet whether current systems are approaching deployment remains unclear. To characterise their scope, validation maturity, and translational readiness, we systematically evaluated 239 peer-reviewed studies published between 2000 and 2025, assessing AI methods, validation practices, and barriers constraining clinical translation. Research was dominated by Parkinson's disease and subthalamic nucleus targeting, with limited coverage of other disorders and targets. Most studies reported encouraging internal performance; however, external validation was rare, evaluations remained predominantly retrospective and single-centre, and more than one-quarter involved small-sample, high-dimensional datasets with elevated overfitting risk. Technology readiness assessment revealed that most systems remain at early-to-intermediate translational stages, constrained more by limited validation than by algorithmic inadequacy, compounded by the biological heterogeneity and dynamic complexity inherent to DBS. Nevertheless, emerging external and prospective studies suggest a field moving toward clinical maturity, with promising applications in targeting, programming, outcome prediction, and adaptive therapy delivery.
Neoadjuvant therapy (NAT) is an important treatment strategy in surgical oncology, but not all patients benefit equally from it. This systematic review is the first to evaluate artificial intelligence (AI) models predicting NAT response from hematoxylin and eosin (H&E)-stained biopsies slides of solid tumors. A systematic search across five databases was performed following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, and study quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2). Out of 235 studies, 25 met the inclusion criteria and were analyzed regarding their AI methodologies, data modalities, and type of NAT. Most studies reported area under the curve (AUC) ranging from 0.70 to 0.90, and approximately 40% included external validation cohorts. In conclusion, AI models show promise in predicting NAT response from pathological slides, but future work should emphasize standardized data acquisition, patient-level validation, data transparency, and code sharing.
Artificial Intelligence (AI) technologies are increasingly prevalent in healthcare, yet without adaptive governance, even well-designed systems risk exacerbating health inequalities. This perspective examines five interconnected governance domains: legal frameworks, evidence generation, regulation and market access, workforce readiness, and public trust. We argue these form a cyclical governance chain in which weaknesses cascade across domains, and identify critical gaps and system-level reforms to ensure AI reduces rather than amplifies disparities.
The lateral pelvis is a critical anatomical region in colorectal, gynecological, and urological surgeries. However, its anatomical complexity and variability pose significant challenges for pelvic lymph node dissection (PLND). This study aimed to develop an artificial intelligence (AI) model to identify key anatomical structures relevant to PLND and evaluate whether AI assistance enhances surgeons' ability to recognize pelvic anatomical features. Thirty-six surgeons representing colorectal, gynecological, and urological specialties, with varying experience levels, reviewed 640 video snippets (0.5 s each) from PLND procedures. The model was trained on 23,259 annotated and 653 unannotated images extracted from 293 PLND procedure videos. Threefold cross-validation yielded Dice similarity coefficients of 0.6483 for the ureter, 0.8654 for the obturator nerve, 0.8619 for the external iliac artery, and 0.8736 for the external iliac vein. Across all structures, AI assistance led to a significant improvement in sensitivity and specificity among participating surgeons (p < .001). Our findings suggest that the proposed AI model may assist surgeons in identifying pelvic anatomical structures across different specialties and experience levels. Further studies using continuous intraoperative workflows will be required to determine its impact on clinical practice.
This narrative review explores the latest advancements in various imaging modalities combined with artificial intelligence (AI) for monitoring response to neoadjuvant therapy (NT) and predicting pathological complete response (pCR). We provide a detailed analysis of the principles, evaluation metrics, strengths, and limitations of each technique. The integration of multimodal imaging and AI is reshaping the paradigm for evaluating NT efficacy in breast cancer, offering robust support for precision medicine.
Food choices shape both human and planetary health; yet, designing foods that are delicious, nutritious, and sustainable remains challenging. Here we show that generative artificial intelligence can learn the structure of the human palate directly from large-scale, human-generated recipe data to create novel foods within a structured design space. Using burgers as a model system, the generative AI rediscovers the classic Big Mac without explicit supervision and generates novel burgers optimized for deliciousness, sustainability, or nutrition. Compared to the Big Mac, its delicious burgers score the same or better in overall liking, flavor, and texture in a blinded sensory evaluation conducted in a restaurant setting with 101 participants; its mushroom burger achieves an environmental impact score more than an order of magnitude lower; and its bean burger attains nearly twice the nutritional score. Together, these results establish generative AI as a quantitative framework for learning human taste and navigating complex trade-offs in principled food design.
Artificial intelligence is being embedded in clinical trial infrastructure, shaping who is identified, stratified, and analysed. Opaque models risk amplifying existing disparities in the evidence base. We argue that embedded transparency, the structural integration of ex ante interpretability, demographic auditability, documented uncertainty handling, and stakeholder-relative explanation, is a necessary, though not sufficient, condition for equitable AI-enabled trials, and propose governance recommendations actionable across regulatory regimes.
Adolescent idiopathic scoliosis (AIS) surgery requires precise fusion segment selection and reliable prediction of postoperative alignment, yet current tools lack individualized, validated solutions. We developed ScoliosisPLAN, an AI-based system integrating a YOLOv8-derived segmentation model (ScolioPlanNet) for personalized fusion planning and a latent diffusion model (ScolioPredNet) for simulating postoperative radiographs. In a retrospective development cohort and prospectively collected internal and external validation cohorts of 1425 patients with ≥2-year follow-up, the system achieved performance comparable to experienced surgeons in replicating fusion planning decisions and predicted key radiographic outcomes within clinically acceptable error margins. ScoliosisPLAN provides an interpretable, data-driven framework linking surgical strategy to outcome prediction, supporting standardized, patient-specific decision-making in AIS care.
Preparing patients for cardiac catheterization requires critical but repetitive tasks. We conducted a prospective evaluation of the AI voice assistant, Sofiya, for pre-procedural calls in our center. A customized Large Language Model was trained on medical knowledge and deployed on an agentic AI framework with conversational AI. Sofiya called patients to provide instructions, collect clinical data, and answer questions or redirect the patient to a nurse. The study consisted of a 90-day stabilization period with collaboration between clinicians and developers (Phase I), followed by a 90-day real-world application led by nurses (Phase II). The primary outcome was the overall rate of successfully completed calls reaching the end of the script with all patients answering all clinical questions. From January 16, 2025, to July 17, 2025, 1431 patients received 1606 calls from Sofiya. Phase I had 806 calls whose completion rate gradually improved, reaching an average of 86.4%. High completed call rate (87.9%) was maintained in Phase II with an additional 800 calls. AI system errors were observed in 48 (6.0%) calls during Phase I and 24 (2.6%) in Phase II. A voice-based AI assistant can successfully augment routine tasks of pre-procedural patient preparation, allowing nurses to focus on clinical aspects of patient care.
Dyspnea is a complex symptom measured using subjective patient-reported ratings. Continuous, automated dyspnea measurements are needed, especially in critical care and trauma settings with impaired patient communication. We prospectively enrolled 54 pulmonary rehabilitation subjects. Participants completed two treadmill walking trials, during which dyspnea measurements were collected at one-minute intervals automatically using physiologic sensors and patient-reported ratings of perceived breathlessness and exertion (RPB and RPE). The sensor data were used to train machine learning models using either 19 or 7 features to generate an objective dyspnea score (ODS). Classification performance was assessed on a held-out test set, compared against patient-reported RPB and RPE. The model trained on 7 features performed best, resulting in a correlation between predicted and actual scores (percent accuracy) of 0.84 (78.7%) for RPE and 0.86 (83.6%) for RPB. Our system incorporating ODS accurately predicts patients' subjective dyspnea scores and is promising for automated, real-time dyspnea measurement.
Retinal vein occlusion (RVO) is a chronic retinal vascular disease that often requires repeated anti-VEGF injections and long-term follow-up. However, predicting treatment responses across different follow-up timepoints remains clinically challenging. To address this issue, we developed an AI system integrating generative adversarial networks (GANs), UNet + +, and ResNet-101 to generate post-treatment OCT and fundus images and support clinical decision-making. A total of 2304 OCT and 576 fundus images from 576 RVO patients were collected at baseline and at weeks 4, 12, and 24 after treatment. The generated images demonstrated favorable visual quality, as evaluated by mean absolute error (MAE), peak signal-to-noise ratio (PSNR), and structural similarity index measure (SSIM). The system further quantified lesion areas and predicted retreatment needs, achieving average AUCs of 0.854 and 0.744 across six models in the internal and external test datasets, respectively. In the reader study, the AI system achieved higher predictive accuracy than retinal specialists while substantially reducing image interpretation time. Clinicians' predictive performance also improved with AI assistance.
Large language models (LLMs) are increasingly evaluated using medical examination datasets, yet most studies emphasize overall accuracy rather than the psychometric structure of test items. We evaluated five LLMs on 199 text-only cardiology residency in-service examination items previously characterized using resident-derived psychometric metrics. Three frontier models were compared with two open-source comparators using a standardized zero-shot, repeated-query protocol and strict-majority scoring. Frontier models achieved substantially higher accuracy than open-source comparators, with Claude Opus 4.6, Gemini 3.1 Flash-Lite, and GPT-5.4 reaching 86.4%, 82.9%, and 81.9%, respectively, compared with 53.3% for MedQwen and 18.6% for Qwen-3.5-35B. Across frontier models, performance increased progressively from hard to easy resident-derived item strata. In multivariable analyses, item difficulty was the only classical psychometric factor consistently associated with AI correctness. IRT-based re-analysis confirmed that higher latent item difficulty was independently associated with lower frontier-model accuracy. Human-AI item-level correlations were modest but exceeded permutation-based null expectations, and frontier-model errors were concentrated among highly ranked human distractors. These findings show that item-level psychometric analysis helps explain variation in frontier LLM performance beyond overall accuracy alone. However, expert-rated rationale assessment and none-of-the-above perturbation testing revealed that strong examination accuracy did not guarantee high-quality explanatory support or reliable recognition of answer absence, indicating that examination performance and answer-selection robustness represent related but distinct dimensions of model behavior.
Plant viruses, once viewed as harmful agricultural pathogens, are now powerful tools in biotechnology. Their nanoscale structure, self-assembly, and biocompatibility enable applications in agriculture, medicine, and environmental sustainability. They serve in gene delivery, genome editing, diagnostics, and nanomaterials for vaccines and drug delivery. Integration with AI, ML, and bioinformatics enhances virus discovery and prediction. Despite challenges, plant viruses are emerging as versatile, sustainable resources for global biotechnological innovations.
This perspective examines how artificial intelligence (AI) is transforming small-molecule development for precision cancer immunomodulation therapy. It outlines AI-driven approaches for de novo design, virtual screening, multi-parameter optimization, and ADMET prediction, targeting immune checkpoints, tumor microenvironment modulation, antigen presentation, and metabolic pathways. The article highlights patient stratification, multi-omics integration, digital twin simulations, translational challenges, and future directions, underscoring AI's potential to deliver effective, personalized immunomodulatory therapeutics.
While multimodal large language models (LLMs) demonstrate significant potential in healthcare applications, their clinical utility is difficult to appraise. Current evaluations of medical-assisting LLMs are often limited by sparse human expertise, narrow specialty scope, and reliance on multiple-choice benchmarks or synthetic vignettes, which can inflate performance and obscure clinical utility. We conducted a multicenter, multidisciplinary study in which more than 400 physicians-spanning seven specialties, varied experience levels, and multiple geographic settings-evaluated LLM-generated free-text responses to real, de-identified clinical cases. In a matched-control design, we also deployed an equivalent number of AI agents configured to mirror physician characteristics to examine whether automated evaluators can supplement or replace human assessment. Our results demonstrated that physician assessments exhibited substantial heterogeneity by clinical seniority and practice environment, leading to notable shifts in relative model rankings across cohorts. While AI agents delivered highly efficient, directionally aligned assessments, they did not fully capture the nuances of human clinical judgment and could not substitute for physician-centered evaluation. Instead, they promise assistive tools that can triage or pre-screen outputs to reduce human burden.
Advances in synthetic biology and tissue engineering have enabled the design and assembly of neural constructs from first principles, renewing interest in cybernetics for understanding control, adaptation, and intelligence. To address shortcomings of artificial intelligence, researchers increasingly draw on the learning, robustness, and energy efficiency of living cognitive systems. Synthetic biological intelligences (SBIs) are beginning to leverage embodied biological computation, while cybernetics provides substrate-agnostic principles to guide minimal cognition research.
In vivo confocal microscopy (IVCM) is a critical ophthalmic examination that provides in vivo cytological and neurological information essential for diagnosing corneal and certain systemic diseases, but its clinical utility is limited by time-consuming interpretation and the need for subspecialty expertise. We developed IVCM-Insight, an artificial intelligence (AI) system integrating image-text contrastive learning with large language models (LLMs) for automated report generation and interactive question answering (QA). Based on 30,368 IVCM images and 4155 paired clinical reports, the model was trained with contrastive alignment, image-conditioned language modeling, and multi-image consistency loss to produce structured diagnostic reports while a domain-adapted LLM supported patient-centered QA. Automated evaluation showed strong agreement with the reference reports: Bilingual Evaluation Understudy (BLEU)-1 to BLEU-4 scores were 0.69, 0.58, 0.47, and 0.41, Recall-Oriented Understudy for Gisting Evaluation (ROUGE-L) was 0.67, Consensus-based Image Description Evaluation (CIDEr) was 1.85, and Metric for Evaluation of Translation with Explicit Ordering (METEOR) was 0.66. In addition, the multi-label classification achieved an accuracy of 0.96 and an F1 score of 0.80. Manual assessment by corneal specialists rated report accuracy (4.17), completeness (4.19), coherence (4.70), and diagnostic support (4.06), with excellent inter-rater reliability; QA outputs achieved high accuracy (4.33), relevance (4.54), and non-harmfulness (4.81). Representative cases, including cytomegalovirus, fungal, and Acanthamoeba keratitis, demonstrated accurate detection of key findings and clinically safe explanations. To our knowledge, IVCM-Insight is the first dedicated AI system for comprehensive IVCM interpretation, with potential to enhance diagnostic efficiency, strengthen physician-patient communication, and broaden access to advanced corneal imaging across care settings.
Atrial fibrillation (AF) is frequently asymptomatic and often remains undetected until complications arise. Although artificial intelligence (AI)-enabled electrocardiography (ECG) can predict incident AF from sinus rhythm ECGs, its influence on physician risk assessment in simulated clinical settings remains uncertain. We developed a deep learning model to predict multi-day AF risk using non-AF 12-lead ECGs. For ECG-labeled outcomes, the model achieved an AUROC of 0.79 in the internal EUMC cohort and 0.74 in the external BIDMC cohort. For Holter-labeled outcomes, AUROC values reached 0.87 in the internal EUMC subset and 0.75 in the prospective PROVISION-AF cohort. To assess decision-support utility, a multinational survey of 70 physicians evaluated how AI-derived risk estimates influenced physician risk assessment and follow-up decisions in structured simulated cases. AI assistance significantly improved physicians' AF risk discrimination (AUROC 0.573 to 0.650) and negative predictive value (0.764 to 0.839), with significant net reclassification improvement for non-AF cases (NRI 0.127, p < 0.001). Performance gains were most notable among non-electrophysiologist cardiologists. In conclusion, AI-derived risk estimates improved physician risk discrimination in a structured simulated survey, particularly in non-specialist settings, supporting their potential role as a digital decision-support tool. Further real-world implementation studies are needed to determine whether these effects translate into improved clinical outcomes or healthcare efficiency.
Paroxysmal nocturnal haemoglobinuria (PNH) is a rare, life-threatening hematologic disease with diagnostic delays exceeding 5 years in 24% of cases. We developed and deployed an artificial intelligence algorithm analyzing structured and unstructured electronic health record data across 14 healthcare organizations in Poland. Screening of 1,307,140 patients identified 356 high-risk individuals; of 119 referred for flow cytometry, 13 were diagnosed (positive predictive value: 10.92%; 95% CI, 9.68%-12.30%), comparing favourably to 6.9% conventional screening hit rate. High-risk patients were significantly older (median 69.5 years) with elevated rates of fatigue (76.4% vs 29.19%), anaemia (72.2% vs 7.61%), and myelodysplastic syndrome (49.2% vs 0.24%; all p < 0.001). Only 2.25% presented with haemoglobinuria versus 45-62% in registry cohorts. Retrospective analysis revealed potentially preventable diagnostic delays of 74-1337 days. Monte Carlo feature selection identified Coombs-negative haemolysis and visit frequency as strongest predictors, supporting the potential utility of AI-assisted screening for identifying atypical PNH presentations.