Vision-language models (VLMs) represent an emerging class of multimodal artificial intelligence (AI) systems that integrate visual information with natural-language understanding and generation. In computational pathology, VLMs provide a framework for aligning histologic morphology from whole slide images (WSIs) with pathology reports, and other text-based knowledge sources. This review summarizes the technical foundations, major applications, evaluation strategies, and deployment considerations of pathology VLMs. Current pathology VLMs support a growing range of use cases, including image-text retrieval, label-efficient classification, visual question answering, abnormality localization, anomaly detection, report generation, and agentic workflow support. These capabilities are enabled by image encoders, text encoders or large language models, multimodal alignment strategies, and, in some systems, generative language components. Despite rapid progress, several barriers remain. Evaluation of pathology VLMs is constrained by limited domain-specific benchmarks, insufficient assessment of visual grounding, overreliance on text-based metrics, vulnerability to hallucination, and uncertain robustness under data shift. Clinical translation also requires validation across institutions, scanners, staining protocols, tissue types, and patient populations, together with workflow integration, regulatory oversight, data privacy, cybersecurity, and pathologist accountability. VLMs are therefore best viewed as assistive systems that may augment rather than replace pathologists. Responsible development will require close collaboration among pathologists, computational scientists, health systems, and regulatory stakeholders to ensure that VLMs improves pathology practice in a safe, interpretable, and clinically meaningful manner.
The rise of multimodal generative AI transforms the intersection of technology and art, offering richer insights into large-scale artworks. While significant research has focused on their creative potential, their ability to represent artworks in latent spaces remains underexamined. We use generative AI, specifically Stable Diffusion, to analyze 500 y of Western paintings by extracting two types of latent information with the model: formal aspects (e.g., colors) and contextual aspects (e.g., subjects). Our findings reveal that contextual information exhibits stronger vector alignment and orientation with conventional artistic periods, styles, and individual artists than formal elements. Also, we show how artistic expression aligns with historical shifts using contextual keywords extracted from paintings. Our generative experiment, infusing prospective contexts into historical artworks, validates this vector alignment and orientation by synthesizing artworks consistent with the stylistic patterns of target periods. This study demonstrates how multimodal AI expands traditional formal analysis by integrating temporal, cultural, and historical contexts to quantify the latent structure of cultural knowledge.
Since pain is a multidimensional and subjective experience, pain assessment remains challenging. With advances in artificial intelligence (AI), automatic pain assessment (APA) systems offer a valuable opportunity for objective pain evaluation. However, most approaches focus on a single modality. In this proof-of-concept study, exploring multimodal fusion strategies in a controlled experimental setting, we present a deep learning framework for multimodal fusion that combines facial, acoustic, and textual information to improve APA in cancer patients. A multimodal dataset was created from video-recorded interviews with oncologic patients. In Phase I, audio, video, and transcripts were segmented at the sentence level and temporally aligned using the Eudico Linguistic Annotator (ELAN) to ensure frame-level correspondence across modalities. In Phase II, modality-specific features were extracted: Facial Action Units from OpenFace, acoustic descriptors (MFCCs, chroma, spectral contrast, and Mel-spectrogram) from a dedicated speech-processing pipeline, and sentence-level textual embeddings from ITA-BERT. During training, the most effective analytical strategy was chosen through knowledge transfer approaches. The ELAN-assisted annotation pipeline streamlined expert labeling. Two architectures were implemented and compared: bimodal autoencoder fusion models and a transformer-based model with pairwise cross-modal attention. These models were trained and evaluated using subject-independent and stratified 5-fold cross-validation. To address the lack of independence between segments, a strictly subject-independent cross-validation strategy was adopted. Knowledge transfer using pretrained large-scale models outperformed traditional feature-based approaches and was applied to multimodal pain detection. Multimodal models achieved performance comparable to the strongest unimodal modality (text), while showing improved balance across modalities, suggesting potential complementary effects. Both multimodal architectures demonstrated high accuracy in distinguishing between pain and non-pain classes. The bimodal autoencoder achieved stable results across folds, with a mean accuracy of about 80 % and balanced error distribution. The pairwise transformer with cross-modal attention achieved similar performance, with smooth training and validation loss curves. No evident divergence between training and validation loss curves was observed across folds, suggesting stable behavior within the cross-validation setting. However, subject-level overfitting cannot be excluded given the limited sample size. Multimodal fusion enhances system robustness by integrating complementary signals. Despite limitations and the need for improvement, multimodal deep learning strategies can support the detection of observable pain-related expressions.
Current artificial intelligence (AI) models for medical imaging predominantly focus on a single imaging modality and a single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training these models typically requires large, well labelled datasets, which are costly and labour intensive to prepare. We aimed to train and evaluate an AI model that can interpret diverse imaging modalities across specialties while maintaining robust performance within each modality. We developed Multimodal, Multi-Disease Medical Imaging Foundation Model (MerMED-FM), a multi-specialty model trained using self-supervised learning and a memory module. MerMED-FM was pretrained on publicly sourced, unlabelled medical images from 12 specialties and seven imaging modalities: chest x-rays, CT, ultrasound, histopathology, colour fundus photography (CFP), optical coherence tomography (OCT), and dermatoscopy. After pretraining, the model was fine-tuned, validated, and evaluated for the diagnosis of a range of diseases on 26 public datasets and five private datasets comprising radiology, histopathology, and ophthalmology images. MerMED-FM was compared against a general-domain vision foundation model, various specialist single-modality foundation models, and a multispecialty foundation model. Models were fine-tuned using 10%, 30%, 50%, and 100% of data, with primary comparative analyses conducted using a 10% label fraction. The primary outcome was the area under the receiver operating characteristic curve (AUROC), which was summarised by imaging modality. MerMED-FM was trained on around 3·3 million images from 53 publicly available, unlabelled datasets, comprising 713 931 chest x-rays, 292 353 CT slices, 389 885 ultrasound frames, 1 017 712 pathology patches, 333 099 CFP images, 176 719 OCT slices, and 401 059 dermatoscopy images. Strong performance was achieved across all modalities at a label fraction of only 10%, with mean AUROC values of 0·844 for chest x-rays, 0·906 for CT, 0·818 for ultrasound, 0·908 for histopathology, 0·810 for CFP, 0·962 for OCT, and 0·827 for dermatoscopy. MerMED-FM has the potential to be a highly adaptable, versatile, cross-specialty foundation model that enables robust interpretation of medical imaging across diverse medical disciplines. National Medical Research Council, Singapore and the Agency for Science, Technology and Research, Singapore.
Tuberculosis (TB) remains a global health crisis, with complex molecular mechanisms that are not fully understood. Traditional pathway analysis methods fail to capture the intricate non-linear relationships within biological networks. We developed a novel artificial intelligence framework integrating graph neural networks (GNNs), transformer architectures, and multimodal deep learning to decipher TB pathway mechanisms. Our approach constructs a comprehensive pathway-gene interaction network from three critical pathways (Tuberculosis hsa05152, Antigen processing hsa04612, and NF-κB signaling hsa04064) and employs three interconnected models: (1) a Graph Convolutional Network for learning pathway-gene relationships, (2) a Transformer encoder for pathway activity prediction, and (3) a multimodal fusion model with attention mechanisms integrating transcriptomic, pathway, and clinical data. The framework was trained and validated on 467 clinical samples and 529 transcriptomic samples from five GEO datasets. The Transformer model achieved strong performance in pathway activity prediction (R 2 = 0.97, MSE = 0.014), demonstrating high accuracy in capturing pathway activation patterns. The multimodal fusion model achieved strong predictive performance (accuracy 88.2%, AUC-ROC 0.90) in clinical outcome prediction, with attention analysis revealing adaptive weighting of different data modalities. Network analysis identified 27 shared genes between Tuberculosis and Antigen processing pathways, and 18 shared genes between Tuberculosis and NF-κB pathways, indicating coordinated immune regulation. Key pathway-gene interactions were identified, including critical roles of IFNG, TNF, IL1B, and NF-κB signaling components. This study represents the first comprehensive application of GNNs and multimodal deep learning to TB pathway analysis. Our framework provides novel insights into TB pathogenesis, identifies potential therapeutic targets, and demonstrates the power of AI-driven approaches for understanding complex disease mechanisms. The interpretability of our models through attention mechanisms enables translation of computational findings into actionable biological insights, with significant implications for precision medicine and personalized TB treatment strategies.
Artificial intelligence (AI) has not seen the clinical uptake that might be expected from a technology that has received so much attention and investment. Coupled to neuroimaging, it is conceivable that AI algorithms can provide better performance in diagnosis and prognosis as well as optimise treatments in a precision medicine regimen, all of which is focused on improving both the patient experience and clinical outcomes. But in practice, little impact has been seen. Why might this be the case? This overview focuses on the possible reasons. First, technical concerns: the way in which AI algorithms are developed and validated in the research setting does not adequately prepare them for deployment in clinics and hospitals. Second, the importance of asking clinical questions with AI algorithms that have meaning and value is often underplayed or not considered. The outputs of AI algorithms mostly, but not always, also need to explain how decisions have been made. Thirdly, operational and ethical considerations loom over the integration of AI algorithms into electronic health record systems, clinical pathways, and legal frameworks. Above all these considerations is the motivation for deployment and particularly whether it is primarily for patient benefit or service economics.
Selection of optimal chemotherapy regimens remains a complex clinical challenge due to interpatient heterogeneity, evolving therapeutic options, and the limitations of population-based clinical guidelines. AI has emerged as a promising tool to support precision oncology by integrating multidimensional data to guide individualized treatment decisions. This systematic review evaluates the role of AI-based models in chemotherapy regimen selection, focusing on their impact on treatment efficacy and adverse drug reactions compared with conventional physician-driven decision-making. A systematic literature search was conducted up to March 21, 2026. Studies evaluating AI-guided chemotherapy selection or treatment decision-support systems in cancer patients were included. The population, exposure, comparison, and outcomes (PECO) framework included cancer patients receiving AI-guided chemotherapy selection versus physician judgment or guideline-based care, with outcomes including survival, treatment response, and toxicity. A total of 1,409 records were identified, with 15 studies meeting the inclusion criteria after screening and eligibility assessment. The included studies encompassed diverse malignancies, including breast, prostate, pancreatic, lung, head and neck, glioblastoma (GBM), hepatocellular carcinoma (HCC), nasopharyngeal carcinoma (NPC), and acute myeloid leukemia (AML). AI models utilized multimodal data sources, such as clinical variables, histopathology, imaging, and multi-omics datasets. Across studies, AI-guided treatment selection was associated with improvements in several clinical outcomes, including overall survival, progression-free survival, and pathological response rates. Several models showed an enhanced ability to identify patients unlikely to benefit from specific chemotherapies, thereby enabling treatment de-escalation. Limited but notable evidence suggested reductions in treatment-related toxicity, particularly cardiotoxicity, when AI-guided strategies were employed. Most studies compared AI performance against physician clinical judgment or guideline-based approaches. AI-assisted chemotherapy regimen selection shows considerable potential to improve treatment efficacy and personalize oncology care while reducing unnecessary toxicity. Although current evidence is largely retrospective and heterogeneous, findings consistently support AI as a valuable adjunct to clinical decision-making. Prospective validation and integration into real-world workflows are essential to establish its role in routine cancer care.
Artificial intelligence (AI) is increasingly used in education, yet less is known about the psychological processes through which AI-supported teaching relates to student engagement, especially in vocational education. Drawing on Self-Determination Theory, this study examined the mediating roles of perceived competence and perceived autonomy in the relationship between AI-supported teaching and student engagement. A 10-week quasi-experimental study was conducted with 148 second-year vocational college students, assigned to either an AI-supported teaching group (n = 74) or a conventional instruction group (n = 74). Multimodal data were collected, including academic assessments, performance evaluations, self-report measures of perceived competence, perceived autonomy, and engagement, as well as behavioral records from AI-supported learning activities. Mediation analysis was performed using bootstrapping with 5,000 resamples. Students in the AI-supported teaching group showed higher levels of student engagement than those in the conventional instruction group (Cohen's d = 0.84). Perceived competence showed a statistically significant indirect association between AI-supported teaching and student engagement [β = 0.22, 95% CI (0.13, 0.31)], supporting H2. In contrast, perceived autonomy did not show a statistically significant indirect effect [β = 0.02, 95% CI (-0.02, 0.06)], and therefore H3 was not supported. The direct association between AI-supported teaching and engagement remained significant after accounting for both mediators. These findings suggest that competence-related experiences may represent an important psychological pathway linking AI-supported teaching and student engagement in structured vocational education contexts. The non-significant role of perceived autonomy indicates that motivational processes in AI-supported learning may vary across instructional settings and should be interpreted in relation to contextual and pedagogical conditions. Given the quasi-experimental design and the context-specific sample, the findings should be interpreted cautiously and further examined through larger, multi-site, andlongitudinal studies.
Conversational artificial intelligence (AI), including text-based chatbots, voice-based agents, multimodal systems, and socially assistive robots (SARs), offers a scalable adjunct to therapist-led dementia care. The post-2022 emergence of large language models (LLMs) has accelerated development, yet few reviews apply a unified conversational AI taxonomy across dementia care. This review synthesized the effectiveness, limitations, and implementation challenges of conversational AI across the dementia care continuum. Six databases (PubMed, Embase, Web of Science, Scopus, IEEE Xplore, ACM Digital Library) were searched for English-language studies (January 2010-March 2026) evaluating conversational AI targeting cognitive, social, or caregiver outcomes. Two reviewers independently screened and extracted data following PRISMA 2020 guidelines; risk of bias used standard tools and findings were synthesized narratively. PROSPERO CRD420261333625. Forty studies (8 randomized controlled trials [RCTs], 32 non-randomized) were included. SARs were the largest category (n = 24; 60.0%), followed by text-based chatbots (n = 12; 30.0%), multimodal systems (n = 3; 7.5%), and voice-based chatbots (n = 1; 2.5%). The strongest cognitive evidence came from a social robot RCT (gain of 3.9 points on a 30-point screening measure (p < 0.001). For caregivers, an international RCT (n = 274) showed significant reductions in depression (d = 0.37) and burden (d = 0.34). LLM-based systems produced an 18-fold increase in conversation duration. Speech recognition failure was the most consistently reported technical barrier. Conversational AI shows directional benefit across cognitive, social, and caregiver outcomes. Critical research gaps remain regarding voice-only randomized evidence and adequately powered LLM trials against usual care.
Respiratory diseases are a major global health challenge. However, identification of respiratory diseases is often limited by subjectivity, environmental noise and inter-clinician variability. This study presents an explainable multimodal deep learning framework for recording-level multiclass classification of respiratory audio signals. The proposed system integrates two complementary representations-a spectro-temporal encoder based on a CNN-BiLSTM-attention architecture and a handcrafted acoustic-feature encoder capturing acoustic descriptors commonly used in respiratory-audio analysis, including MFCCs, zero-crossing rate, spectral centroid, spectral bandwidth, chroma, RMS energy, and spectral rolloff features. These branches are combined through late-stage fusion to leverage both data-driven representation learning and domain-informed acoustic cues. The proposed model was trained and internally evaluated on the Asthma Detection Dataset Version 2, comprising five respiratory categories: bronchial disease, asthma, COPD, healthy, and pneumonia. Mono conversion, resampling to 16 kHz, 100-2000 Hz band-pass filtering, amplitude normalisation, fixed 4 s trimming or zero-padding, training-only augmentation, handcrafted-feature extraction, mel-spectrogram generation, quality control auditing, and stratified recording-level partitioning have been applied in the pre-processing steps. Across five repeated experiments with different random seeds, the proposed hybrid model achieved a mean held-out recording-level test accuracy of 0.9099±0.0163, balanced accuracy of 0.8936±0.0152, macro F1-score of 0.8937±0.0177, macro ROC-AUC of 0.9867±0.0010, and macro PR-AUC of 0.9489±0.0044. Conventional machine learning baseline comparisons showed that the proposed model achieved stronger internal accuracy, balanced accuracy, macro recall, macro F1-score, and macro ROC-AUC than classical machine learning algorithms trained on handcrafted acoustic features, although Random Forest remained competitive in macro PR-AUC. Ablation analysis shows that the deep spectro-temporal branch was the primary contributor to predictive performance, while the handcrafted branch provided complementary interpretable acoustic information rather than consistently improving all classification metrics. Explainability was incorporated using Grad-CAM and Integrated Gradients for spectrogram-based interpretation and SHAP for handcrafted-feature attribution. Domain-shift evaluation on the ICBHI Respiratory Sound Database and a COPD-focused cohort revealed substantial dataset shift effects, including poor healthy-case recognition on ICBHI and seed-dependent COPD recognition in the COPD-focused cohort. Identifier-aware sensitivity analyses showed lower performance than the main recording-level split, suggesting that subject-like or source-level overlap may inflate internal performance estimates. The findings should be interpreted as promising internal held-out recording-level algorithmic performance with limited external transfer, rather than evidence of readiness for clinical use.
Three-dimensional (3D) bioprinting has reached a complexity limit where empirical, parameter-by-parameter optimization no longer scales. The dominant mode of artificial intelligence (AI) integration remains AI-augmented, where AI is treated as an analytical addition to a conventional pipeline. We argue that the field is approaching a discontinuous transition towards AI-native bioprinting, in which AI represents the operational layer of system intelligence, not an ancillary tool. A systematic analysis of 365 publications on the intersection of bioprinting and AI (2015-2026), performed through 18 queries organized by the four search axes of the PubMed database, shows that the intersection grew 136 times during the decade, with an acceleration of 3.16 times only between 2024 and 2025. Mapping the publications to the six functional domains reveals a marked asymmetry: clinical translation counts 154 papers, while cell viability prediction-the biological foundation that every closed-loop system requires-counts only three. We define AI-native bioprinting as a system architecture that combines continuous learning, multi-modal sensing fused through visual, mechanical and biological signals, and biologically closed control loops. We present a conceptual shift from printing accuracy to biological intelligence as a success criterion. The transition requires open datasets, consensus biological metrics, inter-laboratory validation, and early regulatory engagement.
Advances in monitoring systems featuring wearable sensors, computer vision, and artificial intelligence (AI) have been increasingly used in sports science and rehabilitation practices as a means of movement pattern analysis, injury prevention, and training optimization. These technologies are becoming essential components of athlete-performance analysis and rehabilitation-monitoring systems designed to support biomechanical assessment, athlete development, and movement-quality evaluation. Athlete-performance analysis and rehabilitation monitoring increasingly rely on intelligent multimodal sensing systems capable of continuously evaluating movement quality, biomechanical patterns, training execution, and recovery progress. Human activity recognition (HAR) serves as a key enabling technology for these applications by providing automated assessment of human movement using wearable and vision-based sensing modalities. Therefore, the purpose of this study was to develop and evaluate an attention-based multimodal framework that integrates wearable inertial sensing and RGB video analysis for robust athlete-performance assessment and rehabilitation monitoring through accurate recognition of human movement patterns. Athlete-performance analysis and rehabilitation monitoring combining inertial sensor data and RGB-based visual information was introduced. Inertial signals were segmented with adaptive windowing, whereas silhouette refinement was performed to analyze motion structures from visual inputs in support of athlete-performance analysis and rehabilitation monitoring. Temporal, spatial, and motion features such as trajectory, orientation, and skeleton-based space-time representations were calculated from multimodal inputs. The proposed framework was designed to capture complex movement dynamics associated with rehabilitation exercises and sports-related motion patterns across heterogeneous sensing environments. Extracted features were then combined and optimized with a multimodal feature fusion approach, while the Ranger optimization algorithm was utilized during the process. An attention-based deep learning classifier was implemented to classify movement activities. The results showed that the proposed framework reached accuracy scores of 88.40% and 87.96% on the VIDIMU dataset and the UTD-MHAD dataset respectively. Recognition performance across both inertial and vision-based modalities provided greater robustness than single-modality solutions. The integration of wearable sensing and computer vision modalities further improved the ability of the framework to analyze complex movement behaviors under varying execution conditions and environmental variations. The proposed multimodal framework provides a foundation for intelligent athlete-performance and rehabilitation-monitoring systems by integrating wearable sensing, computer vision, and attention-based artificial intelligence for robust movement analysis. The findings highlight its potential to support biomechanical assessment, movement-quality evaluation, training-performance monitoring, rehabilitation tracking, and injury-risk management in modern sports and healthcare environments.
English proficiency is vital for non-native speakers' career development, yet classroom instruction alone cannot meet practical demands, making informal digital learning of English (IDLE) increasingly important. Artificial intelligence (AI), with conversational and multimodal functions, offers new opportunities for IDLE. However, existing research on AI-mediated IDLE has predominantly focused on language majors and often relied on a single methodological lens, neglecting STEM undergraduates and the complex network dynamics among motivational factors. However, research has largely focused on language majors, leaving STEM majors underexplored. Guided by the Hedonic-Motivation System Adoption Model (HMSAM), this study analyzed data from 413 Chinese STEM majors using partial least squares structural equation modeling (PLS-SEM, SmartPLS 4.0) and psychological network analysis (PNA, R 4.5.3). PLS-SEM results showed that enjoyment was the strongest direct predictor of AI-IDLE, followed by focused immersion, perceived usefulness, and curiosity. Control contributed indirectly via focused immersion, while boredom was non-significant. Perceived ease of use influenced AI-IDLE only through cognitive and emotional pathways. The model explained 58.1% of the variance. PNA further identified enjoyment, focused immersion, and control as central nodes, while the link between perceived usefulness and AI-IDLE was non-significant. These findings suggest that Chinese STEM undergraduates' AI-IDLE is primarily driven by intrinsic hedonic motivations rather than utilitarian evaluations. The study provides empirical support for designing AI tools that enhance enjoyment and control to foster STEM students' extracurricular English engagement.
Medication-related harm remains a major patient-safety challenge, substantially driven by fragmented pharmacovigilance ecosystems, inconsistent drug nomenclature, heterogeneous interaction knowledge bases, and the absence of unified multimodal medication-safety infrastructures. Existing clinical decision-support systems frequently remain interaction-centered, with limited integration of complementary safety domains such as lactation risk assessment, intravenous compatibility evaluation, and regulatory toxicity overlays. This study aimed to develop and evaluate SafeRx, a provenance-aware pharmaceutical knowledge integration platform designed to harmonize heterogeneous medication-safety data within a unified canonical framework. SafeRx integrated Romanian regulatory product summaries, Danish drug-drug interaction repositories, OpenFDA boxed warning datasets, LactMed lactation safety records, and Stabilis intravenous compatibility data through automated extraction pipelines, ontology-aware canonical substance normalization, AI-assisted semantic enrichment, and pharmacist-supervised governance workflows. The platform consolidated 2190 canonical substances linked to 9463 validated identifier mappings, 9932 canonical interaction pairs enriched with tier-based prioritization and mechanism-aware annotations, 514 regulatory boxed warning records, 1895 lactation safety entries, and 9088 intravenous compatibility records. SafeRx demonstrates the feasibility of constructing an interoperable, explainable, and pharmacist-supervised medication-safety infrastructure capable of integrating heterogeneous regulatory, pharmacologic, reproductive, and physicochemical safety domains within a unified clinically navigable framework for research, pharmacovigilance, and clinical decision-support applications. Future studies should evaluate prospective clinical implementation, electronic health record interoperability, and the real-world impact of tier-based prioritization on prescribing workflows, alert fatigue, and medication-safety outcomes.
Recent advances in wearable, vision-based, trajectory, physiological, and multimodal sensing technologies, together with deep learning, have enabled continuous, objective, and individualized assessment of sport performance and athlete health. Unlike prior reviews that primarily focus on a single sensing modality, sport, or algorithmic series, this review integrates wearable, vision-based, trajectory, physiological, and multimodal sensing streams with deep learning models across both performance analysis and athlete health monitoring, thereby clarifying modality-task-model relationships and translational limitations. This review synthesizes recent progress in sensor-based sports intelligence, focusing on how heterogeneous data streams are transformed into performance- and health-related decision support. The reviewed applications include athlete and ball perception, multi-object tracking, pose estimation, action recognition, trajectory and tactical analysis, training-load and fatigue monitoring, injury-risk prediction, rehabilitation monitoring, and return-to-play support. Deep learning architectures, including CNNs, LSTMs, GRUs, TCNs, Transformers, attention mechanisms, graph neural networks, and multimodal fusion models, are discussed in relation to their suitability for visual, temporal, spatial, physiological, and multisource data. This review further identifies key challenges, including data heterogeneity, annotation scarcity, limited cross-sport and cross-device generalization, real-time deployment constraints, model interpretability, privacy protection, and ethical governance. Moving forward, research efforts should focus on the development of standardized datasets, reliable multimodal data fusion strategies, self-supervised and transfer learning approaches, and deployment on edge or cloud computing platforms. Additionally, enhancing interpretability through explainable AI and implementing closed-loop, individualized monitoring systems are critical. By synthesizing advances in sensing technologies, deep learning methodologies, and real-world applications, this review aims to provide a practical reference for optimizing athletic performance, preventing injuries, guiding rehabilitation, and supporting long-term health management of athletes.
Individuals with psychiatric disorders frequently experience comorbid cardiometabolic conditions, complicating treatment and worsening health outcomes. Both psychiatric and cardiometabolic disorders have been individually associated with alterations in brain structure. Yet, it remains unclear whether these associations stem from a shared genetic basis that underlies their frequent co-occurrence. We analyzed genome-wide association summary statistics from large international consortia of individuals of European ancestry, including psychiatric disorder GWAS with case-control sample sizes ranging from ~18,000 to ~158,000 cases, cardiometabolic disease GWAS with up to ~242,000 cases, and cortical morphology GWAS from UK Biobank comprising ~39,000 individuals. We applied complementary multivariate, causal, and mediation genetic analyses to disentangle genetic factors underlying brain alterations and comorbidity. Here we show that patterns of genetic overlap differ across disorders. Schizophrenia exhibits substantial polygenic overlap with cortical thickness and type 2 diabetes, despite low genetic correlation. In contrast, attention-deficit/hyperactivity disorder (ADHD) is more strongly correlated with cardiometabolic disease but shows limited overlap with cortical morphology. Notably, cortical surface area partly mediates the genetic association between ADHD and type 2 diabetes. Pathway analyses highlight metabolic stress processes in ADHD as well as neurodevelopmental and immune processes in schizophrenia. These findings indicate that psychiatric-cardiometabolic comorbidity arises through both shared and disorder-specific genetic pathways. This work clarifies the genetic architecture of multimorbidity and highlights opportunities for trait-targeted prevention strategies in psychiatry. Many people with psychiatric disorders also experience physical health problems, such as heart disease or diabetes. It remains unclear whether these links are due to shared genetic factors. In this study, we used large genetic datasets from hundreds of thousands of participants to investigate how psychiatric disorders, brain structure, and cardiometabolic diseases are genetically connected. We applied statistical methods to identify shared genetic influences and to explore whether one trait may partly influence another. We found that schizophrenia and ADHD show distinct genetic patterns: schizophrenia shares more genetic factors with brain structure and diabetes, while ADHD is more strongly linked to metabolic pathways. These findings highlight that different mental disorders may involve different biological routes, which could inform prevention and treatment strategies in the future.
Brain-computer interface (BCI) technology represents a critical frontier in neurorehabilitation. This study aims to systematically analyze the global research landscape, hotspot distribution, and evolving trends of BCI interventions for upper limb rehabilitation in stroke survivors between 2016 and 2025. Bibliometric analysis and systematic mapping were conducted using data from the Web of Science Core Collection and PubMed. Literature was retrieved using terms related to "stroke," "brain-computer interface," and "upper limb rehabilitation." Screening followed the PRISMA guidelines. Visualization and quantitative mapping were performed using CiteSpace (v.6.4.R2) and VOSviewer (v.1.6.20) to evaluate publication volume, international collaboration, and keyword co-occurrence clusters. Annual publications increased steadily from 37 in 2016 to 104 in 2025, with 65.6% published since 2020. The United States (n = 144), China (n = 83), and Italy were the most productive countries. Keyword analysis revealed a paradigm shift from functional electrical stimulation toward robotics-assisted therapy, motor imagery, and AI-driven decoding. Significant burst strengths were observed for "closed-loop systems," "generative AI," and "multi-modal feedback," indicating these as the current primary frontiers. BCI research for post-stroke recovery is transitioning from experimental signal processing to intelligent, multi-modal, and personalized clinical systems. Bibliometric evidence confirms that integrating BCI with robotic-assisted rehabilitation or functional electrical stimulation (FES) has become the mainstream clinical trend. Future efforts must focus on improving EEG signal stability and developing user-friendly hardware to facilitate the transition of BCI from research settings to daily clinical practice. China has emerged as the second most productive country, though international cooperation with European institutions remains an area for further growth.
The significant risks posed by per- and polyfluoroalkyl substances (PFAS) to water quality and public health have attracted increasing attention as a class of emerging contaminants. Efficient onsite assay of PFAS in environmental waters is critical for risk traceability, early warning, and safeguarding water security, yet it is hindered by challenges in accuracy, practicality, and intelligence. Here, we proposed a universal artificial intelligence (AI)-assisted molecularly imprinted polymer (MIP) gate-controlled the enzyme-like activity of nanozyme strategy-driven multimodal onsite assay for PFAS in environmental waters, with perfluorooctanoic acid (PFOA) selected as the model target. Briefly, MIP-encapsulated Fe-doped coordination polymer (MIP@Fe-BDC) nanozyme was prepared, and MIP@Fe-BDC possessed peroxidase-like (POD-like) activity. Owing to the selective binding capability conferred by MIP, only PFOA could inhibit POD-like activity of MIP@Fe-BDC. This inhibition prevented the oxidation of colorless 3,3',5,5'-tetramethyl-benzidine (TMB) to its blue oxidized form (oxTMB), thereby interrupting the colorimetric and photothermal signal enhancement as well as the fluorescence signal quenching triggered by oxTMB, resulting in a linear response between PFOA and colorimetric/fluorescence/photothermal signals. To achieve onsite detection, a low-cost MIP@Fe-BDC-based test paper was developed. Multimode images were collected via a smartphone and a thermal imager, and then analyzed using a residual neural network with 18-layer (ResNet18) model for real-time quantitative feedback, in which the whole detection process was completed in just 8.0 min. Moreover, these results were consistent with those obtained using the liquid chromatography-mass spectrometry (LC-MS) method, implying the superior accuracy. This work provides an innovative and universal solution for the portable, low-cost, rapid, efficient, and intelligent onsite surveillance, traceability, and early warning of PFAS in environmental waters, which is of great significance for preventing and controlling environmental water pollution and protecting public health.
Artificial intelligence (AI) is being increasingly used in educational and mental health contexts, yet many emotion-related applications still prioritize detection, classification, and automated feedback over contextual understanding. This study uses the Stanford AI Index Reports (2021-2025) as an exploratory discourse corpus to examine how prominent AI reports frame technology, application domains, and governance, and to consider what this framing implies for emotion regulation through sport and exercise in higher education. Across 1813 report pages, we applied BERTopic with multilingual sentence embeddings (paraphrase-multilingual-MiniLM-L12-v2), UMAP dimensionality reduction, HDBSCAN clustering, and class-based TF-IDF, followed by dynamic and hierarchical topic analysis and theory-informed synthesis. Of 22 topics generated, 13 relevant to the study focus were retained and validated through keyword inspection, representative-text review, and independent expert agreement. The analysis indicated a three-layer structure: a technology core, an application-expansion layer, and an ethics-and-governance layer. Health and education themes grew most across reports, with medicine/health rising from 14 to 105 and school pathways from 3 to 105 segment occurrences between 2021 and 2025, whereas sport, exercise, embodied activity, and campus support appeared only indirectly. As prominence reflects raw frequency across five reports, trends are read descriptively. We propose a human-technology-environment framework comprising multimodal contextual profiling, autonomy-supportive task adaptation, feedback-reflection-practice loops, peer and campus support integration, and human-in-the-loop governance. The study does not test intervention effects; its contribution is conceptual and agenda-setting, clarifying a gap between mainstream AI discourse and the embodied, relational, and ecological conditions through which sport and exercise may support students' emotion regulation.
Artificial intelligence (AI) has rapidly emerged as a transformative tool in virology, offering new opportunities for the detection, classification, and surveillance of viral pathogens. Recent advances in machine learning, deep neural networks, and multimodal data analysis now enable the identification of viral signatures from genomic sequences, medical images, environmental samples, and social-media-derived epidemiological signals. This review provides a comprehensive overview of state-of-the-art AI methodologies applied to viral pathogen research, with a particular focus on image-based diagnostics, automated quality assessment of virology-related digital content, and predictive modelling for outbreak monitoring. We discuss how convolutional and transformer-based architectures are being used to classify infected tissues, detect viral particles, and support laboratory workflows. Furthermore, we highlight the emerging role of AI in evaluating the reliability of user-generated images and short videos related to infectious diseases, an area increasingly relevant in the age of misinformation. Challenges such as dataset bias, limited annotated virological images, ethical concerns, and the need for standardized quality-assessment pipelines are critically examined. Finally, we outline future research directions, including hybrid AI-biological models, AI-supported viral surveillance in healthcare environments, and the integration of explainable AI to enhance clinical trust.