The current global status of breastfeeding is marked by both progress and challenges. Digital health interventions (DHIs) have emerged as a promising new strategy for improving breastfeeding practices, yet evidence regarding their impact on breastfeeding outcomes remains limited. To evaluate the impact of DHIs on breastfeeding practices and outcomes. Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we searched PubMed, Web of Science, the Cochrane Library, Embase, CINAHL, Scopus, IEEE Xplore, and grey literature databases, from inception to April 2026. We included randomized controlled trials (RCTs) and quasi experimental studies that enrolled pregnant women or lactating mothers using DHIs versus standard care, waiting lists, or placebo. Outcomes were exclusive breastfeeding (EBF) rates, self efficacy, and knowledge. Studies recruiting mothers with infectious diseases or severe substance use disorders, non English publications, unavailable full texts or data, animal studies, letters, conference proceedings and studies rated as "high risk" were excluded. Study quality was evaluated using RoB 2.0 and ROBINS-I tools. Evidence certainty was assessed using GRADE (Grading of Recommendations, Assessment, Development, and Evaluation). We performed analysis using Review Manager 5.4 and R-4.6.0 software, estimated using the Hartung-Knapp-Sidik-Jonkman (HKSJ) random-effects model. Fifty-two studies involving 11,704 participants were included. DHIs were found to improve exclusive breastfeeding rates (at <3 months: relative risk [RR] 1.33, 95% CI 1.18-1.50; 95% prediction interval [PI] 0.85-2.08; P<.001, I²=84%; at 3-6 months: RR 1.46; 95% CI 1.22-1.75; 95% PI 0.83-2.55; P<.001, I2=75%; at ≥6 months: RR 1.64, 95% CI 1.27-2.11; 95% PI 0.74-3.62; P<.001, I2=86%) and breastfeeding self-efficacy (standardized mean difference [SMD] 0.67, 95% CI: 0.31-1.03; 95%PI -1.17-2.51; P < .001, I² = 95%). No statistically significant effects were observed regarding breastfeeding knowledge. The exploratory results of subgroup analysis suggested that antenatal intervention delivery, intervention duration ≥ 3 months, study settings in developing countries, and participants including preterm parturients, adult women, and cesarean delivery mothers could improve the exclusive breastfeeding rate. In contrast, larger effect sizes for breastfeeding self-efficacy were observed among women with intervention duration < 3 months, full-term delivery, and singleton pregnancy. Overall, among the RCTs, 12 were low risk and 30 had some issues. Among the quasi-experimental studies, 3 were low risk and 7 as moderate risk. Evidence certainty ranged from low to moderate. DHIs show promise in improving exclusive breastfeeding rates and breastfeeding self-efficacy. Compared with previous reviews, this review provides a more comprehensive evaluation of DHIs and identifies potential effect modifiers through subgroup analyses, offering updated evidence for digital breastfeeding support. However, moderate risk of bias, substantial heterogeneity, wide 95%PI, and low-to-moderate GRADE certainty warrant cautious interpretation. Further high-quality studies are needed to confirm the effectiveness and optimize the implementation of DHIs in clinical practice. PROSPERO CRD 420251233352; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251233352.
Evidence shows limited use of routine data to support health actions in lower and middle-income countries (LMICs). To address that, there are several ongoing efforts to strengthen routine health information systems which are largely crippled by the fragmented nature of the health system. This scoping review aims to identify the key enablers of routine health data use and examine gaps in the institutionalization of data use practices across different levels of the health system in LMICs. We scoped up four literature sources, that is, PubMed, IEEE Xplore, the Public Library of Science (PLOS), and Google Scholar for the article published between January 2019 and December 2023. Three reviewers independently screened article titles, abstracts, and full text against inclusion and exclusion criteria, following the Arksey and O'Malley and PRISMA framework to guide the review process. We classified our findings into three health system levels: that is facility, district, and national/sector-wide. Out of 380 screened articles, 41 were selected for inclusion, where most of the articles (48.8%), were on the health facility level. The study found that factors influencing the use of routine data vary across health system levels, with human resource capacity and friendly data collection tools being more important at the facility level. In contrast, the availability, and capacity to use analytical tools and governance structures were more reported at the district and national levels. Further, capacity building in health information systems, IT infrastructure, and regular supervision were crucial across all health system levels. We argue for focused interventions to be designed to institutionalise routine data use practices for better healthcare outcomes across health system levels.
Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.
Falls among older adults are a leading cause of morbidity and loss of independence. Wearable sensors combined with machine learning (ML) offer opportunities for objective fall risk evaluation, but low model transparency limits clinical adoption. Interpretable and explainable artificial intelligence (XAI) methods can address this constraint, yet their application in wearable sensor-based fall risk assessment has not been systematically examined. A PRISMA 2020-compliant systematic review was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore. Studies were eligible if they included older adults, employed wearable sensors, hybrid sensor systems, or structured clinical assessment instruments, applied AI/ML for fall risk assessment (not detection), and incorporated intrinsic interpretability or post-hoc XAI. Data were extracted on population characteristics, sensor modalities, outcome definitions, ML algorithms, explainability strategies, validation methods, and predictive performance. Eleven studies (2019-2025, total n = 5,484) met inclusion criteria. Inertial Measurement Units predominated. Fall risk definitions were heterogeneous, spanning retrospective fall history, clinical balance scales, and prospective diaries. Explainability was implemented almost exclusively at the global level, with only two studies providing both global and local explanations. On comparable tasks, interpretable models achieved accuracy of 0.65-0.92 and AUC of 0.70-0.92, suggesting interpretability carried no consistent performance penalty over more complex designs. All studies relied on internal validation only; none performed external validation or real-time deployment. None recruited P&O users or incorporated device-specific predictors. Current models favour interpretable architectures and achieve moderate-to-high performance, but are constrained by heterogeneous outcome definitions, absence of external validation, and global-only explainability that limits individual-level clinical utility. The evidence base does not yet support clinical deployment. Extending these findings to prosthetics and orthotics users, a clinically important downstream application, will require device-specific datasets, asymmetry-adjusted thresholds, and instance-level explanations. These represent the priority directions for the next stage of this research agenda.
Cardiovascular diseases (CVDs) remain the leading cause of mortality worldwide, creating an urgent need for accurate, trustworthy, and clinically deployable artificial intelligence (AI) systems capable of supporting complex diagnostic decision-making. Although AI has advanced considerably in cardiovascular diagnosis, existing evidence remains fragmented across algorithms, data modalities, and isolated application domains, limiting a comprehensive understanding of clinically integrated AI systems. This study presents a PRISMA 2020-guided systematic review and proposes a clinically grounded six-layer taxonomy that organizes cardiovascular AI according to diagnostic objectives, data modalities, modeling paradigms, data integration complexity, interpretability and trustworthiness, and deployment maturity. A systematic search of PubMed, Scopus, Web of Science, IEEE Xplore, and ScienceDirect identified 226 records, of which 76 primary empirical studies met the predefined eligibility criteria and were included in the comparative evidence synthesis. The review demonstrates the evolution of cardiovascular AI from conventional machine learning applied to structured clinical data toward deep learning for physiological signals and medical imaging, followed by multimodal AI systems integrating heterogeneous clinical information. Comparative synthesis across the proposed taxonomy highlights substantial progress in predictive performance while revealing persistent challenges related to external validation, dataset representativeness, workflow integration, explainability, privacy, governance, and prospective clinical deployment. The review further distinguishes clinically validated technologies from emerging paradigms, including federated learning, foundation models, and agentic AI. Overall, the proposed taxonomy provides a unified framework for organizing contemporary cardiovascular AI research and offers a practical roadmap for evaluating the maturity, trustworthiness, and clinical readiness of next-generation intelligent diagnostic systems.
We propose a multi-channel ExG recording system integrated with body channel communication (BCC), targeting wearable multimodal human-machine interfaces, such as extended reality (XR) and brain-computer interfaces (BCIs). To suppress noise from the BCC transmitter that appears as common-mode interference (CMI) at the ExG analog front end (AFE), we use time-multiplexed acquisition scheme that allocates separate time slots for ExG recording and BCC transmission. This temporal separation preserves ultra-low-noise ExG acquisition, enabling reliable recording of EEG and EOG in addition to EMG and ECG. Because the BCC transmitter must transmit multi-channel ExG data within a limited slot, it requires a high data rate. Using PAM-4 signaling, the proposed system achieves a 10-Mbps data rate. This modulation scheme is enabled by a bias-electrode-free ExG AFE, which increases the BCC signal amplitude by approximately 2.13×, thereby providing sufficient amplitude margin for reliable PAM-4 level discrimination. In addition, a charge-pump-based CMI cancellation loop incorporating a multi-channel least-mean-square (LMS) filter, together with a DC servo loop, mitigates inter-channel mismatch and suppresses differential-mode artifacts. Fabricated in 180-nm CMOS, the chips achieve a total common-mode rejection ratio (CMRR) of 101.1 dB and suppress electrode DC offsets up to 500 mV, while maintaining stable operation under 18-Vpp common-mode interference and 15% electrode mismatch.
This paper studies the problem of text-video retrieval, where the goal is to learn accurate cross-modal alignment between videos and text. This problem is challenging because of the matching ambiguity caused by the inherent gap between the heterogeneous video and text modalities. In particular, the differences in the information granularity and abstraction levels between the two modalities hinder a reliable sample-level alignment. Moreover, redundant visual content, sparse textual descriptions, and temporal variability in videos introduce additional uncertainty, resulting in ambiguous matching and suboptimal performance. In this paper, we propose a novel method named Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR), which models video-text pairs as probability distributions and captures uncertainty through the evidence theory. Specifically, we perform distribution-level representation learning to resolve the semantic ambiguity of video-text pairs. To improve the alignment further, we introduce a distribution-based embedding refinement module to ameliorate the semantic consistency across modalities. The proposed PE2LR is able to pull positive sample pairs closer in the embedding space, while pushing the negative pairs apart. Comprehensive experiments on several benchmark datasets (including MSRVTT, DiDeMo, and ActivityNet Captions) demonstrate that our PE2LR achieves state-of-the-art search performance. The code will be made publicly available upon acceptance of the paper.
We investigated the effect of posture on the link between cerebral circulation and cortical activity without applying cognitive tasks. We computed the zero-lag mutual information (MI) between spontaneous variations of mean cerebral blood velocity (MCBv) and the series of spectral powers computed over electroencephalographic (EEG) channels in traditional frequency bands. Estimation of MI was performed according to a fully linear approach and to a technique able to describe nonlinear components of the relationship as well. Time-shifting surrogate approach was utilized to reject the null hypothesis of uncoupling. Analysis was carried out in 27 healthy young individuals (age: 33±8 yrs; 13 males; 14 females) at rest in supine position (REST) and during active standing (STAND). Percentage of rejection of the null hypothesis of uncoupling peaked to 52% and did not vary across the MI estimates and experimental conditions. STAND did not affect MI regardless of the method utilized to its estimate. Results were consistent across brain areas and EEG frequency bands. In healthy young subjects in the absence of task-related activity, neurovascular coupling (NVC) can be assessed from spontaneous fluctuations of MCBv and EEG spectral powers, posture is not a confounding factor, and nonlinear components negligibly contribute to the information exchange between MCBv and resting-state brain activity. The significant values of MI suggest that NVC can be assessed from spontaneous variability of MCBv and EEG spectral power series and was not affected by orthostatic position.
Clustering aims to uncover heterogeneous features within data samples and partition them into meaningful groups. This article first establishes a theoretical connection between biorthogonal nonnegative matrix factorization (Bi-ONMF) and biorthogonal nonnegative tensor factorization (Bi-ONTF) under the t-product framework. Building on this foundation, we propose a novel Bi-ONTF model integrated with anchor graph learning (BNTF-AGL) for multiview clustering (MVC). The model employs sparse embedding learning to select anchor points, thereby eliminating redundant connections across multiview anchor graphs. To effectively capture complementary information among views, we introduce the tensor Schatten $p$ -norm as an approximation of the tensor tubal rank, which promotes consistency among cluster assignments across views. An adaptive augmented Lagrangian method (ALM) method is developed to optimize the proposed model, and we establish that every accumulation point of the resulting iterative sequence is a stationary Karush-Kuhn-Tucker (KKT) point under stated assumptions. Extensive experiments on nine real-world datasets demonstrate that the proposed method achieves competitive or superior clustering performance.
Semantic segmentation is critical for intelligent robotics to understand complex environments. While CNN-based models on RGB images achieve high performance, their accuracy drops in fast-motion or low-light scenes. Fortunately, event cameras, with high temporal resolution and low latency, offer robust perception in such challenging conditions. Many event-image fusion methods attempt to combine the complementary strengths of both modalities, but most adopt simple fusion strategies without considering intermodal correlations or designing computationally expensive architectures, resulting in degraded accuracy and high energy costs. To overcome these limitations, we propose a lightweight spiking neural network (SNN)-based event-image fusion network (Spike-EIFNet) that leverages the complementary strengths of multimodal fusion and energy-efficient spike-driven computation. In particular, to reduce computation cost for lightweight, Spike-EIFNet adopts a dual-branch SNN encoder to process events and images in parallel. Then, to improve the segmentation accuracy with enhanced feature interaction, we introduce a spike-driven cross-modal fusion (SCMF) module, consisting of a modality-aware fine-grained extraction (MFE) stage to capture dynamic cues from events and spatial details from images, followed by a cross-modal interaction and fusion (CIF) stage for effective feature alignment. Finally, a lightweight feature enhancement (LFE) module is proposed to further refine feature representations and facilitate deep-shallow feature fusion. Extensive experiments demonstrate that Spike-EIFNet achieves 67.34% and 58.09% mean intersection over union (mIoU) on the DDD17 and DSEC-Semantic datasets while consuming $72.83\times $ and $100.26\times $ less energy, respectively. Compared with ANN-based methods, Spike-EIFNet significantly reduces energy consumption; among SNN-based methods, it achieves the highest segmentation accuracy with a favorable accuracy-efficiency tradeoff. Code is available at: https://github.com/Chensyfighting/Spike-EIFNet.
Three-dimensional data are now central to computer vision, robotics, autonomous driving, medical imaging, augmented reality, and computer graphics. Yet, 3-D deep learning is difficult because geometric data are often large, irregular, and mostly empty. Dense voxel convolutional neural networks (CNNs) provide a simple extension of image CNNs to 3-D, but their memory and computation grow cubically with resolution. Sparse 3-D CNNs address this problem by storing and processing only the active parts of space, such as surface voxels or occupied cells in a point cloud. This article gives an accessible introduction to voxel-based CNNs, sparse convolution, and the data structures that make sparse 3-D networks practical, with special attention to octrees, hash tables, and the practical lessons and pitfalls that matter in real systems.
Data videos are an increasingly popular storytelling medium, effectively communicating data-driven insights to diverse audiences. However, designing compelling data videos requires not only domain knowledge of the data but also expertise in cinematic design, posing substantial challenges for non-experts in design. To study how AI can support non-experts in addressing such challenges during early-stage design, we introduce ideate-through-cocreate, a design concept in which users' ideation is scaffolded through co-creation with generative AI. We instantiate this concept by proposing Data Video Designer (DVD), a proof-of-concept system that supports the ideation of data video opening scenes. DVD facilitates non-experts' ideation processes by generating exemplary video scripts and visuals from high-level topic and data summaries, referencing relevant design guidelines, and integrating accurate visualizations derived from their own data into the generated scenes. Through this workflow, DVD turns abstract cinematic guidelines and AI-generated materials into actionable ideation support. A workshop and a comparative study demonstrated the usefulness of the structured co-creation support provided by DVD within the scope of opening-scene design and revealed the creative potential in balancing the data authenticity and stylized expression. These findings provide initial evidence for the potential of structured human-AI co-creation to support data video ideation.
Ultrasound imaging, with its advantages of portability, real-time feedback, and non-invasiveness, has become an indispensable modality in computer-aided diagnosis. However, inherent challenges in ultrasound images, such as speckle noise, low contrast, and blurry boundaries, significantly hinder segmentation accuracy. Existing methods attempt to alleviate these issues by deepening network architectures, enlarging convolutional kernels, or introducing self-attention mechanisms to enhance feature representation. Nevertheless, these strategies often compromise computational efficiency, making it difficult to achieve a desirable balance between segmentation accuracy and inference speed, thereby limiting their applicability in resource-constrained clinical scenarios such as portable ultrasound devices. To address these challenges, this paper proposes a lightweight and efficient segmentation network, termed LS2Net. In the encoder, to mitigate the impact of speckle noise, LS2Net incorporates wavelet convolution to transform spatial features into the frequency domain for targeted processing, which not only suppresses noise but also enlarges the receptive field and captures global contextual information. Meanwhile, pixel difference convolution is employed to capture fine-grained details under low-contrast conditions, enhancing the perception of local textures. Furthermore, we design a Residual Refinement Skip Connection (RRSC) module, which leverages the discrepancy between down-sampled and up-sampled features to preserve informative components while filtering redundant information, thereby facilitating more effective feature reconstruction. In the decoder, a Multi-Receptive Field Coarse-to-Fine Module (MRCFM) is introduced to further integrate multi-scale contextual information through hybrid receptive fields, enabling more precise delineation of object boundaries. LS2Net demonstrates superior computational efficiency and state-of-the-art (SOTA) performance across five ultrasound datasets, achieving a favorable balance between segmentation accuracy and inference speed. In addition, it delivers competitive performance on other imaging modalities, highlighting its versatility. Moreover, LS2Net exhibits strong generalization capability in cross-dataset evaluations, validating its robustness and scalability. Overall, LS2Net provides a practical and deployable solution for real-time clinical applications, particularly suitable for portable or resource-limited devices. The source code is publicly available at: https://github.com/CYYJL/LS2Net.
Chronic conditions such as cardiovascular disease, diabetes, and cancer require sustained lifestyle changes and self-management, yet traditional care models often provide limited support for long-term behavior change. Digital health technologies, particularly virtual agents, computer-generated characters simulating human-like interactions through verbal and nonverbal cues, offer new ways to provide personalized, scalable, and continuous support. However, the ways in which distinct components within such digital health technologies, including those used in chronic care interventions, are chosen and combined remain underreported. We conducted a systematic scoping review to map how behavior change techniques (BCTs), health data types, and delivery channels are rationalized, combined, and applied in virtual agent-delivered interventions for chronic condition management. The review followed established scoping review frameworks and adhered to PRISMA-ScR reporting guidelines. A search was performed across PubMed, Scopus, PsycInfo, WebofScience, and IEEE Xplore in September 2024. Twenty-one studies met the inclusion criteria. We examined the rationales reported by authors for intervention design, categorized as theory-driven, practice-driven, empirically-driven, mixed, or not explicitly stated. Few studies explained why they selected specific techniques or how health data and delivery channels were intended to interact. Across studies, BCTs were identified but often not explicitly labelled. The most common agent-delivered techniques were self-monitoring, feedback, instruction on how to perform a behavior, and prompts and cues. These techniques were typically supported by subjective self-reports (e.g., symptoms, behaviors), objective data (e.g., step counts, blood pressure), adherence data (e.g., activity completion) and user preference data (e.g., preferred timing of reminders). Delivery channels comprised smartphone or tablet apps. This review provides the first systematic map of how BCT-health data-delivery channel combinations are applied in virtual agent interventions for chronic condition management. It highlights foundational design patterns and reporting gaps, emphasizing the need for transparent, theory-informed reporting to guide future development of adaptive, evidence-based digital health tools.
Cyclic peptide drugs show great potential in antiviral, antibacterial, anticancer, and immunomodulatory therapies, yet accurate prediction of their membrane permeability remains challenging. Existing approaches based on SMILES, molecular graphs, or 3D structures have inherent limitations: SMILES lack spatial information, graphs inadequately capture stereochemistry, and 3D methods are sensitive to conformational variability. Moreover, current multimodal fusion strategies often fail to effectively integrate heterogeneous molecular information. To address these challenges, we propose MultiMol, a multimodal framework that integrates SMILES sequences, molecular images, molecular graphs, and 3D conformations through tailored pre-training tasks and a scalable fusion mechanism. Experiments show that MultiMol consistently outperforms existing methods in cyclic peptide permeability prediction. Visualization and interpretability analyses further demonstrate its strong feature extraction and generalization capabilities. MultiMol also prioritizes promising KRAS-targeting cyclic peptides, supporting its practical utility in virtual screening. The code is available at https://github.com/chaoxiuxiu/multi-mol.
Extended reality (XR) can be a valuable tool for knowledge workers. This idea is reflected in the advertisements for current XR devices, which promote them as productivity tools. However, a future work culture in which XR plays a significant role involves a variety of research areas, many of which are arguably not well explored yet. Through expert discussions, we identified 10 challenges in this emerging field, which are presented and discussed here. These challenges cover a diverse set of topics ranging from technological aspects, such as hardware or interaction, to content design and collaboration. They also include more human centric-topics concerned with work-life boundaries, health effects, social aspects, and accessibility. We also address challenges in industry and regarding privacy and security. Therefore, this article contributes an overview of research directions and open challenges for realizing the potential of XR for the future of work.
Federated semi-supervised learning (FSSL) for medical image segmentation has been extensively studied in recent years. Due to the requirement of specialized knowledge and equipment for annotating medical data, only a very limited number of medical institutions have a small amount of labeled data. However, existing federated semi-supervised segmentation methods primarily focus on fully supervised clients to improve average performance and often overlook the contributions of unsupervised clients. To effectively leverage unsupervised clients and extract more valuable information, we propose pseudo-global based sequential contribution estimation for federated semi-supervised segmentation, abbreviated as FedPSC. FedPSC first estimates the performance contributions of all clients using the difference in validation performance, and then builds a pseudo-global model based on this. Sequentially, the pseudo-global model and the union generated by the exclusion effect are used to estimate the gradient contribution of the client, thereby establishing a relationship between the two contributions. Besides, we also introduce a gradient direction exponential moving average method, which aims to train client models by integrating both global general knowledge and local personalized insights. Experimental results on two commonly used datasets confirm the effectiveness of our proposed method. Additional experiments and analysis are also provided to give in-depth understanding of FedPSC.
Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.
Existing deep reinforcement learning (DRL) methods for autonomous underwater vehicle (AUV) path planning face two practical challenges: 1) dependency on manual reward engineering and 2) hyperparameter sensitivity in dynamic marine environments. This article presents a novel AUV path planning framework incorporating generative adversarial imitation learning (GAIL) and DRL algorithm that automates reward function synthesis through adversarial learning from expert demonstrations. The proposed architecture introduces a hierarchical reward mechanism that concurrently optimizes global trajectory planning and local motion constraints. By eliminating manual reward engineering, our approach reduces training complexity while maintaining policy convergence stability. Extensive experimentation demonstrates superior performance with 93.7% faster training convergence and 72.7% higher path convergence optimality compared to conventional DRL baselines. Two-tier validation confirms operational effectiveness: 1) Gazebo simulations achieve maximum 100% success rate in dynamic scenarios and 2) field deployments for submarine pipeline inspection attain 1 m average tracking accuracy. The results demonstrate that GAIL-DRL trained AUVs exhibit enhanced path planning stability while satisfying real-time planning requirements for marine transportation systems.
Multi-modal magnetic resonance imaging (MRI) plays a crucial role in brain tumor diagnosis. However, the substantial physiological sensitivity discrepancies across imaging modalities create a severe domain gap that challenges current cross-modality segmentation methods. Many unsupervised do main adaptation (UDA) approaches reduce this gap through image style translation or distribution alignment, which have shown promising results but may struggle to preserve modality specific anatomical and pathological information. In contrast, few-shot segmentation (FSS) provides a promising alternative paradigm with strong cross-domain generalization capability, avoiding explicit domain alignment. Therefore, in this work, we propose a novel Structure-aware Attention Prototype Network (SAPNet) for cross-modality few-shot brain tumor segmentation, which fully exploits task-agnostic, multi-scale features from the large vision model. Specifically, our SAPNet consists of three main components. First, we introduce a masked support feature reconstruction (MSFR) module to encourage the network to understand anatomically meaningful structures. Second, a patch-level attention-to-prototype alignment (APA) module is proposed to adaptively balance the aggressive and conservative segmentation tendencies by cross-attention (CA) and prototype based learning. Third, a lightweight multi-scale decoder with the contrastive embedding space is employed to enhance fine grained pixel-level prediction. Extensive experiments on three public benchmarks, BraTS 2020, VS-SEG, and BraTS 2023 PEDdatasets, demonstrate that SAPNet consistently outperforms state-of-the-art UDA and FSS methods, exhibiting strong gener alization and robustness.