Peptide-based vaccines, enabled by bioinformatics and machine learning (ML), have emerged as one of the most promising approaches for rapid, safe, and cost-effective vaccine design against infectious diseases. Unlike conventional approaches that depend heavily on whole-pathogen cultures or recombinant protein expression, peptide vaccines can be designed in silico and synthesized quickly. Rational and targeted in silico approaches for the discovery of peptide-based vaccine candidates include B-cell and T-cell epitope prediction, immunogenicity, antigenicity, allergenicity, autoimmunity, population coverage, sequence conservation, molecular docking, molecular dynamics simulation, in silico cloning, and immunological simulation analyses. The combination of these comprehensive computational methods can effectively generate high-quality vaccine candidates for subsequent validation via in vitro and in vivo experiments. This review contextualizes the historical trajectory of peptide-based vaccinology, from early linear epitope discoveries in the 1960s to multi-epitope constructs and clinically tested candidates such as UB-612 and PepGNP-Covid19. It examines critical challenges in immunoinformatics, including performance gaps in epitope prediction tools, complexities in human leucocyte antigen (HLA) mapping, and the need for extensive manual intervention in pipelines. Artificial intelligence-driven approaches, spanning deep learning, and interpretable ML, are positioned to transform epitope prediction, reduce human error, and standardize reproducibility. These advances have the potential to support global outbreak response targets such as the Coalition for Epidemic Preparedness Innovations (CEPI) 100 Days Mission and the World Health Organization (WHO) Research and Development (R&D) Blueprint. However, their performance remains constrained by data quality, dataset imbalance, limited benchmark standardization, and persistent underrepresentation of many HLA alleles and population groups. Key Points Peptide-based vaccines, accelerated by bioinformatics and machine learning, offer a potentially rapid, relatively safe, and cost-effective alternative to traditional vaccine design, enabling in silico development and swift synthetic manufacturing. Computational methods such as B-cell and T-cell epitope prediction, immunogenicity analysis, and molecular simulations allow for rational and targeted vaccine candidate discovery, enhancing quality and efficiency. The field has evolved from early linear epitope discoveries in the 1960s to sophisticated multi-epitope constructs and clinically tested candidates like UB-612 and PepGNP-Covid19. Major challenges in immunoinformatics include performance limitations in epitope prediction tools, complexities in HLA mapping, and the necessity for manual intervention in data pipelines. Artificial intelligence-driven models, including deep learning and interpretable machine learning, promise to overcome these challenges by improving prediction accuracy, reducing errors, and supporting global epidemic response efforts such as CEPI's 100 Days Mission and the WHO R&D Blueprint.
The European Union (EU) enforces strict regulations on the traceability and labeling of genetically modified organisms (GMOs), including genome-edited (GE) lines produced through new genomic techniques (NGTs). Identifying GE organisms created by single nucleotide variations (SNVs) is however challenging, as a single SNV alone cannot unambiguously define a GE line. Recently, we introduced the concept of generating a genetic fingerprint to distinguish a specific GE rice line. This proof-of-concept approach integrated whole-genome sequencing (WGS)-based characterization with the Illumina technology, the public 3 K Rice Genomes (3KRG) database, and statistical feature-selection tools, to select and combine key genetic elements, including GE on-target site(s) and cultivar-specific 2-SNV barcodes, into a unique genetic fingerprint. In the present study, we expand this concept into a generalized data-driven framework allowing identification of multiple rice lines. Supported by newly developed bioinformatics and statistical feature-selection-based pipelines, this optimized strategy enables the generation of genetic fingerprints irrespective of a rice cultivar's inclusion in publicly available databases like 3KRG. In addition, this refined strategy can leverage WGS data generated from both Illumina and Oxford Nanopore Technologies (ONT) platforms for fingerprint generation and GE line identification. Using two distinct in-house GE rice lines from different cultivars, along with various publicly available WGS datasets, we demonstrated the robustness, scalability, and specificity of this approach for reliable GE rice line identification. Our findings provide a methodological foundation for data-driven traceability of GE rice lines, reinforcing regulatory compliance, supporting intellectual property (IP) protection, and contributing to the responsible implementation of EU GMO/NGT legislation.
Gene network enrichment analysis (GNEA) offers a robust approach for interpreting the complex molecular mechanisms underlying phenotypic variability. Despite its utility, prevailing GNEA methodologies are predominantly optimized for binary phenotypes, leading to substantial information loss when applied to continuous biological traits such as drug sensitivity, cancer progression, etc. Additionally, traditional enrichment strategies often exhibit conceptual inconsistency between their null models and test hypotheses, as they rely on permuting phenotype labels rather than genes. To address these limitations, we introduce cell line-specific gene network enrichment analysis (CellGNEA), a computational strategy designed to identify pathway-level molecular interactions associated with continuous phenotypes in a cell line-specific manner. CellGNEA constructs gene regulatory networks tailored to individual cell lines and assesses molecular interplays within these networks by integrating multiple network-derived metrics, including clustering coefficient, PageRank, and regulatory effects. Associations between gene networks and continuous phenotypes are quantified using a Kolmogorov-Smirnov-based statistic, with statistical significance determined via a gene permutation strategy that aligns the null model with the tested hypothesis. Monte Carlo simulation studies indicate that CellGNEA exhibits robustness and enhanced sensitivity in detecting network enrichment linked to continuous phenotypes. Applications of CellGNEA to drug sensitivity-specific gene networks enable the identification of leukemia-related pathways with molecular interactions significantly associated with therapeutic response. Notably, our analysis revealed that imatinib, quizartinib, and ruxolitinib consistently correlate with network-level remodeling across acute myeloid leukemia, myelodysplastic syndrome, and chronic myeloid leukemia pathways, and identified resistance-associated genes including PTPN11, MS4A1, and BTK. Overall, CellGNEA establishes a systematic, scalable framework for functional network analysis of continuous phenotypes, facilitating comprehensive characterization of cell line-specific biological properties and providing valuable insights for systems biology and precision medicine.
Inter-sample tumor heterogeneity poses significant challenges to metastatic cancer treatment. Although multiomics analyses provide nuanced molecular insights into tumor heterogeneity, current molecular pathway analysis tools focus on group-based comparisons, which may overlook differential single-sample perturbations. Here, we present normalized single-sample single-omic pathway analysis (SOPA), and its extension, normalized single-sample integrated multiomics pathway analysis (SIMPA), as a bioinformatics pipeline for performing supervised differential pathway analysis. The pipeline utilizes custom algorithms to analyze differential pathway activity in single samples, comparing the molecular profile in each sample to a range of controls. In single -omics analysis, SOPA shows advantages compared to standard tools such as single sample gene set enrichment analysis and gene set variation analysis in identifying single sample deviations from predefined controls. For integrated multiomics, we show that in predefined-control contexts, SIMPA provides an effective alternative over unsupervised tools such as multiomics gene set analysis (MOGSA) and PAthway Deviation scores using Multiple Factor Analysis (padma), addressing tumor heterogeneity. Particularly, SIMPA unveiled particular tumor subgroups with dysregulated immune and metabolic pathway activity marked by variable immune infiltration and survival differences, which were missed by MOGSA and padma. Overall, SOPA and SIMPA are valuable in the supervised analysis of single samples and allow for the investigation of complex multiomics data to gain personalized hypothesis-generating insights. The flexibility of this pipeline allows implementation in preclinical and clinical research, offering significant advantages over prior pathway analysis tools in studying systems biology. The Python package for SOPA and SIMPA is freely accessible at https://github.com/hasanalsharoh/SIMPApy/.
Antibodies play central roles in immune defense and are widely used as therapeutic agents. However, the high structural and sequence diversity of antigen-binding loops, combined with limited experimental data and weak co-evolutionary signals, makes it difficult to develop generalizable predictive models. In this work, we investigate test-time fine-tuning strategies to improve protein language model (pLM) performance in low-data settings, with a focus on antibody-related tasks. Systematic evaluations across tasks show that carefully constrained fine-tuning greatly enhances performance while preserving generalization. In particular, depth-selective fine-tuning consistently outperforms full-depth fine-tuning, with optimal performance achieved when tuning 50%-75% of model layers for medium- to small-sized pLMs. We introduce AbTune, a test-time fine-tuning framework that leverages this depth-controlled adaptation strategy. Across antibody structure prediction, mutation effect prediction, and binding affinity prediction, AbTune outperforms both standard pLM baselines and task-specific predictors, achieving the best performance among the evaluated baselines on two of the three tasks. To gain insight into the adaptation process and identify optimal AbTune protocols, we analyzed representation shifts, examined how sequence properties influence fine-tuning dynamics, and evaluated metrics that capture potential overfitting. Our results show that fine-tuning depth, duration, and perplexity jointly influence performance and must be carefully controlled to achieve optimal results.
Chromatin organization shapes gene regulation by linking distal elements across megabase scales, yet most predictive genomics models still treat the genome as linear, without incorporating 3D structure. Hi-C provides genome-wide chromatin conformation information, but its contact maps are population-averaged, distance-biased, and noisy, obscuring biologically specific contacts. We present CHROME, a framework built on a self-avoiding polymer ensemble null model that identifies physically specific, nonrandom Hi-C contacts. By integrating these contacts into graph representations, CHROME enables efficient information transfer across spatially connected loci. It integrates sequence, chromatin accessibility, or pretrained embeddings into a graph attention architecture to predict cell-line-specific ChIP-seq profiles, improving performance over matched local encoder baselines. In a held-out cell line, CHROME demonstrates improved performance in selected settings, suggesting potential for cross-cell-type transfer. The resulting graph embeddings also enhance prediction on tissue-specific eQTL and ClinVar variant pathogenicity, compared with local sequence-based embeddings. Beyond predictive performance, CHROME provides interpretability through attention-derived neighbor-to-center contributions that reveal how spatially connected loci influence local regulatory activity over multi-megabase distances. Together, these results highlight the value of incorporating physically validated chromatin interactions for improving regulatory prediction and variant interpretation.
Binding of T-cell receptors (TCRs) and their cognate peptide-major histocompatibility complex (pMHC) target is determined by both TCR$\alpha $ and TCR$\beta $ chains. However, not all TCR$\alpha $ and TCR$\beta $ can bind to each other. Predicting their pairing is crucial for understanding the TCR-pMHC interaction and developing effective de novo TCRs. Here, we show that in the general TCR repertoire, TCR$\alpha $ and TCR$\beta $ chain compositions are independent. However, in pMHC-binding TCRs, clear associations between TCR$\alpha $ and TCR$\beta $ chains are found, also for TCRs binding to the same pMHC. The association between the CDR3 amino acid composition and $V$, $J$ usage of TCR$\alpha $ and TCR$\beta $ reveals distinct binding patterns between specific $V$ and $J$ genes, as well as negative correlations between the charge and polarity of the TCR$\alpha $ and TCR$\beta $ chains, but positive associations between their molecular weights. These associations are used for the development of a prediction model for TCR$\alpha $ and TCR$\beta $ pairing. We present here TCR-BARN (TCR Beta-Alpha chains paiRing using Nlp) that employs an initial embedding for each amino acid in the TCR alpha and beta CDR3 sequences, followed by long short-term memory (LSTM) networks to capture sequence dependencies. The $V$ and $J$ genes are represented using one-hot encoding. LSTM outputs are concatenated and passed through a fully connected feedforward layer for binding prediction. TCR-BARN reaches an area under the curve $>0.65\pm 0.007$ for epitope-bound TCRs. TCR-BARN can be used for generating cognate TCRs resembling natural TCRs and evaluating the generated TCR quality.
Accurate identification of ultra-low-frequency tumor-derived variants is critical for circulating tumor DNA (ctDNA)-based minimal residual disease (MRD) profiling. However, current ctDNA analysis workflows largely operate as predefined, sequential pipelines without explicit mechanisms for continuous monitoring of variant-calling performance or adaptive control, thereby limiting detection stability in genomically heterogeneous regions. To address this limitation, we developed MRDsteer, an autonomous closed-loop agent driven by artificial intelligence (AI). MRDsteer monitors the analytical reliability of variant calling during the analysis process using multidimensional quality metrics, such as filtration ratio and strand bias. When the estimated reliability falls below an actionable threshold, MRDsteer triggers localized re-calling only in high-risk genomic regions, instead of repeating the entire analysis. In this way, MRDsteer provides closed-loop control by continuously assessing variant-calling reliability and applying targeted intervention when needed. Comparative analyses in simulated and real-world datasets showed that MRDsteer improved the stability and sensitivity of ctDNA variant detection. Under challenging conditions, including ultra-low variant allele frequencies, MRDsteer demonstrated improved detection performance compared with representative baseline methods. In clinical cohorts, MRDsteer improved ctDNA-based MRD stratification and strengthened progression-free survival separation in the K438 cohort, including both non-small cell lung cancer (NSCLC) and nasopharyngeal carcinoma (NPC) subgroups. These results suggest that MRDsteer may provide a robust and clinically useful computational strategy for sensitive MRD detection and longitudinal ctDNA monitoring. https://github.com/aAT0047/MRDsteer.git.
Workflow Management Systems (WMSs) make it easier to run bioinformatics analyses by combining tools into reusable workflows. However, selecting the right WMS can be challenging, particularly for users from noncomputational backgrounds. Many reviews of WMSs focus on technical details or only consider usability from the perspective of developers. The first objective of this scoping review was to identify the characteristics of bioinformatics WMSs that are most relevant for life scientists. The second objective was to explore how these characteristics have previously been evaluated. The review included 21 papers published since 2018 that evaluated bioinformatics WMSs based on criteria relating to the user experience. Published papers and websites describing 55 currently available WMSs were also included to identify characteristics highlighted by their developers. Twelve themes emerged from these evaluation criteria and WMS characteristics: Basic Computing, Functions, Security, Scalability, Cost/Efficiency, Sustainability, Usability, Learnability, Reproducibility, FAIRness (Findability, Accessibility, Interoperability, Reusability), Flexibility, and Support. Evaluations focusing on the needs of noncomputational users preferred Graphical User Interfaces and platforms that provided plenty of guidance for users. Papers prioritizing the needs of developers instead favoured text-based interfaces and flexible platforms that gave users greater control. In addition to these contrasting views on what was considered a positive characteristic, differences in how criteria were defined and scored meant that evaluations could not be compared between papers and would be impossible for users to repeat on new or updated WMSs. Users do not currently have a clear approach to follow when selecting a WMS.
Modern bioinformatics faces escalating challenges stemming from both the inherent computational hardness of many fundamental problems and the rapidly growing scale and complexity of biological data, increasingly limiting the effectiveness of classical computational approaches. Quantum computing has emerged as a promising paradigm for addressing these challenges by enabling alternative problem representations and novel search strategies for exploring complex solution spaces. This systematic review provides a structured overview of the emerging field of quantum bioinformatics and aims to supplement recent reviews on this topic in the journal by providing an updated and structured synthesis of current research. We systematically collect and organize existing studies across 10 bioinformatics domains to identify research trends, dominant themes, and recurring methodological patterns. The review examines quantum and hybrid quantum-classical approaches, problem formulations, and encoding strategies, with particular attention to the constraints of noisy intermediate-scale quantum devices, including noise, limited scalability, and the need for error mitigation. We further synthesize reported limitations, open challenges, and prospective research directions.
Single-cell RNA sequencing has emerged as a transformative tool, enabling precise phenotype prediction and the detailed identification of disease-associated cell subpopulations. However, many existing computational approaches still rely on predefined cell-type annotations during model training. This dependence makes their predictive performance highly sensitive to subjective annotation quality, labeling inconsistencies, and dataset-specific biases, ultimately hindering their generalizability across diverse patient cohorts. To address these challenges, we propose scCap, an annotation-free framework that leverages knowledge-augmented clustering for robust phenotype prediction. Specifically, the framework first constructs initial clusters from raw gene expression profiles and subsequently refines them within the embedding space of a pretrained single-cell foundation model, allowing the clusters to better reflect broader biological organization while preserving fine-grained cellular heterogeneity. The resulting knowledge-augmented clusters are then integrated into a hierarchical multiple instance learning framework with dual-level attention, enabling interpretable predictions at both the cell and cluster levels. Evaluated across three public scRNA-seq datasets, scCap consistently outperforms baseline models in predictive accuracy. Furthermore, scCap identifies disease-associated subpopulations previously reported in the literature without relying on predefined cell-type annotations. These results demonstrate that scCap provides a robust and interpretable framework for annotation-free phenotype prediction.
Adverse drug reactions caused by molecules binding unintended targets are a major concern in drug discovery. Early identification of such interactions during drug design and development minimizes risks and enhances therapeutic efficacy. While experimental approaches are time-consuming and resource-intensive, in silico virtual screening offers a faster, cost-effective strategy to anticipate off-target effects early in drug design. Here, we present a point cloud-based virtual screening method, designed to perform off-target identification based solely on the chemical composition of the primary drug binding site. By screening multidimensional point clouds representing all potential human binding sites, this approach identifies alternative targets based on shape and physicochemical properties. Notably, it operates independently of the protein's overall structure or sequence. This focus on the binding site broadens the search space to structurally unrelated proteins and enables screening without requiring lead molecule information. Using an experimentally validated test set, we demonstrated the method's ability to identify alternative targets across protein families and predict drug promiscuity, achieving a Top-10 recall of 24% for validated, strongly modulated targets. While ligand-agnostic sequence- (BLASTP) and structure-based (Foldseek) methods show higher overall recall, point cloud-based screening uniquely recovers 13.9% (79.4% of its total recovered hits) of known off-targets with <30% sequence identity at Top-100, which are missed by both BLASTP and Foldseek. Beyond known targets, we found several high-ranking candidates not yet annotated as drug-targets but showing notable cavity similarity despite being structurally unrelated, which we present as testable hypotheses for follow-up.
The emergence of high-throughput sequencing technologies has generated unprecedented amounts of molecular data, posing significant challenges for analysis and interpretation. In the past 15 years, deep learning (DL) has revolutionized many data analysis fields. However, a key limitation of most DL approaches is their reliance on massive amounts of data that often require labeling. Self-supervised learning (SSL) uses large-scale unlabeled data to learn meaningful representations for specific tasks, allowing training of the models without labels. While SSL has primarily been applied to fields like natural language processing and medical image analysis, SSL can also be applied to molecular data to learn meaningful representations of molecular sequences that can be used for downstream tasks. Despite the growing interest in SSL applications using molecular data over the recent years, no comprehensive review has been published on this topic, so far, making it timely to address. This paper aims to provide researchers in DL and bioinformatics with a clear view of SSL omics applications to foster future work in the domain. This review examines the principles of SSL, such as foundation models, and it discusses the application of SSL to various omics data types, summarizes information from 17 studies, and categorizes applications by data type, detailing common tasks, model architectures, and repositories. Key applications such as DNABERT and Nucleotide Transformer are highlighted, demonstrating the contributions of SSL in understanding gene regulation. Future directions for SSL in omics are outlined, emphasizing the potential for integrating multi-omics data and developing more sophisticated pretext tasks.
Single-cell RNA sequencing (scRNA-seq) has emerged as a transformative technology for decoding cellular heterogeneity and state diversity within complex tissues through high-throughput transcriptomic profiling of individual cells. scRNA-seq clustering is a critical task for analyzing scRNA-seq data, which resolves high-dimensional expression profiles into interpretable cellular types and states. Although the Transformer, as a powerful foundation model, offers strong representation learning capabilities, its use in single-cell analysis is limited by the absence of a biologically meaningful and computationally scalable tokenization mechanism. Existing methods typically construct tokens through similarity-based neighbor selection, a process that is highly sensitive to metric quality and incurs substantial computational overhead, limiting applicability to large-scale datasets. Here, we introduce a hash-driven tokenization mechanism, scHashFormer, in which we design a novel hash encoder with a learnable hash window size and train it using self-supervised learning to realize similar cells with the same hash codes. Identical hash codes define hash buckets from which token sequences are constructed. By aggregating information from the constructed sequence, similar cells are brought closer in the embedding space, ensuring more effective clustering. Extensive experiments on multiple scRNA-seq datasets demonstrate that scHashFormer achieves competitive clustering effectiveness and scalability. The resulting embeddings enhance performance in downstream tasks, including trajectory preservation and gene differential expression analysis.
Pancreatic ductal adenocarcinoma (PDAC) is a highly lethal malignancy with limited therapeutic options. In this study, we introduce a target-based deep learning framework to investigate the anticancer activity of eravacycline (Erav), a United States Food and Drug Administration (FDA)-approved antibacterial agent previously identified in our work as a potential anticancer candidate through computational screening. We developed a novel two-phase in silico yeast-based prediction model to explore potential mechanisms of action, followed by in vitro and in vivo experimental validation. DNA polymerase kappa (POLK) and mutant p53 emerged as the top-ranked candidate targets. In the studied mutant p53 PDAC model, Erav treatment significantly reduced mutant p53 protein levels and was associated with marked downregulation of POLK protein expression. POLK is a previously underexplored DNA polymerase that has been reported to be overexpressed in multiple cancer types. In a subcutaneous xenograft model, Erav treatment resulted in a 76% reduction in tumor volume. Our findings demonstrate an association between Erav treatment and reduced POLK protein expression in the studied mutant p53 PDAC model, supporting POLK as a prioritized candidate for further investigation and providing preliminary mechanistic insight into Erav activity. This integrative computational-experimental pipeline offers a robust strategy for accelerating drug repurposing in oncology.
Graph-based deep learning has emerged as a powerful framework for modeling drug-target interactions (DTIs), enabling the integration of molecular, structural, and systems-level information within a unified representation. In this review, we survey graph-based DTI models across biomedical network-, sequence/hybrid-, and structure-based paradigms, which form a continuum from large-scale association inference to structure-resolved interaction modeling with increasing mechanistic specificity. Beyond architectural advances, we introduce an output-driven perspective in which models are evaluated according to how well their predictions align with the informational and decision-making requirements of different stages of the drug discovery pipeline. Within this framework, attention mechanisms and semi-supervised learning are discussed as key developments that enhance feature prioritization and data efficiency in data-limited settings. We further examine how model outputs support applications ranging from target identification and drug repurposing to structure-guided lead optimization. Finally, we analyze key benchmarking challenges, including data leakage, sequence redundancy, and structural bias, and discuss emerging directions such as multimodal integration and the use of predicted protein structures. Together, this review provides a unified perspective on the design, evaluation, and translational application of graph-based DTI models.
Single-cell spatial multi-omics technologies enable the simultaneous acquisition of multimodal molecular profiles and spatial location information in situ, providing a novel perspective for spatial domain identification and functional characterization of tissues. However, existing methods still suffer from several limitations, including insufficient denoising capability for single-cell data, reliance on static graph structures, and inadequate exploitation of the complementary relationships between spatial information and molecular features. To address these challenges, we propose AGCLD, an adaptive graph contrastive learning method with denoising for spatial domain identification. Specifically, a modality-specific denoising variational autoencoder is first employed to learn robust latent representations, thereby effectively mitigating noise interference. A differentiable graph generator is then introduced to adaptively construct spatial adjacency graphs and expression similarity graphs, alleviating the bias introduced by fixed neighborhood assumptions. Finally, AGCLD utilizes a dual-graph attention network to encode the spatial adjacency and expression similarity graphs, yielding spatial and feature embeddings, and incorporates a contrastive learning mechanism to align the dual-view representations, thereby enhancing representation consistency. Extensive experiments on five spatial multi-omics datasets demonstrate that AGCLD outperforms state-of-the-art methods, including SpatialGlue, in spatial domain identification tasks.
Variant-set association analysis is a powerful strategy for genetic studies of whole-genome sequence (WGS) data, especially for rare variants. By aggregating variant signals, variant-set analysis can improve statistical power, result interpretability, and study replicability. Motivated by the evidence that 3D genome architecture plays a critical role in regulating gene transcription, several works have incorporated 3D genome architecture into gene-based association tests and demonstrated great promise. In this work, we extend the idea of 3D-genome guided test from gene-centric to gene-agnostic, whole-genome testing by introducing an Hi-C informed kernel association test (i.e. HiC-KAT). We present a principled procedure that converts Hi-C contact confidence into borrowing weights and integrates these weights into genetic similarity kernels so that higher-confidence interacting loci contribute more to the association test of the target variant set. We use a controlling parameter to adaptively determine the appropriate degree of information borrowing from its interacting loci during association testing. We assess the performance of HiC-KAT using simulations and illustrate its advantage in detecting rare-variant sets using WGS data from the ARIC study in the Trans-Omics for Precision Medicine program.
The exponential growth of biomedical literature creates a cognitive bottleneck in drug target discovery, particularly for identifying therapeutically relevant mechanisms beyond established pathways. In ocular neovascularization, anti-VEGF therapies are standard of care, yet non-response and resistance remain critical unmet needs. We present an integrated computational framework combining Graph Retrieval-Augmented Generation (GraphRAG)-based literature mining, pathway co-localization analysis, and deep learning-based druggability assessment for systematic target prioritization. Using 5562 angiogenesis-related PubMed abstracts, we constructed a vascular knowledge graph (17 842 nodes; 9555 edges) and applied pathway co-localization with vascular endothelial growth factor A (VEGF-A) as a biological filter. As a validation step, the workflow recovered four targets-fibroblast growth factor 2, transforming growth factor-beta 1, interleukin-1 beta, and matrix metalloproteinase-9-already supported by clinical or advanced preclinical development, demonstrating concordance with expert-driven selection. Iterative querying subsequently identified two additional mechanistically supported candidates, fibroblast growth factor 1 and hepatocyte growth factor, sharing receptor tyrosine kinase-centered pathways with VEGF-A but lacking clinical evaluation in ocular neovascularization. Deep learning-based structural analysis (DeepSite and PocketMiner) identified high-confidence ligandable pockets for all six candidates. This work demonstrates how GraphRAG can systematically mine existing literature to recover known targets and surface literature-supported candidates that may be underprioritized for translational development. Rather than claiming de novo discovery, we emphasize the framework's utility as a scalable, transparent, and reproducible methodology for overcoming citation bias and literature overload. The workflow is generalizable to other complex, literature-rich disease domains.
Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.