Distinguishing pathogenic mutations from benign polymorphisms remains a critical challenge in precision medicine. EnTao-GPM, developed by Fudan University and BioMap, addresses this through three innovations: (1) Cross-species targeted pre-training on disease-relevant mammalian genomes (human, pig, mouse), leveraging evolutionary conservation to enhance interpretation of pathogenic motifs, particularly in non-coding regions; (2) Germline mutation specialization via fine-tuning on ClinVar and HGMD, improving accuracy for both SNVs and non-SNVs; (3) Interpretable clinical framework integrating DNA sequence embeddings with LLM-based statistical explanations to provide actionable insights. Validated against ClinVar, EnTao-GPM demonstrates superior accuracy in mutation classification. It revolutionizes genetic testing by enabling faster, more accurate, and accessible interpretation for clinical diagnostics (e.g., variant assessment, risk identification, personalized treatment) and research, advancing personalized medicine.
Community science observational datasets are useful in epidemiology and ecology for modeling species distributions, but the heterogeneous nature of the data presents significant challenges for standardization, data quality assurance and control, and workflow management. In this paper, we present a data workflow for cleaning and harmonizing multiple community science datasets, which we implement in a case study using eBird, iNaturalist, GBIF, and other datasets to model the impact of highly pathogenic avian influenza in populations of birds in the subantarctic. We predict population sizes for several species where the demographics are not known, and we present novel estimates for potential mortality rates from HPAI for those species, based on a novel aggregated dataset of mortality rates in the subantarctic.
Highly pathogenic avian influenza (HPAI) has expanded its host range with recent detections in dairy cattle, raising critical concerns regarding within-herd persistence and cross-species spillover. This study develops a stochastic $SEI_sI_aR-B$ compartmental model to analyse HPAI transmission, explicitly accounting for environmental pathogen reservoirs and noise intensities through Wiener processes. The positivity and boundedness of solutions are established, and the disease-free and endemic equilibria are analytically derived. The basic reproduction number is determined using the next-generation matrix method. Numerical simulations confirm that the model dynamics are consistent with theoretical analysis and illustrate how stochastic fluctuations significantly influence disease persistence. Furthermore, sensitivity analysis using Latin Hypercube Sampling (LHS) and Partial Rank Correlation Coefficients (PRCC) identifies the transmission rate from asymptomatic infectious cattle ($β_a$) as the primary driver of transmission. The model effectively captures the dynamics of environmental variability affecting HPAI spread, suggesting that effective control strategies must prioritise the e
Clinical variant classification of pathogenic versus benign genetic variants remains a challenge in clinical genetics. Recently, the proposition of genomic foundation models has improved the generic variant effect prediction (VEP) accuracy via weakly-supervised or unsupervised training. However, these VEPs are not disease-specific, limiting their adaptation at the point of care. To address this problem, we propose DYNA: Disease-specificity fine-tuning via a Siamese neural network broadly applicable to all genomic foundation models for more effective variant effect predictions in disease-specific contexts. We evaluate DYNA in two distinct disease-relevant tasks. For coding VEPs, we focus on various cardiovascular diseases, where gene-disease relationships of loss-of-function vs. gain-of-function dictate disease-specific VEP. For non-coding VEPs, we apply DYNA to an essential post-transcriptional regulatory axis of RNA splicing, the most common non-coding pathogenic mechanism in established clinical VEP guidelines. In both cases, DYNA fine-tunes various pre-trained genomic foundation models on small, rare variant sets. The DYNA fine-tuned models show superior performance in the held-
Antiterminators are essential components of bacterial transcriptional regulation, allowing the control of gene expression in response to fluctuating environmental conditions. RNA-binding antiterminators are particularly important regulatory proteins that play a significant role in preventing transcription termination by binding to specific RNA sequences. These RNA-binding antiterminators have been extensively studied for their roles in regulating various metabolic pathways. However, their role in modulating the physiology of pathogens requires further investigations. This review focuses on these RNA-binding proteins in both Gram-positive and Gram-negative bacteria, particularly on their structures, mechanism of action, and target genes. Additionally, the involvement of the antitermination mechanisms in bacterial pathogenicity will be discussed. This knowledge is crucial for understanding the regulatory mechanisms that govern bacterial pathogenicity, opening up exciting prospects for future research, and potentially new alternative strategies to fight against infectious diseases.
Pathogen identification is pivotal in diagnosing, treating, and preventing diseases, crucial for controlling infections and safeguarding public health. Traditional alignment-based methods, though widely used, are computationally intense and reliant on extensive reference databases, often failing to detect novel pathogens due to their low sensitivity and specificity. Similarly, conventional machine learning techniques, while promising, require large annotated datasets and extensive feature engineering and are prone to overfitting. Addressing these challenges, we introduce PathoLM, a cutting-edge pathogen language model optimized for the identification of pathogenicity in bacterial and viral sequences. Leveraging the strengths of pre-trained DNA models such as the Nucleotide Transformer, PathoLM requires minimal data for fine-tuning, thereby enhancing pathogen detection capabilities. It effectively captures a broader genomic context, significantly improving the identification of novel and divergent pathogens. We developed a comprehensive data set comprising approximately 30 species of viruses and bacteria, including ESKAPEE pathogens, seven notably virulent bacterial strains resistan
The highly pathogenic avian influenza (HPAI) H5 clade 2.3.4.4b has triggered an unprecedented global panzootic. As the frequency and scale of HPAI H5 outbreaks continue to rise, understanding how wild birds contribute to shape the global virus spread across regions, affecting poultry, domestic and wild mammals, is increasingly critical. In this review, we examine ecological and evolutionary studies to map the global transmission routes of HPAI H5 viruses, identify key wild bird species involved in viral dissemination, and explore infection patterns, including mortality and survival. We also highlight major remaining knowledge gaps that hinder a full understanding of wild birds role in viral dynamics, which must be addressed to enhance surveillance strategies and refine risk assessment models aimed at preventing future outbreaks in wildlife, domestic animals and safeguard public health.
The goal of this study is to develop a computational model of the progression of changes in mitochondrial phenotype resulting from infection with pathogenic mycobacteria. This ultimately will enable a large-scale virulence screen of mutant bacterial libraries. Mycobacterium tuberculosis (Mtb) is an intracellular pathogen, but only a small number of its genes have been studied for roles in intracellular host cell survival and replication. Mitochondria are the powerhouse of the host cell and play critical roles in cell survival when attacked by certain pathogens. When Mtb bacteria invade host cells, they induce changes in mitochondrial morphology, making mitochondria a novel target for image processing and machine learning to determine virulence associations of genes in Mtb and potentially other related intracellular pathogens. By hypothesizing mitochondria as an instance of a dynamic and interconnected graph, we demonstrate a statistical approach for quantitatively recognizing novel mitochondrial phenotypes induced by invading pathogens.
Prominently accountable for the upsurge of COVID-19 cases as the world attempts to recover from the previous two waves, Omicron has further threatened the conventional therapeutic approaches. Omicron is the fifth variant of concern (VOC), which comprises more than 10 mutations in the receptor-binding domain (RBD) of the spike protein. However, the lack of extensive research regarding Omicron has raised the need to establish correlations to understand this variant by structural comparisons. Here, we evaluate, correlate, and compare its genomic sequences through an immunoinformatic approach with wild and mutant RBD forms of the spike protein to understand its epidemiological characteristics and responses towards existing drugs for better patient management. Our computational analyses provided insights into infectious and pathogenic trails of the Omicron variant. In addition, while the analysis represented South Africa's Omicron variant being similar to the highly-infectious B.1.620 variant, mutations within the prominent proteins are hypothesized to alter its pathogenicity. Moreover, docking evaluations revealed significant differences in binding affinity with human receptors, ACE2 a
Estimating risk factors for incidence of a disease is crucial for understanding its etiology. For diseases caused by enteric pathogens, off-the-shelf statistical model-based approaches do not consider the biological mechanisms through which infection occurs and thus can only be used to make comparatively weak statements about association between risk factors and incidence. Building off of established work in quantitative microbiological risk assessment, we propose a new approach to determining the association between risk factors and dose accrual rates. Our more mechanistic approach achieves a higher degree of biological plausibility, incorporates currently-ignored sources of variability, and provides regression parameters that are easily interpretable as the dose accrual rate ratio due to changes in the risk factors under study. We also describe a method for leveraging information across multiple pathogens. The proposed methods are available as an R package at \url{https://github.com/dksewell/dare}. Our simulation study shows unacceptable coverage rates from generalized linear models, while the proposed approach empirically maintains the nominal rate even when the model is misspec
Pathogenic bacteria present a large disease burden on human health. Control of these pathogens is hampered by rampant lateral gene transfer, whereby pathogenic strains may acquire genes conferring resistance to common antibiotics. Here we introduce tools from topological data analysis to characterize the frequency and scale of lateral gene transfer in bacteria, focusing on a set of pathogens of significant public health relevance. As a case study, we examine the spread of antibiotic resistance in Staphylococcus aureus. Finally, we consider the possible role of the human microbiome as a reservoir for antibiotic resistance genes.
The spread of harmful mis-information in social media is a pressing problem. We refer accounts that have the capability of spreading such information to viral proportions as "Pathogenic Social Media" accounts. These accounts include terrorist supporters accounts, water armies, and fake news writers. We introduce an unsupervised causality-based framework that also leverages label propagation. This approach identifies these users without using network structure, cascade path information, content and user's information. We show our approach obtains higher precision (0.75) in identifying Pathogenic Social Media accounts in comparison with random (precision of 0.11) and existing bot detection (precision of 0.16) methods.
We suggest that a particular form of social hierarchy, which we characterize as 'pathogenic', can, from the earliest stages of life, exert a formal analog to evolutionary selection pressure, literally writing a permanent developmental image of itself upon immune function as chronic vascular inflammation and its consequences. The staged nature of resulting disease emerges 'naturally' as a rough analog to punctuated equilibrium in evolutionary theory, although selection pressure is a passive filter rather than an active agent like structured psychosocial stress. Exposure differs according to the social constructs of race, class, and ethnicity, accounting in large measure for observed population-level differences in rates of coronary heart disease across industrialized societies. American Apartheid, which enmeshes both majority and minority communities in a social construct of pathogenic hierarchy, appears to present a severe biological limit to continuing declines in coronary heart disease for powerful as well as subordinate subgroups: 'Culture', to use the words of the evolutionary anthropologist Robert Boyd, 'is as much a part of human biology as the enamel on our teeth'.
Evidence-informed policy on infections requires estimates of their effects on health. However, pathogenic variation, whereby occurrence of adverse outcomes depends on the infecting strain, might complicate the study of many infectious agents. Here, we consider the interpretation of epidemiologic studies on effects of infections on health when there is heterogeneity in strain-specific effects and information on strain composition is unavailable. We use potential outcomes and causal inference theory for analyses in the presence of multiple versions of treatment to argue that oft-reported quantities in these studies have a causal interpretation that depends on population frequencies of infecting strains. Moreover, as in other contexts where the treatment-variation-irrelevance assumption might be violated, transportability requires additional considerations, beyond those needed for non-compound exposures. This discussion, that considers potential heterogeneity in strain-specific effects, will facilitate interpretation of these studies, and for the reasons mentioned above, also highlights the value of pathogen subtype data.
Foodborne diseases remain a major public-health burden, and the gastric acid barrier serves as the body's primary chemical defense against ingested microbes. Yet experimentally investigating pathogen survival within this environment is highly challenging. Although recent computational stomach models have provided insights into gastric disorders, none have coupled fluid flow, acid transport, and pathogen population kinetics in a realistic stomach to assess gastric acid barrier function. Here, we develop an imaging-based stomach model that tracks 10,000 massless particles representing pathogen colonies ingested with a liquid meal as they are advected through a dynamic, spatially heterogeneous pH field. The model incorporates acid secretion, peristaltic mixing, and gastric tone-driven emptying. Using this framework, we quantify how hypomotility and altered gastric tone influence pathogen survival. Motility emerges as the dominant factor governing pathogen fate. The hypomotile stomach exhibits weaker mixing, retaining nearly 50% of the initial pathogen population alive 6 minutes after ingestion, compared with less than 30% in healthy cases. It also produces broader acid-dose distributi
Pathogen genome data offers valuable structure for spatial models, but its utility is limited by incomplete sequencing coverage. We propose a probabilistic framework for inferring genetic distances between unsequenced cases and known sequences within defined transmission chains, using time-aware evolutionary distance modeling. The method estimates pairwise divergence from collection dates and observed genetic distances, enabling biologically plausible imputation grounded in observed divergence patterns, without requiring sequence alignment or known transmission chains. Applied to highly pathogenic avian influenza A/H5 cases in wild birds in the United States, this approach supports scalable, uncertainty-aware augmentation of genomic datasets and enhances the integration of evolutionary information into spatiotemporal modeling workflows.
Infections depend on interactions between pathogen and host proteins, but comprehensively mapping these interactions is challenging and labor intensive. Many biological networks have hierarchical, scale-free structure, so we developed a deep learning framework, ApexPPI, that represents protein networks in hyperbolic Riemannian space to capture these features. Our model integrates multimodal biological data (protein sequences, gene perturbation experiments, and complementary interaction networks) to predict likely interactions between pathogen and host proteins through multi-task hyperbolic graph neural networks. Mapping protein features into hyperbolic space led to much higher accuracy than previous methods in predicting host-pathogen interactions. From tens of millions of possible protein pairs, our model identified thousands of high-confidence interactions, including many involving human G-protein-coupled receptors (GPCRs). We validated dozens of these predicted complexes using AlphaFold 3 structural modeling, supporting the accuracy of our predictions. This comprehensive map of host-pathogen protein interactions provides a resource for discovering new treatments and illustrates
Organisms have evolved immune systems that can counter pathogenic threats. The adaptive immune system in vertebrates consists of a diverse repertoire of immune receptors that can dynamically reorganize to specifically target the ever-changing pathogenic landscape. Pathogens in return evolve to escape the immune challenge, forming an co-evolutionary arms race. We introduce a formalism to characterize out-of-equilibrium interactions in co-evolutionary processes. We show that the rates of information exchange and entropy production can distinguish the leader from the follower in an evolutionary arms races. Lastly, we introduce co-evolutionary efficiency as a metric to quantify each population's ability to exploit information in response to the other. Our formalism provides insights into the conditions necessary for stable co-evolution and establishes bounds on the limits of information exchange and adaptation in co-evolving systems.
Epidemic spreading over populations networks has been an important subject of research for several decades, and especially during the Covid-19 pandemic. Most epidemic outbreaks are likely to create multiple mutations during their spreading over the population. In this paper, we study the evolution of a pathogen which can mutate continuously during the epidemic spreading. We consider pathogens whose mutating parameter is the mortality mean-time, and study the evolution of this parameter over the spreading process. We use analytical methods to compute the dynamic equation of the epidemic and the conditions for it to spread. We also use numerical simulations to study the pathogen flow in this case, and to understand the mutation phenomena. We show that the natural selection leads to less violent pathogens becoming predominant in the population. We discuss a wide range of network structures and show how different effects are manifested in each case. We also applied our theory in the context of the Covid-19 pandemic, using relevant epidemiological data collected for this outbreak. We provided explanations for the variants spreading processes observed throughout this pandemic.
During the recent pandemic, a rise in COVID-19 cases was followed by a decline in influenza. In the absence of cross-immunity, a potential explanation for the observed pattern is behavioral: non-pharmaceutical interventions (NPIs) designed and promoted for one disease also reduce the spread of others. We study short-term and long-term dynamics of two pathogens where NPIs targeting one pathogen indirectly influence the spread of another - a phenomenon we term behavioral spillover. We examine how perceived risk of and response to one disease substantially alters the spread of other pathogens, revealing how waves of different pathogens emerge over time as a result of behavioral interdependencies and human response. Our analysis identifies the parameter space where two diseases simultaneously co-exist, and where shifts in prevalence occur. Our findings are consistent with observations from the COVID-19 pandemic, where NPIs contributed to significant declines in infections such as influenza, pneumonia, and Lyme disease.