共找到 20 条结果
Identification of causal genes and pathways is a critical step for understanding the genetic underpinnings of rare diseases. We propose novel approaches to gene prioritization and pathway identification using DNA language model, graph neural networks, and genetic algorithm. Using HyenaDNA, a long-range genomic foundation model, we generated dynamic gene embeddings that reflect changes caused by deleterious variants. These gene embeddings were then utilized to identify candidate genes and pathways. We validated our method on a cohort of rare disease patients with partially known genetic diagnosis, demonstrating the re-identification of known causal genes and pathways and the detection of novel candidates. These findings have implications for the prevention and treatment of rare diseases by enabling targeted identification of new drug targets and therapeutic pathways.
Many rare genetic diseases exhibit recognizable facial phenotypes, which are often used as diagnostic clues. However, current facial phenotype diagnostic models, which are trained on image datasets, have high accuracy but often suffer from an inability to explain their predictions, which reduces physicians' confidence in the model output.In this paper, we constructed a dataset, called FGDD, which was collected from 509 publications and contains 1147 data records, in which each data record represents a patient group and contains patient information, variation information, and facial phenotype information. To verify the availability of the dataset, we evaluated the performance of commonly used classification algorithms on the dataset and analyzed the explainability from global and local perspectives. FGDD aims to support the training of disease diagnostic models, provide explainable results, and increase physicians' confidence with solid evidence. It also allows us to explore the complex relationship between genes, diseases, and facial phenotypes, to gain a deeper understanding of the pathogenesis and clinical manifestations of rare genetic diseases.
Adenosine receptors are G-protein-coupled receptors involved in a wide range of physiological and pathological phenomena in most mammalian systems. All four receptors are widely expressed in the central nervous system, where they modulate neurotransmitter release and neuronal plasticity. A large number of gene association studies have shown that common genetic variants of the adenosine receptors (encoded by the ADORA1, ADORA2A, ADORA2B and ADORA3 genes) have a neuroprotective or neurodegenerative role in neurologic/psychiatric diseases. New genetic studies of rare variants and few novel associations with depression or epilepsy subtypes have recently been reported. Here, we review the literature on the genetics of adenosine receptors in neurologic and/or psychiatric diseases in humans, and discuss perspectives for further genetic research. We also provide an update on the genetic structures of the four human adenosine receptor genes and their regulation - a topic that has not been extensively addressed. Our review emphasizes the importance of (i) better characterizing the genetics of adenosine receptor genes and (ii) understanding how these genes are regulated.
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-consuming and expensive. Efficient computational methods are therefore needed to prioritize genes within the list of candidates, by exploiting the wealth of information available about the genes in various databases. Here we propose ProDiGe, a novel algorithm for Prioritization of Disease Genes. ProDiGe implements a novel machine learning strategy based on learning from positive and unlabeled examples, which allows to integrate various sources of information about the genes, to share information about known disease genes across diseases, and to perform genome-wide searches for new disease genes. Experiments on real data show that ProDiGe outperforms state-of-the-art methods for the prioritization of genes in human diseases.
Prioritizing disease-associated genes is central to understanding the molecular mechanisms of complex disorders such as Alzheimer's disease (AD). Traditional network-based approaches rely on static centrality measures and often fail to capture cross-modal biological heterogeneity. We propose NETRA (Node Evaluation through Transformer-based Representation and Attention), a multimodal graph transformer framework that replaces heuristic centrality metrics with attention-driven relevance scoring. Using AD as a case study, gene regulatory networks are independently constructed from microarray, single-cell RNA-seq, and single-nucleus RNA-seq data. Random-walk sequences derived from these networks are used to train a BERT-based model for learning global gene embeddings, while modality-specific gene expression profiles are compressed using variational autoencoders. These representations are integrated with auxiliary biological networks, including protein-protein interactions, Gene Ontology semantic similarity, and diffusion-based gene similarity, into a unified multimodal graph. A graph transformer assigns NETRA scores that quantify gene relevance in a disease-specific and context-aware ma
Rare diseases are collectively common, affecting approximately one in twenty individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in DNA sequencing, development of new computational and experimental approaches to prioritize genes and genetic variants, and increased global exchange of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize, and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Further, all data generated, currently representing ~7500 individuals from ~3000 families, is rapidly made available to researchers worldwide via the Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL) to catalyze global efforts to develop approaches for genetic diagnoses in rare diseases (https://gregorconsortium.org/data). The majority of these families have undergone prior clinical genetic testing
The identification of disease-associated genes has recently gathered much attention for uncovering disease complex mechanisms that could lead to new insights into the treatment of diseases. For exploring disease-susceptible genes, not only experimental approaches such as genome-wide association studies (GWAS) have been used, but also computational methods. Since experimental approaches are both time-consuming and expensive, numerous studies have utilized computational techniques to explore disease genes. These methods use various biological data sources and known disease genes to prioritize disease candidate genes. In this paper, we propose a gene prioritization method (NRSSPrioritize), which benefits from both local and global measures of a protein-protein interaction (PPI) network and also from disease similarity knowledge to suggest candidate genes for colorectal cancer (CRC) susceptibility. Network Propagation, Random Walk with Restart, and Shortest Paths are three network analysis tools that are applied to a PPI network for the purpose of scoring candidate genes. Also, by looking through diseases with similar symptoms to CRC and obtaining their causing genes, candidate genes a
The study of human genes and diseases is very rewarding and can lead to improvements in healthcare, disease diagnostics and drug discovery. In this paper, we further our previous study on gene disease relationship specifically with the multifunctional genes. We investigate the multifunctional gene disease relationship based on the published molecular function annotations of genes from the Gene Ontology which is the most comprehensive source on gene functions.
A major challenge in biomedical data science is to identify the causal genes underlying complex genetic diseases. Despite the massive influx of genome sequencing data, identifying disease-relevant genes remains difficult as individuals with the same disease may share very few, if any, genetic variants. Protein-protein interaction networks provide a means to tackle this heterogeneity, as genes causing the same disease tend to be proximal within networks. Previously, network propagation approaches have spread signal across the network from either known disease genes or genes that are newly putatively implicated in the disease (e.g., found to be mutated in exome studies or linked via genome-wide association studies). Here we introduce a general framework that considers both sources of data within a network context. Specifically, we use prior knowledge of disease-associated genes to guide random walks initiated from genes that are newly identified as perhaps disease-relevant. In large-scale testing across 24 cancer types, we demonstrate that our approach for integrating both prior and new information not only better identifies cancer driver genes than using either source of information
Published biomedical information has and continues to rapidly increase. The recent advancements in Natural Language Processing (NLP), have generated considerable interest in automating the extraction, normalization, and representation of biomedical knowledge about entities such as genes and diseases. Our study analyzes germline abstracts in the construction of knowledge graphs of the of the immense work that has been done in this area for genes and diseases. This paper presents SimpleGermKG, an automatic knowledge graph construction approach that connects germline genes and diseases. For the extraction of genes and diseases, we employ BioBERT, a pre-trained BERT model on biomedical corpora. We propose an ontology-based and rule-based algorithm to standardize and disambiguate medical terms. For semantic relationships between articles, genes, and diseases, we implemented a part-whole relation approach to connect each entity with its data source and visualize them in a graph-based knowledge representation. Lastly, we discuss the knowledge graph applications, limitations, and challenges to inspire the future research of germline corpora. Our knowledge graph contains 297 genes, 130 dise
Network-based computational approaches to predict unknown genes associated with certain diseases are of considerable significance for uncovering the molecular basis of human diseases. In this paper, we proposed a kind of new disease-gene-prediction methods by combining the path-based similarity with the community structure in the human protein-protein interaction network. Firstly, we introduced a set of path-based similarity indices, a novel community-based similarity index, and a new similarity combining the path-based similarity index. Then we assessed the statistical significance of the measures in distinguishing the disease genes from non-disease genes, to confirm their availability in predicting disease genes. Finally, we applied these measures to the disease-gene prediction of single disease-gene family, and analyzed the performance of these measures in disease-gene prediction, especially the effect of the community structure on the prediction performance in detail. The results indicated that genes associated with the same or similar diseases commonly reside in the same community of the protein-protein interaction network, and the community structure is greatly helpful for th
Identifying disease genes from human genome is an important and fundamental problem in biomedical research. Despite many publications of machine learning methods applied to discover new disease genes, it still remains a challenge because of the pleiotropy of genes, the limited number of confirmed disease genes among whole genome and the genetic heterogeneity of diseases. Recent approaches have applied the concept of 'guilty by association' to investigate the association between a disease phenotype and its causative genes, which means that candidate genes with similar characteristics as known disease genes are more likely to be associated with diseases. However, due to the imbalance issues (few genes are experimentally confirmed as disease related genes within human genome) in disease gene identification, semi-supervised approaches, like label propagation approaches and positive-unlabeled learning, are used to identify candidate disease genes via making use of unknown genes for training - typically in the scenario of a small amount of confirmed disease genes (labeled data) with a large amount of unknown genome (unlabeled data). The performance of Disease gene prediction models are l
The intricate relationship between genetic variation and human diseases has been a focal point of medical research, evidenced by the identification of risk genes regarding specific diseases. The advent of advanced genome sequencing techniques has significantly improved the efficiency and cost-effectiveness of detecting these genetic markers, playing a crucial role in disease diagnosis and forming the basis for clinical decision-making and early risk assessment. To overcome the limitations of existing databases that record disease-gene associations from existing literature, which often lack real-time updates, we propose a novel framework employing Large Language Models (LLMs) for the discovery of diseases associated with specific genes. This framework aims to automate the labor-intensive process of sifting through medical literature for evidence linking genetic variations to diseases, thereby enhancing the efficiency of disease identification. Our approach involves using LLMs to conduct literature searches, summarize relevant findings, and pinpoint diseases related to specific genes. This paper details the development and application of our LLM-powered framework, demonstrating its p
Motivation: Predicting gene-disease associations (GDAs) is the problem to determine which gene is associated with a disease. GDA prediction can be framed as a ranking problem where genes are ranked for a query disease, based on features such as phenotypic similarity. By describing phenotypes using phenotype ontologies, ontology-based semantic similarity measures can be used. However, traditional semantic similarity measures use only the ontology taxonomy. Recent methods based on ontology embeddings compare phenotypes in latent space; these methods can use all ontology axioms as well as a supervised signal, but are inherently transductive, i.e., query entities must already be known at the time of learning embeddings, and therefore these methods do not generalize to novel diseases (sets of phenotypes) at inference time. Results: We developed INDIGENA, an inductive disease-gene association method for ranking genes based on a set of phenotypes. Our method first uses a graph projection to map axioms from phenotype ontologies to a graph structure, and then uses graph embeddings to create latent representations of phenotypes. We use an explicit aggregation strategy to combine phenotype em
Alzheimer's disease is the most common cause of dementia. It is the fifth-leading cause of death among elderly people. With high genetic heritability (79%), finding disease causal genes is a crucial step in find treatment for AD. Following the International Genomics of Alzheimer's Project (IGAP), many disease-associated genes have been identified; however, we don't have enough knowledge about how those disease-associated genes affect gene expression and disease-related pathways. We integrated GWAS summary data from IGAP and five different expression level data by using TWAS method and identified 15 disease causal genes under strict multiple testing (alpha<0.05), 4 genes are newly identified; identified additional 29 potential disease causal genes under false discovery rate(alpha < 0.05), 21 of them are newly identified. Many genes we identified are also associated with some autoimmune disorder.
Neurodegenerative diseases are characterized as the progressive loss of neural cells, e.g. neurons, glial cells. Ageing, monogenic variations, viral infections, and many other factors are determined and speculated as causes for them. While many individual genes, such as APP for Alzheimer disease and HTT for Huntington disease, and biological pathways are studied for neurodegenerative diseases, system-wide pathogenesis studies are limited. In this study, we carried out a meta-analysis of RNA-Seq studies for three neurodegenerative diseases, namely Alzheimer's disease, Parkinson's disease and Amyotrophic Lateral Sclerosis (ALS) to minimize the batch effect derived differences and identify the similarly altered factors among studies. Our main assumption is that these three diseases share some pathological pathway pattern. For this purpose, we downloaded publicly available Alzheimer's disease (84 patients + 33 controls = 117 individuals), Parkinson's disease (28 patients + 43 controls = 71 individuals) and ALS (2 studies: 46 patients + 25 control = 71 individuals) RNA-Seq data from Sequence Read Archive (SRA) database. The significantly differentially expressed genes common to these st
Prion diseases are invariably fatal and highly infectious neurodegenerative diseases affecting humans and animals. By now there have not been some effective therapeutic approaches to treat all these prion diseases. In 2008, canine mammals including dogs (canis familials) were the first time academically reported to be resistant to prion diseases (Vaccine 26: 2601--2614 (2008)). Rabbits are the mammalian species known to be resistant to infection from prion diseases from other species (Journal of Virology 77: 2003--2009 (2003)). Horses were reported to be resistant to prion diseases too (Proceedings of the National Academy of Sciences USA 107: 19808--19813 (2010)). By now all the NMR structures of dog, rabbit and horse prion proteins had been released into protein data bank respectively in 2005, 2007 and 2010 (Proceedings of the National Academy of Sciences USA 102: 640--645 (2005), Journal of Biomolecular NMR 38:181 (2007), Journal of Molecular Biology 400: 121--128 (2010)). Thus, at this moment it is very worth studying the NMR molecular structures of horse, dog and rabbit prion proteins to obtain insights into their immunity prion diseases. This article reports the findings of th
Rare disease diagnosis increasingly relies on integrating genomic, phenotypic and transcriptomic evidence, yet these signals remain difficult to reconcile within a common interpretive framework. Here we present RareCollab, an LLM-powered framework for multimodal reasoning in Mendelian disease diagnosis that integrates more than 100 diagnostic evidence signals across DNA, RNA, phenotype, curated variant-level knowledge, and in-silico pathogenicity evidence. This design enables large language models to operate as calibrated, interpretable reasoning modules rather than as a single end-to-end ranker. We applied RareCollab to 890 patients from three cohorts, including 119 Undiagnosed Diseases Network probands with paired DNA and RNA data, constituting a large systematic benchmark for multimodal rare disease diagnosis under paired genomic and transcriptomic evaluation. In this real-world multimodal benchmark, RareCollab prioritized 94% of diagnostic genes within the top 10. Across recall thresholds from top 1 to top 10, it consistently outperformed proprietary phenotype-driven LLM baselines including Claude Sonnet 4.6 and GPT-5-mini by more than 25% on average and surpassed established s
One of the important issues in oncology is finding the genes that perturbation the cell functionality, and result in cancer propagation. The genes, namely driver genes, when they mutate in expression, result in cancer through activation of the mutated proteins. So, many methods have been introduced to predict this group of genes. These are mostly computational methods based on the number of mutations of each gene. Recently, some network-based methods have been proposed to predict Cancer Driver Genes (CDGs). In this study, we use a network-based approach and relative importance of each gene in the propagation and absorption of genes anomalies in the network to recognize CDGs. The experimental results are compared with 19 previous methods that show our proposed algorithm is better than the others in terms of accuracy, precision, and the number of recognized CDGs.
Livestock mobility, particularly that of small and large ruminants, is one of the main pillars of production and trade in West Africa: livestock is moved around in search of better grazing or sold in markets for domestic consumption and for festival-related activities. These movements cover several thousand kilometers and have the capability of connecting the whole West African region thus facilitating the diffusion of many animal and zoonotic diseases. Several factors shape mobility patterns even in normal years and surveillance systems need to account for such changes. In this paper, we present a procedure based on temporal network theory to identify possible sentinel locations using two indicators: vulnerability (i.e. the probability of being reached by the disease) and time of infection (i.e. the time of first arrival of the disease). Using these indicators in our structural analysis of the changing network enabled us to identify a set of nodes that could be used in an early warning system. As a case study we simulated the introduction of F.A.S.T. (Foot and Mouth Similar Transboundary) diseases in Senegal and used data taken from 2020 Sanitary certificates (LPS, laissez-passer