共找到 20 条结果
Biomedical research effort is distributed highly unevenly across human genes, with a small subset dominating the scientific literature while thousands remain sparsely studied. Whether this imbalance reflects intrinsic biological importance or historically reinforced research bias remains unclear. Understanding how research attention relates to gene properties is essential for more systematic exploration of the human genome. Here, we quantify gene-level publication patterns and integrate sequence features, evolutionary constraint, gene age, expression, and disease associations across stratified gene sets. Using standardized MANE Select annotations, we show that publication counts follow a strongly heavy-tailed distribution. Highly studied genes cluster within a narrow GC-content regime and exhibit lower nonsynonymous substitution rates and lower dN/dS ratios, consistent with stronger long-term evolutionary constraint. In contrast, genes sampled from below rank 10,000 are enriched for evolutionarily younger loci and display moderately elevated dN and dN/dS values, reduced expression magnitude, and increased tissue specificity. At the disease level, research attention concentrates within a limited number of dominant domains, particularly cancer, respiratory, and vascular diseases, whereas congenital and rare disease categories remain comparatively underrepresented. Genes associated with orphan diseases show significantly reduced publication counts. Together, these results demonstrate that research attention is systematically structured across evolutionary, molecular, and disease dimensions. The least-studied genes represent a distinct and underexplored portion of the genome, characterized by features that may reduce experimental tractability. These findings highlight the need for bias-aware research prioritization strategies to broaden discovery and ensure more comprehensive characterization of human genes.
Microsporidia are a fungi-related rapidly evolving lineage of obligate intracellular parasites with poorly known cell biology. Cellular processes in eukaryotes heavily rely on a group of proteins known as the Ras GTPase superfamily, so a comprehensive analysis of these proteins in Microsporidia is expected to provide fundamental insights into the microsporidian cell biology. Adopting a complex bioinformatic approach to cope with the high divergence of microsporidian genes, we reconstructed the evolutionary history of the Ras superfamily of GTPases in Microsporidia and their closest relatives. Our results demonstrate that gene loss along the whole microsporidian phylogeny has been a vastly dominating factor shaping the Ras superfamily gene complements in extant Microsporidia. This trend has most massively affected Hepatospora eriocheir with a mere 12 Ras superfamily genes, the smallest number recorded so far for an autonomously reproducing eukaryote. Also notable is the loss of the virtually ubiquitous Rab GTPase Rab5 and its dedicated regulators in two different microsporidian lineages, or the Rho family GTPases of most Microsporidia having lost C-terminal prenylation. Most unexpectedly, two different microsporidian species lack detectable orthologs of the beta subunit of the signal recognition particle (SRP) receptor, a loss unknown from any other eukaryote. Strikingly, the loss of SRβ is apparently compensated for by the SRP receptor alpha subunit of these species uniquely possessing a predicted transmembrane domain, potentially anchoring the protein in the ER membrane. Our findings thus define new extremes in the impact of reductive evolution on core processes of the eukaryotic cell.
Erwinia amylovora, the causative agent of fire blight, poses a significant threat to global pome fruit production. This study presents a comprehensive genomic analysis of 317 E. amylovora strains and 227 Erwinia phages to elucidate virulence evolution, phage-host dynamics, and the genomic signatures of the co-evolutionary arms race. Our analysis suggests that a substantial portion of E. amylovora's virulence factors (VFs) share evolutionary origins with diverse plant, human, and animal pathogens, underscoring widespread horizontal gene transfer. We identified bacterial phage hydrolases‑like proteins that share phylogenetic and domain-level similarities with phage endolysins. These observations are consistent with the possibility that some bacterial hydrolases originated from phage-derived ancestors, although functional repurposing remains to be experimentally validated. Crucially, our analysis identifies systematic, non-random associations between bacterial defense systems (e.g., RM, CRISPR-Cas, TA) and mobile anti-defense genes. Statistical correlations show strong patterns of co-occurrence and mutual exclusivity, which are consistent with an ongoing phage-bacteria arms race. These patterns provide a genomic basis for generating hypotheses about co-evolutionary dynamics. These findings may advance our understanding of E. amylovora pathogenicity and phage interactions, offering foundational insights for developing targeted phage-based biocontrol strategies against this devastating plant pathogen. Experimental validation of the predicted virulence factors and defense correlations is warranted to confirm their biological roles.
Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.
Krascheninnikovia arborescens is a drought-tolerant subshrub of the Amaranthaceae family that is endemic to China and plays an important role in desert ecosystems. However, little is known about the structure and evolutionary dynamics of its organellar genomes. In this study, we assembled and characterized the mitochondrial and chloroplast genomes of K. arborescens using long-read sequencing data. The mitochondrial genome was assembled as a master circular genome representation of 387,891 bp and contains 62 annotated genes, whereas the chloroplast genome exhibits a typical quadripartite structure of 152,039 bp with 90 genes. The mitochondrial genome harbors abundant repetitive sequences and multiple plastid-derived insertions, indicating a dynamic structural organization. In contrast, gene content remains highly conserved, and all core protein-coding genes show signatures of purifying selection, particularly those involved in ATP synthesis and respiratory metabolism. Predicted RNA editing sites differ substantially between the two organelles, suggesting distinct post-transcriptional modification patterns. Phylogenetic analyses based on shared organellar genes consistently place K. arborescens within Amaranthaceae and support its evolutionary relationships within Caryophyllales. These results reveal a combination of structural dynamism and functional conservation in the organellar genomes of K. arborescens. This study provides a foundation for future comparative and evolutionary studies of Amaranthaceae and expands genomic resources for this family.
BACKGROUND STAPHYLOCOCCUS AGNETIS: is an emerging pathogen primarily associated with bovine mastitis and avian lameness. Despite increasing reports of its occurrence across animal hosts, its genomic diversity and the distribution of antimicrobial resistance (AMR) and virulence-associated genes remain insufficiently characterized. RESULTS: The species S. agnetis possesses an open pan-genome, dominated by cloud gene families enriched in defense mechanisms and genomic plasticity, consistent with gene flux. Evolutionary reconstruction indicated that purifying selection and gene loss are the main signatures of evolutionary dynamics in the S. agnetis pan-genome, with extensive gene loss particularly affecting cell wall biogenesis functions. Notably, significant gene gain events were observed at early-diverging internal nodes of the phylogeny, suggesting that gene acquisition occurred during the early diversification of S. agnetis. AMR profiling identified a limited repertoire of AMR genes. However, the detection of a plasmid-borne AMR gene and the distribution of plasmids highlight the potential for plasmid-mediated dissemination of AMR in S. agnetis. Virulence profiling identified 28 chromosomally located putative virulence-related genes, predominantly homologous to S. aureus, including core adherence factors and sporadically distributed enterotoxin homologs suggestive of acquisition via horizontal gene transfer (HGT). CONCLUSIONS: Collectively, this study provides comprehensive insights into the genomic diversification of S. agnetis and highlights its emerging AMR traits and putative virulence potential in animal-associated settings.
Understanding the organization and evolution of metabolic networks is essential for uncovering how organisms adapt to changing environments. Whereas free-living bacteria typically maintain robust and redundant metabolic systems, endosymbiotic bacteria undergo extreme genome reduction during their adaptation to intracellular life. This process results in highly streamlined and interconnected metabolic networks, in some cases smaller than the theoretical minimum required for sustaining independent cellular function. Using a large-scale comparative framework, we analyzed 101 genomes of insect endosymbiotic bacteria by computing two metabolic network models: metabolite- and reaction-based. We found strong correlations between genome size and key topological properties, including clustering coefficient, network diameter, and number of nodes, indicating that genome reduction directly constrains metabolic network architecture. Despite extensive gene loss, endosymbiotic metabolic networks retain scale-free organization, suggesting the preservation of essential connectivity and robustness. Furthermore, clustering analyses revealed that network topology reflects phylogenetic relationships across bacterial taxa, demonstrating that metabolic organization retains evolutionary signals even in the most reduced genomes. Our findings show that the metabolic networks of insect endosymbiotic bacteria preserve clear evolutionary imprints, revealing a deep connection between genomic reduction, network structure, and phylogenetic history. The complementary use of metabolite- and reaction-based models provide a powerful framework for exploring how symbiotic evolution reshapes metabolic systems while maintaining essential biological organization.
Food ingestion is fundamental for animal survival and growth, with the cessation of feeding upon nutrient fulfillment being tightly regulated by a variety of satiety factors. Notably, sulfakinin/cholecystokinin (SK/CCK)-type neuropeptide signaling has been identified as an inhibitory regulator of food intake across the animal kingdom. However, its regulatory mechanism in feeding in deuterostome invertebrates remains unclear. Here, we characterized SK/CCK-type signaling in a deuterostome invertebrate, the sea cucumber Apostichopus japonicus (phylum Echinodermata). A single SK/CCK-type precursor in A. japonicus generates two mature peptides (AjSK/CCK1, AjSK/CCK2) that activate a shared receptor (AjSK/CCKR), triggering Ca2+ mobilization via the Gαq-dependent pathway and extracellular signal regulated kinase 1/2 (ERK1/2) phosphorylation. Both peptides induce dose-dependent contraction of longitudinal muscles, while AjSK/CCK2 additionally elicits sustained contraction of the posterior intestine, an effect absent in other gut regions. Long-term injection of both peptides reduces food intake and significantly downregulates orexin-type neuropeptide genes (AjOrexin1P, AjOrexin2P) in the circumoral nerve ring (CNR) and intestine. Unlike mammals, where CCK inhibits feeding by contracting the pyloric sphincter to delay gastric emptying, SK/CCK-type peptides in sea cucumbers exert their anorexic effect in part by selectively contracting the posterior intestine, thereby inhibiting intestinal emptying. This divergence in action sites highlights the evolutionary adaptability of SK/CCK-type signaling as a conserved inhibitory regulator of feeding across bilaterian animals. Elucidating these mechanisms in the economically important A. japonicus may inform development of appetite-promoting agents for sustainable aquaculture.
Given its high mortality and broad societal impact, the COVID-19 pandemic is arguably one of the most consequential public health crises of the twenty-first century. Although previous studies have identified several genes associated with COVID-19 susceptibility, relatively little is known about the genes contributing to severe COVID-19, including their evolutionary histories. In the current study, we analyzed IL-4, TLR2, CCL2, and SLC11A1-immunity genes that have previously been implicated in severe COVID-19 and other immune-related diseases-in globally diverse populations from the 1000 Genomes Project. We also tested for associations between genetic variation at these genes and clinical COVID-19 phenotypes in nearly 4000 laboratory-confirmed COVID-19-positive individuals across two datasets from the GEN-COVID Multicenter Study in Italy. Based on our analyses, we identified striking signatures of positive selection within and around all four genes, including extensive haplotype structure, elevated population differentiation, and significant selection coefficients consistent with both ancient and more recent adaptive events. Notably, many of these signals were population-specific, highlighting the role of local selection in shaping immune gene diversity. Several selected alleles were also present in Neanderthal and/or Denisovan genomes, reflecting both shared ancestral polymorphism and archaic introgression within these genes. Functional predictions based on in silico analyses further revealed that a subset of selected alleles maps to transcription factor binding sites and is predicted to influence binding affinity. In addition, our genotype-phenotype analyses uncovered coding variants in TLR2 that were correlated with COVID-19 severity and a related comorbidity, with estimated effect sizes ranging from moderate to large. Interestingly, these significantly associated alleles occur at rare or low frequency in western European and East Asian populations but are absent in populations of African and South Asian descent, indicative of a relatively recent origin. Overall, our study provides new insights into the evolution of biologically relevant immunity genes in modern human populations and identifies genetic variation that may contribute to differences in risk for severe COVID-19.
The Murinae subfamily, one of the most diverse and widely distributed rodent groups, is an important model in ecology, evolutionary biology, and biomedical research. However, deep-level phylogenetic relationships, especially the evolutionary status of key tribes, remain highly contentious. Although mitogenomes are widely used in mammalian evolutionary studies, an integrated analytical framework that combines comparative genomics, selection pressure, codon-usage dynamics, and morphological evidence is currently lacking. Here, we aim to reassess the deep phylogenetic relationships within Murinae, with a particular focus on verifying the proposed sister relationship between the tribes Micromyini and Vernayini. We sequenced and assembled the complete mitogenomes of a single individual of Leopoldamys neilli and Micromys pygmaeus. Through a systematic comparison of mitogenomes from 117 Murinae species-representing 13 tribes and 50 genera, we identified both conserved and lineage-specific genomic features. Murinae mitogenomes were highly conserved in structure and gene order, with length variation primarily attributable to changes in the control region. Nucleotide composition was largely uniform across most lineages, though Micromyini and Vernayini displayed distinct compositional profiles, suggesting possible differences in evolutionary pressures. All protein-coding genes were under strong purifying selection, yet evolutionary rates varied nearly tenfold among genes (e.g., ATP8 versus COX1). Codon-usage bias was predominantly influenced by natural selection and correlated with phylogenetic relationships. Critically, our large‑scale phylogenomic analyses did not support the previously proposed sister‑group relationship between Micromyini and Vernayini. Instead, we recovered Vernayini as the sister group to the remaining Murinae excluding Hapalomys-a topology that received maximal statistical support (PP = 1.00, BS = 100%). This revised phylogenetic hypothesis is supported by multidimensional evidence: (i) principal component analysis of morphological traits reveals significant divergence between the two tribes; (ii) COX1‑based genetic distances (18.93 ± 0.01%) substantially exceed both the generic (12.61%) and subfamilial (18.08%) thresholds reported for Murinae; and (iii) distinct codon‑usage patterns further differentiate these tribes at the molecular level. This study provides the first large‑scale, integrated analysis of comparative mitogenomic data, codon usage bias, and morphological data in Murinae. We supply new genomic resources, reassess the phylogenetic framework of Murinae, definitively clarify that Micromyini and Vernayini are not sister groups, and demonstrate the power of combining genomic, compositional, and phenotypic data in systematic research. These findings offer a foundation for future phylogenomic studies while highlighting the importance of validating mitochondrial‑based hypotheses with additional data sources, such as nuclear markers or morphological traits.
Animal gastrointestinal tracts generally evolved towards a diverse and spatially structured organ system for efficient food digestion. In it, food is chemically broken down and bacterial load reduced by gastric acid in the stomach, acting as a "gatekeeper" for microbes entering the intestines where chyme nutrients and water are absorbed. The natural microbiota across gastrointestinal tract zones support digestion, compete with ingested pathogens and acts itself as an immune stimulus. Despite its important role, several lineages of fish, such as pipefishes, have secondarily lost their stomach and evolved agastric digestion, with unknown consequences to their intestines' microbiomes. Here, we test how stomach loss might affect the microbiome by investigating the fore-, mid- and hindgut's autochthonous microbiota of the Baltic Sea broadnosed pipefish, Syngnathus typhle, and comparing it to the stomach, fore- and hindgut's autochthonous microbiota of the sympatric and ecologically similar three-spine stickleback, Gasterosteus aculeatus. Using 16S-rRNA gene sequencing and qPCR, we show that microbial abundance is high in the stomach, accompanied by high alpha diversity, but low in the intestine of G. aculeatus, although microbial diversity remains at intermediate levels - a pattern almost inversed in S. typhle. G. aculeatus' stomach has the most distinct microbiota across gastrointestinal zones; however, this species' intestines' microbes are also found in S. typhle. In contrast, the pipefish's hindgut is the most distinct zone, and many microbes shared across its whole intestine are not found in the G. aculeatus. Our data supports the notion of the stomach and its distinct microbiome being an immunological gatekeeper for the gut, but also suggests that S. typhle might benefit from the additional microbes as many indicator taxa are suspected to act as mutualistic symbionts. Stomach-loss may therefore be a trade-off between improved chemical digestion capabilities and an immunological gate-keeper vs. improved microbial digestion and increased immune stimulation.
Similarities and differences in the self-assembly of actin filaments from different species inform our understanding of its evolution. However, this basic knowledge is largely incomplete. Here, we systematically characterize assembly kinetics for actin from two yeast species that are five hundred million years apart in evolution, Saccharomyces cerevisiae and Schizosaccharomyces pombe, and compare them to the well-studied rabbit muscle actin from which they diverged a billion years ago. We find that, in the ATP state, both yeast actins behave strikingly like mammalian actin at filament barbed ends. In contrast, yeast actin filaments in both the ADP·Pi and the ADP states depolymerize several-fold faster than their mammalian counterparts, and they release inorganic phosphate over 20-fold faster. We show that the absence of methylation on histidine 73 largely accounts for this faster aging of yeast actin filaments. We also reveal biochemical and mechanical differences between the actins of the two yeasts. Our findings suggest that actins are more diverse and biochemically specialized across species than previously recognized.
WRKY transcription factors are major regulators of plant stress responses and development, yet their evolutionary dynamics across major cereals and dicot needs further characterization. Previous studies cataloged WRKY genes individually in single species, but no comprehensive comparative analysis integrating phylogenomic, syntenic, and compositional analyses across the monocot-dicots divide has been conducted. This knowledge gap limits our ability to identify conserved functional constraints versus lineage-specific evolutionary innovations in WRKY regulatory networks. A high-resolution comparative genomic analysis of 547 WRKY genes was performed across seven plant genomes: Arabidopsis thaliana, Oryza sativa subspecies japonica (126 genes), indica (109 genes), and the previously uncharacterized O. glaberrima (51 genes), Brachypodium distachyon (56 genes), Zea mays (105 genes), and Triticum aestivum (52 genes). Phylogenetic analysis revealed distinct group proportions, with Group III representing 70% of total WRKY genes in cereals compared to only 20.8% Group I and 12.8% Group II, representing a monocot-specific expansion. Group III genes constitute the dominant WRKY classification across all cereal species examined, with pronounced enrichment in cereals (mean 64.9%; range 59.6-69.8%) representing a 2.6-fold difference relative to Arabidopsis (25.0%). This cereal-specific expansion is mechanistically driven by tandem duplication events significantly enriched for Group III genes (Fisher's exact test, p = 0.007), maintained under strong purifying selection (mean Ka/Ks = 0.141). Synteny analysis identified 218 collinear gene pairs between rice and Brachypodium, 186 with Zea mays, 164 with Triticum aestivum, and 142 with Arabidopsis, indicating lineage-specific conservation patterns. Evolutionary rate analysis revealed highly conserved WRKY domains (Ka/Ks = 0.08-0.12) juxtaposed against rapidly evolving flanking regions (Ka/Ks = 0.42-0.78), suggesting strong purifying selection on DNA-binding function. t-SNE analysis identified 22 bridge genes with intermediate compositional profiles spanning the monocot-dicot divide, distributed across four cereal lineages and exhibiting structural properties consistent with a directional evolutionary trajectory from ancestral Group I to derived Group III configurations. Notably, O. glaberrima showed reduced Group I representation (9.8%) and elevated Group III proportion (80.4%), indicating lineage-specific retention patterns during independent domestication. This comprehensive analysis establishes a quantitative framework for dissecting WRKY gene family evolution in cereals, identifies stress-responsive orthologs prioritized for crop improvement, and demonstrates that ancient polyploidy, recent segmental duplication, and differential selection pressure collectively shape cereal regulatory architecture. The study provides a foundation for targeted breeding strategies to enhance climate resilience in major cereal and dicot crops.
The primary symbiont Candidatus Portiera is essential for nutrient provisioning in whiteflies. Genomic instability is a hallmark of Bemisia tabaci-associated Portiera, but the specific molecular evolutions and their metabolic consequences compared to Portiera from other whiteflies remain unclear. To overcome the limited sampling of previous studies, we assembled novel Portiera genomes from seven additional B. tabaci cryptic species. Comparative genomic, phylogenetic, and species delimitation analyses were conducted with other publicly available Portiera genomes. Branch-model selection analysis identified differentially evolved genes in the B. tabaci-associated Portiera, which were significantly enriched in amino acid biosynthetic pathways. The composition of essential amino acids biosynthetic pathways was systematically analyzed across Portiera, host nuclear, and secondary symbiont genomes. B. tabaci-associated Portiera formed a monophyletic lineage with larger genomes, lower coding density, and accelerated evolutionary rates, classified as a single species distinct from Portiera in other whiteflies. Twenty-two genes showed significantly different evolutionary rates with enrichment in amino acid biosynthetic pathways. Key genes for lysine and arginine biosynthesis were lost or pseudogenized in B. tabaci-associated Portiera but remained intact in other whiteflies like Trialeurodes vaporariorum. The synthesis of most other essential amino acids was similarly incomplete across all Portiera, relying on host or secondary symbiont genes for pathway completion. The B. tabaci-associated Portiera represents a unique symbiotic metabolic architecture where host horizontally transferred genes may potentially compensate for Portiera's genomic erosion, contrasting with the more autonomous Portiera in other whiteflies. This study reveals divergent evolutionary trajectories and metabolic integration strategies in whitefly symbiotic systems.
The genus Allium L. (Amaryllidaceae J.St.-Hil.) comprises more than 1,000 species distributed primarily across the Northern Hemisphere and represents one of the most taxonomically and economically important lineages of monocots. Despite extensive molecular research based on nuclear and plastid markers, phylogenetic relationships within the genus remain incompletely resolved, largely due to limited locus sampling and pervasive phylogenetic discordance. We implemented a nuclear phylogenomic approach using the Angiosperms353 target-enrichment probe set to investigate evolutionary relationships and sources of gene tree conflict within Allium. A total of 47 accessions representing 43 species were analyzed using 302 loci. Concatenation-based and multispecies coalescent analyses support the division of Allium into three evolutionary lineages but reveal gene-tree discordance, particularly within the third lineage. Subgenera Allium, Cepa, and Polyprason were recovered as non-monophyletic, with several species showing conflicting placements across analytical frameworks. Multiple complementary analyses demonstrate that incomplete lineage sorting (ILS) is a major contributor to this discordance. Phylogenetic network inference and D-statistic tests detected gene flow among several lineages, indicating reticulate evolution contributed to the genus's complex genomic structure. Divergence time estimates place the origin of Allium in the early Eocene (~ 52 Mya), with major diversification in the Miocene. Rapid lineage diversification during this period likely promoted the persistence of ancestral polymorphisms and contributed to widespread phylogenomic conflict. Our results provide a nuclear phylogenomic framework for Allium and demonstrate that the combined roles of ILS and gene flow have shaped the evolutionary history of one of the largest monocot genera.
Fibroblast growth factors (FGFs) are crucial for animal development, growth and physiological regulation. Early vertebrate evolution was shaped by complex whole-genome duplications (WGDs); after a shared first event (1RV), jawed vertebrates underwent a second distinct WGD (2RJV), while cyclostomes (lampreys and hagfishes) experienced a different, lineage-specific genome expansion (2RCY). Despite their key evolutionary position, the FGF gene family in cyclostomes has remained largely uncharacterized, contrasting with extensive research in jawed vertebrates and their invertebrate chordate relatives. To illuminate this knowledge gap, we conducted a comprehensive genomic survey of FGF genes across four cyclostome species, five jawed vertebrates, and an amphioxus outgroup, leveraging newly available cyclostome genomes. Our analysis surprisingly reveals a significantly reduced FGF repertoire in lampreys (17 genes) and hagfishes (12 genes) compared to jawed vertebrates (22-32 genes). This finding suggests extensive FGF gene loss in cyclostomes following their unique genome duplication history. Phylogenetic and synteny analyses confirm that all eight ancestral FGF subfamilies were first established in the common ancestor of vertebrates. Significantly, we report the discovery of a novel FGF gene, Fgf25, within the FGF4/5/6/25 subfamily, uniquely retained in actinopterygians but lost in sarcopterygians. We also propose a new evolutionary model for Fgf3, suggesting its origin via tandem duplication of an unknown Fgf gene after the 1RV but before the second genome duplication, ultimately leading to the conserved Fgf3-Fgf4-Fgf19 gene linkage in jawed vertebrates. This detailed characterization of the cyclostome FGF repertoire provides insights into early vertebrate FGF evolution and a valuable resource for future investigations into cyclostome evolution and development.
Orthologous gene inference is a crucial technical challenge in evolutionary biology. It typically depends on sequence similarity searches and employs a graph clustering method to infer homologous gene families. However, the all-vs-all sequence similarity search is time-consuming for large-scale genome datasets. In this work, we present LSGFA, a method that detects subgraphs based on the similarity of protein domain architectures and then performs graph clustering within each subgraph, corresponding to sequences that share similar compositions of protein domains. LSGFA carries out four steps in the analysis workflow: protein domain annotation, initial clustering based on Pfam domain architecture, SSN-based clustering, and detection of pan-genomic patterns. Benchmarking against five state-of-the-art tools (OrthoFinder, Roary, PanTA, Panaroo, and PGAP2) across multiple datasets demonstrates that LSGFA achieves a balanced trade-off between computational efficiency and biological accuracy. It takes less time than OrthoFinder while identifying more core genes than high-speed heuristic tools, and its orthogroup inference results show strong consistency with OrthoFinder. Due to the high proportion of proteins with known domain architectures in prokaryotes, LSGFA is particularly well-suited for prokaryotic genomes, where it significantly reduces computational time while yielding accurate homologous gene inference.
Glutathione S-transferases (GSTs) are a crucial gene superfamily for plant stress adaptation. However, their evolutionary trajectories and genomic organizational principles across the plant kingdom remain poorly understood. Through a large-scale comparative genomic analysis of 74 plant species, we identified 4,355 GST genes and classified them into 16 subfamilies. Phylogenetic reconstruction revealed massive and lineage-specific expansion of the stress-responsive Tau and Phi subfamilies in land plants, in contrast to the high conservation of ancient subfamilies (e.g., Theta, Zeta). Structural analysis suggested clade-specific motifs in Tau members associated with functional diversification. In polyploid barnyardgrass, genome duplication led to a disproportionate increase in GST copies: Tau and Phi genes showed unbalanced retention and formed selective clusters on homeologous group 1 and 2 chromosomes, while ancient subfamilies maintained dosage stability. Our study elucidates divergent evolutionary dynamics within the GST family. The lineage-specific expansion and selective retention of clustered Tau/Phi genes during polyploidization suggest a potential genomic signature that may be associated with adaptive evolution in plants. This study provides a comprehensive genomic resource and framework for future functional studies of GSTs in plant stress biology.
ACC-synthase (1-aminocyclopropane-1-carboxylate synthase), also known as the ACS gene, plays a pivotal role in ethylene production, which is of great importance in the fruit ripening process for producing saleable yield (marketable fruit). The ACS gene family presumably controls stress responses, plant growth and development, and particularly fruit ripening. Computational biology was used as an essential tool to identify seven ACS genes in Carica papaya (red hermaphrodite) using an RNA-seq database (NCBI GEO). Further, the phylogenetic relationships of ACS genes determined gene family resemblance in the genomes of Hordeum vulgare, Musa acuminata, C. papaya, and Arabidopsis thaliana; therefore, the identified gene families were further classified into four distinct clades (Type-I, Type-II, Type-III, and Type-IV) in alignment with the well-established Arabidopsis classification. Moreover, encompassing gene structure, domain motifs, cis-element phylogenetic profiling, synteny, and transcriptomic profiling unveiled latent structural and functional attributes within CpACS genes. Through segmental duplication of CpACS, insights into evolutionary duplication events were predicted. The paralogous behavior of ACS genes in C. papaya and a comprehensive transcriptomic analysis demonstrated both up- and down-regulation patterns in response to ethylene treatment at different time points during the fruit ripening process, using the papaya manual handbook V2 (2021). Gene expression showed upregulation of two essential CpACS genes, CpACS5 and CpACS6. RT-qPCR validates the expression of these important genes during fruit ripening. However, one gene, CpACS7, is expressed in the later stages of fruit development. Our results demonstrated novel avenues for understanding the expression pathways of the ACS gene family in red hermaphrodite papaya, and most of these genes were linked to regulating various abiotic stresses, plant growth, and fruit development.
The Amorphophallus konjac is an important specialty cash crop in China and is rich in konjac glucomannan (KGM); however, its long-term exposure to abiotic and biotic stresses has hindered the development of the industry. Lipoxygenase (LOX) is a key enzyme in plant fatty acid metabolism and stress response, playing a vital role in plant growth and development, regulation of secondary metabolism, and responses to biotic and abiotic stresses. To date, systematic studies on the LOX gene family in A. konjac remain lacking. This study aims to identify the A. konjac LOX gene family at the whole-genome level, analyze its sequence characteristics, evolutionary relationships, and expression patterns under various stresses, and provide candidate genes for stress-tolerant molecular breeding. Based on the LOX conserved domain, an HMM model was constructed, and 11 AkLOX family members were identified from the A. konjac genome. According to catalytic sites, the gene family was classified into three subfamilies: 9-LOX, 13-LOX Type I, and 13-LOX Type II, which were unevenly distributed on 4 chromosomes. The encoded proteins contained 848-945 amino acids with molecular weights ranging from 97,110.54 Da to 102,962.26 Da, all of which were unstable hydrophilic proteins. Phylogenetic and collinearity analyses revealed 6 homologous gene pairs between A. konjac and Amorphophallus albus, one pair between A. konjac and Cucumis sativus/Arabidopsis thaliana, and two pairs between A. konjac and Oryza sativa/Solanum tuberosum, indicating the phylogenetic conservation of the LOX family. Promoter cis-element analysis identified 23 types of cis-regulatory elements, most of which were related to plant growth and development, hormone responses, as well as biotic and abiotic stress responses. qRT-PCR results showed that the expression levels of AkLOX1/4/8 were significantly up-regulated under low temperature, drought, and MeJA treatments, supporting the reliability of the cis-element analysis. Under Pectobacterium carotovorum subsp. carotovorum (Pcc) stress, the expression levels of AkLOX2/3/11 were significantly up-regulated, showing an obvious stress-specific response pattern. This study performed genome-wide identification and expression analysis of the AkLOX gene family in A. konjac, and clarified their sequence characteristics, evolutionary relationships, and stress response patterns. It was confirmed that members of this gene family are widely involved in biotic and abiotic stress responses in A. konjac, providing candidate genes for subsequent research on A. konjac resistance to Pcc, low temperature, drought and other stresses.