The Vertebrate Genomes Project (VGP) aims to produce complete and near-error-free reference genomes for all ~70,000 extant vertebrate species1. Organized in four phases, it progressively targets all vertebrate orders, families, genera, and eventually all species. Here we present the completion of VGP Phase I, delivering reference genomes for ~95% of vertebrate orders, along with additional lineages within those orders, totaling 816 species and 1.6 trillion base pairs of main haplotype sequence. These genomes were assembled and annotated over an 8-year period (2018-2026) of rapid advances in genome sequencing, assembly, and annotation methods2-4, alongside the growth of associated consortium initiatives and international collaborations5-9. They represent some of the highest-quality vertebrate genomes currently available, and most have become the primary reference for their respective species in public databases. Comparative analyses across a subset of 579 species when we reached a threshold of 85% of orders allowed us to reconstruct the genome of the last common ancestor of all vertebrates 500 million years ago, identify diverse modes of sex chromosome evolution, reveal clade-specific three-dimensional genome architecture, discover methylated epigenetic landscapes across vertebrates, and provide a framework for studying gene and pseudogene evolution, immune loci, cancer-associated genes, and other trait-associated loci. Approximately a quarter of this subset are listed as Vulnerable to Critically Endangered by the IUCN Red List of Threatened Species, and have enabled more advanced genomic investigations of extinction risk. VGP Phase I delivers a reference backbone for vertebrate genomics, enabling discoveries that would otherwise remain out of reach across evolution, conservation, and medicine.
The European roe deer (Capreolus capreolus) is one of the most widespread ungulates in Europe, with a phylogeographic structure mainly shaped by Pleistocene glacial cycles and secondary contacts with the Siberian roe deer (C. pygargus). We sequenced 52 complete mitogenomes of C. capreolus from Slovenia, Poland and France, and combined them with 24 publicly available sequences of C. capreolus and C. pygargus, yielding an alignment of 76 genomes representing 59 haplotypes (42 from C. capreolus and 17 from C. pygargus). Phylogeographic structure was assessed using a median-joining network, and divergence times were estimated using a time-calibrated Bayesian phylogeny based on mitochondrial coding regions, incorporating published ancient C. pygargus mitogenomes. We additionally screened mitochondrial protein-coding genes for selection. The haplotype network recovered the three major European roe deer clades (Eastern, Central, and Western) and detected Central-clade haplotypes in France. Two Polish haplotypes (Cp9 and Cp10), detected in C. capreolus, clustered within the C. pygargus mitochondrial lineage, supporting mitochondrial introgression. Time-calibrated phylogenies placed introgressed haplotypes within established C. pygargus lineages. Selection analyses provided limited evidence for episodic positive selection restricted to a small number of codons. Whole mitogenomes improve resolution of roe deer phylogeography and reveal introgressed maternal lineages, while time-calibrated phylogenies and selection tests add evolutionary context for interpreting mtDNA diversity in genus Capreolus.
Krascheninnikovia arborescens is a drought-tolerant subshrub of the Amaranthaceae family that is endemic to China and plays an important role in desert ecosystems. However, little is known about the structure and evolutionary dynamics of its organellar genomes. In this study, we assembled and characterized the mitochondrial and chloroplast genomes of K. arborescens using long-read sequencing data. The mitochondrial genome was assembled as a master circular genome representation of 387,891 bp and contains 62 annotated genes, whereas the chloroplast genome exhibits a typical quadripartite structure of 152,039 bp with 90 genes. The mitochondrial genome harbors abundant repetitive sequences and multiple plastid-derived insertions, indicating a dynamic structural organization. In contrast, gene content remains highly conserved, and all core protein-coding genes show signatures of purifying selection, particularly those involved in ATP synthesis and respiratory metabolism. Predicted RNA editing sites differ substantially between the two organelles, suggesting distinct post-transcriptional modification patterns. Phylogenetic analyses based on shared organellar genes consistently place K. arborescens within Amaranthaceae and support its evolutionary relationships within Caryophyllales. These results reveal a combination of structural dynamism and functional conservation in the organellar genomes of K. arborescens. This study provides a foundation for future comparative and evolutionary studies of Amaranthaceae and expands genomic resources for this family.
In recent years, many high-quality reference genome sequences for arthropod species have been generated. Although most genome papers describe their protocols and metrics, no consensus exists on the data that should be included in genome reports. Here, we review current standards across seven key stages of an arthropod genome project (budgeting, sourcing and vouchering, sample preparation and sequencing, genome assembly, analysis reproducibility, databasing, and genome annotation) and identify persistent gaps in standards as well as their implementation. To assess current standards reporting in the community, we surveyed 100 arthropod genome papers published in 2024. The use of long reads to assemble highly contiguous arthropod genomes is now standard practice when adequate input DNA is available, and basic assembly contiguity and conserved gene content statistics are consistently reported. However, there is less standardization in pre- and post-assembly procedures and metrics. When comparing Darwin Tree of Life (DToL) genome notes to other journals, publications from the latter group were less likely to describe compliance with ethical collection practices, sample vouchering, post-assembly curation steps, and assembly quality metrics beyond basic contiguity and completeness values. Genome annotation practices are highly variable: some genome note formats do not explicitly require annotation, and while the reporting rate of protein-coding gene annotations is higher in non-DToL publications, the submission rate of annotations to centralized sequence databases is much lower. Our findings highlight critical opportunities to harmonize reporting standards and promote their dissemination, ensuring that future arthropod genomes are both comparable and maximally reusable for large-scale comparative and applied research.
Fermented vegetables in Yunnan Province, China, harbor abundant microbial diversity. However, the development of indigenous starter cultures remains under-utilized. Genomic information regarding Leuconostoc (L.) mesenteroides isolates from this region is particularly scarce. To assess the genomic characteristics of eight L. mesenteroides isolates from traditional Yunnan fermented vegetables, we performed whole-genome sequencing and conducted a comparative analysis with 21 publicly available vegetable-derived genomes. Comparative genomic analysis revealed marked variation in genome size and plasmid content, and pangenome analysis indicated an open configuration. Core-genome multilocus sequence typing (cgMLST) of the eight indigenous isolates showed high allelic diversity, indicating a genetically heterogeneous and non-clonal population. Phylogenomic analysis revealed that the evolutionary relationships among the 29 strains were not strictly correlated with their vegetable sources, suggesting an influence from other factors, such as geographic origin and region-specific processing methods. Similar to the profiles of the 21 publicly available genomes, inactive prophages, intrinsic vancomycin resistance genes, and genomic island fragments were detected in eight isolates, whereas no known virulence genes were identified. Bacteriocin gene clusters varied among strains, while stress tolerance and probiotic-related genes were conserved. Overall, these results provide genomic indications relevant to the safety, adaptability, and fermentation potential of indigenous L. mesenteroides from Yunnan. However, because these functional traits are inferred solely from genomic predictions, subsequent experimental validation is essential to confirm their phenotypic properties and technological efficacy.
The products of isocyanide synthase (ICS) biosynthetic gene clusters (BGCs) have been implicated in microbial interactions, pathogenesis, and metal homeostasis. While several ICS BGCs have been described as mediating metal-associated ecologies, the evolutionary history of these clusters is unexplored among Lecanoromycetes, a clade comprised predominantly of symbiotic, lichen-forming fungi (LFF), which are known to thrive in both metal-contaminated and scarce environments. Analyzing nearly 4,000 fungal genomes, including 90 Lecanoromycetes, we identified a significant 3-fold enrichment of ICS-encoding genes in lichenized fungi compared with non-lichenized counterparts. This expansion includes six distinct clades enriched in LFF. Evolutionary reconstruction uncovered a widespread "split" variant of the copper-responsive (crmA) pathway, where the ICS- and NRPS-like components are encoded on separate genes, contrasting to the canonical "fused" crmA megasynthase. Metabolic characterization and genetic deletions in Fusarium graminearum confirmed that this split architecture is functionally equivalent to the fused form. Our chemical analysis suggests the first evidence of a potential leucine-derived isocyanide metabolite. Phylogenetic reconstruction indicates that the fused crmA arose from a split ancestor whose NRPS-like subdomain evolved from a canonical thioester reductase architecture likely via domain replacement. Redefinition of the crmA pathway to include the split ICS/NRPS-like variant reveals that crmA is one of the most prevalent ICSs in the fungi. We developed a website (https://isocyanides.fungi.wisc.edu/) that facilitates the exploration and downloading of all major results in our study. Our work demonstrates how exploring understudied fungal lineages can define new specialized metabolism lineages and reshape our understanding of the evolution of broadly conserved biosynthetic pathways.
The Dendrocalamus giganteus complex comprises D. giganteus, D. calostachyus and D. sinicus, being the largest known, iconic bamboo species and is economically, ecologically and culturally significant in Southeast Asia, serving as a pillar of daily life of indigenous people. However, lack of understanding of its genetic diversity pattern and population history has hindered effective germplasm conservation and development as sustainable non-timber forest resources. Here, we present a population genomic study of the giant bamboos with whole-genome resequencing of 284 accessions of three closely related species across their potentially native geographical ranges in Myanmar and Yunnan Province, China. We identified seven highly supported phylogenetic clades for the populations of the D. giganteus complex, and all populations exhibit low levels of genetic diversity while a high degree of genetic differentiation among them. One of them in the remote northern and northwestern Myanmar was found as the hotspot of genetic diversity of the giant bamboos. Tajima's D value and demographic history inference suggested the occurrence of population bottleneck, leading to a sharp decline in Ne during the last glacial period. Strikingly, D. sinicus displayed the lowest genetic diversity in the complex likely due to predominance of selfing or inbreeding. Overall, our study provides valuable insights into the evolutionary history and population genetics of the D. giganteus complex, serving as an important foundation for developing effective conservation strategies for the giant bamboos in Southeast Asia.
Tropical regions are biodiversity-rich, yet remain underrepresented in the availability of genomic resources, as is evident in the Western Ghats of India, a biodiversity hotspot with high endemism. Here, we present high-quality, de novo genome assemblies for seven birds, representing seven families distributed in the Western Ghats: Black-naped Monarch (Monarchidae: Hypothymis azurea), Indian Yellow Tit (Paridae: Machlolophus aplonotus), Brown-cheeked Fulvetta (Leiothrichidae: Alcippe poioicephala), Malabar Trogon (Trogonidae: Harpactes fasciatus), Blue-bearded Bee-eater (Meropidae: Nyctyornis athertoni), Malabar Whistling-Thrush (Muscicapidae: Myophonus horsfieldii), Orange-headed Thrush (Turdidae: Geokichla citrina). Using a hybrid Oxford Nanopore long reads-Illumina short reads approach, we assembled genomes with sizes ranging from 1.03 to 1.13 Gbp. All assemblies demonstrated high contiguity and completeness (BUSCO scores > 97%, UCEs > 4799). Repeat masking identified ~10% of the genomes as interspersed repeats. Of the predicted protein-coding genes, an average of 9619 per species received high-confidence functional annotation hits. Comparative analysis showed our assemblies had significantly higher contiguity than the median of existing avian genomes on NCBI (Wilcoxon test, p = 0.00226). Our genome assemblies fill a key geographic and taxonomic gap in the genomic data and provide a foundational resource for evolutionary and ecological research in the Old-World tropics.
Reference genome assemblies are essential infrastructure for investigating phylogeny, and population/conservation genetics of wild organisms. Birds serve as model vertebrates in ecology and evolutionary biology due to their well-documented natural histories and extensive community science data. We release a set of 350 newly assembled avian genomes, which, when combined with 97 previously published genomes, represent 447 of the bird species recorded in Denmark, the Faroe Islands, and Greenland-the largest regional dataset of a vertebrate group to date. These genomes are published for various research activities. This data release advances the global effort to build comprehensive and accessible biodiversity genomic resources for the research community.
Complete, haplotype-resolved genome assemblies have provided unprecedented insight into the evolution of structurally complex, rapidly evolving regions of human genomes; however, population-scale pangenome resources of our closest relatives, chimpanzees and bonobos (genus, Pan), are necessary to ascertain the origins and evolutionary context of these loci. Here, we sequence and assemble 58 haplotypes from four distinct Pan clades to high contiguity (median contig NG50=54 Mb), including eight near-T2T genomes. These genomes reveal previously intractable genetic variation increasing estimates of genome-wide genetic diversity 6-37% across populations compared to short-read estimates. We identify recurrent structural polymorphisms across species impacting genes associated with immune response and host-pathogen interaction and find that structural variants (SVs) are 170- to 260-fold more likely than single nucleotide variants (SNVs) to exhibit high-impact effects across species. Contrasting SV patterns across primates we find that transposable element mutation rates differ by as much as threefold between species. We show that human disease-associated short tandem repeat (TR) loci have uniquely expanded in humans sensitizing our species to these TR-expansion disorders. Physically phased haplotypes enable reconstruction of genome-wide genealogical histories, uncovering ancient, functional genetic variation maintained by balancing selection, as well as signatures of recent adaptation in chimpanzee subspecies. Several malaria-associated loci exhibit ancient structural polymorphism, including the African great ape-specific glycophorin (GYP) gene expansion. We characterize the sequence, structure, and composition of diverse glycophorin haplotypes in humans and chimpanzees. We identify independent malaria-protective GYPA-B fusion events in humans and novel chimpanzee glycophorin genes resulting from both ancient and recent fusion events demonstrating parallel adaptations to pathogen resistance across hominins. Together, our resource highlights the critical importance of nonhuman primate population-scale pangenomics for understanding the evolution of complex genome structures and the biodiversity of our endangered closest living relatives.
Microsatellites within genomes play crucial roles in regulating gene expression, DNA replication, and chromosomal structure and function. Analyzing the composition and distribution patterns of microsatellites in closely related species not only reveals their evolutionary dynamics and adaptive mechanisms but also provides essential technical support for applications in genetic breeding, species conservation, and disease research. As one of the world's most captivating animal groups, the landscape patterns of microsatellites across feline genomes remain to be systematically characterized. This study utilized high-quality genomic data to conduct a systematic comparative analysis of microsatellite landscape distribution patterns across the genomes of 13 felid species. The findings revealed that microsatellite abundance and distribution exhibit species-specific characteristics, with a non-random genomic distribution and a negative correlation between microsatellite abundance and repeat length. The predominant distribution pattern followed the sequence: single > double > quadruple > triple > quintuple > sextuple nucleotide repeats. Microsatellite abundance peaked in intergenic regions, whereas trinucleotide repeats were more prevalent within exons. Coding regions showed a marked preference for trinucleotide and hexanucleotide repeats. Enrichment analysis of GO and KEGG pathways indicated that coding sequences containing microsatellites were primarily involved in transcription and translation processes. Our study elucidates the distribution patterns and characteristics of microsatellites across diverse feline species, providing significant insights into their evolutionary mechanisms and functional roles. Furthermore, these findings establish a valuable reference and foundational dataset for the future development of high-quality, species-specific microsatellite markers in felids.
Subterranean ecosystems host highly specialized and often cryptic biodiversity, yet even intensively studied landscapes may conceal deeply divergent vertebrate lineages. Here we describe Demogorgonichthys arcanus gen. et sp. nov., a new genus and species of cave-obligate fish discovered in Bobcat Cave, a long-monitored karst system on Redstone Arsenal in northern Alabama, USA. Phylogenetic analyses based on mitochondrial nd2, complete mitochondrial genomes, and the nuclear gene rhodopsin place D. arcanus within Amblyopsidae but reveal deep divergence from all described genera, including extensive lineage-specific degeneration of a vision-related gene. Notably, D. arcanus occurs in syntopy with the Southern Cavefish (Typhlichthys subterraneus) despite lacking a close phylogenetic relationship, providing evidence for multiple independent evolutionary origins of cave adaptation within a single groundwater system. This discovery highlights persistent detection bias in groundwater ecosystems and demonstrates that cryptic vertebrate diversity can persist even in well-characterized environments. Extreme endemism and restriction to a single cave-aquifer system further underscore the vulnerability of subterranean biodiversity and the importance of integrating evolutionary and conservation perspectives.
Ecologists and conservation biologists rely on genetic diversity as a key essential biodiversity variable (EBV) used to track population health and dynamics, and utilize the population parameter θ (estimated by the average pairwise genomic distance) as a key metric of diversity. While whole-genome-sequencing (wgs) is increasingly affordable, it will be considerable time before the full diversity of life is represented by high-quality assembled genomes; even then, constant monitoring will still require repeated sampling of populations. In contrast, genome skimming (low-coverage, short-read wgs) is highly cost-effective but challenging to analyze because the coverage is too low for assembly and reliable error correction. Mature methods, such as Mash, exist for estimating pairwise genomic distances based on the Jaccard similarity of k-mer sets computed using sketching techniques. Some, such as Skmer, additionally model the impacts of low coverage. These methods have been successfully applied to assembly-free species identification and phylogenetics; however, their use in population genetics has been limited. This is because these methods implicitly treat genomes as haploid and heterozygosity confounds true estimates of genomic distance for diploid organisms. In this paper, we address this problem through a number of technical advances. First, we use coalescent theory to mathematically derive how the Jaccard index between two diploid samples changes with the scaled population size parameter (θ). Next, we derive an estimator that computes θ from the Jaccard index, in addition to several auxiliary variables, which we also estimate from the genome skims. The resulting method, DipSkmer, enables more accurate estimates of coverage, sequencing error, and pairwise nucleotide distance for diploid samples. Analyses of both simulated and empirical datasets show that for diploids and low distances (e.g., < 2%), DipSkmer produces the most accurate pairwise distance estimates, outperforming existing alignment-free methods such as Mash and Skmer, and closely approximates ANGSD, a reference and alignment-based tool. The code for DipSkmer is available at https://github.com/echarvel3/ReSkmer/tree/DipSkmer-REFACTOR. Simulation scripts and environments are available at https://github.com/echarvel3/dipskmer_scripts.
Long-term studies of isolated animal populations have greatly improved the understanding of various evolutionary processes. However, potentially elevated inbreeding in those compared to wild populations is a common concern. Conventionally, inbreeding has been investigated using reconstructed pedigrees, but nowadays it can be done directly at the genomic level. Here, we utilise genomic data from an intensively studied isolated rhesus macaque (Macaca mulatta) population on the small island Cayo Santiago (Puerto Rico), founded in 1938 with wild animals from India. We quantified inbreeding levels by inferring runs of homozygosity (ROH), i.e., long identical haplotypes inherited from both parents. We identified ROH in 97 ~5x-coverage genomes from Cayo Santiago and, for comparison, in 79 rhesus macaque genomes from five wild populations from China. Notably, this conventionally considered low-coverage data proved sufficient to infer ROH > 4 centimorgans long after imputing the genomes using a reference panel. Our results revealed that the ROH-derived effective population size on Cayo Santiago, 420 individuals, falls within the ranges we inferred in wild populations. Moreover, a general scarcity of individuals with long ROH in both the Cayo and wild populations indicates very few cases of close-kin breeding, suggesting that mechanisms to avoid close-kin breeding operate in rhesus macaques, in both wild and isolated populations. Taken together, our results suggest that Cayo Santiago remains a representative study population.
Freshwater biodiversity is declining at alarming rates, and the biological characteristics of small freshwater fishes make them particularly vulnerable to habitat degradation and fragmentation. Pygmy perches are small Australian fishes whose populations have recently undergone rapid declines due to human-induced pressures. They also exhibit paedomorphic miniaturization, an evolutionary phenomenon whose consequences for the persistence and diversification of freshwater fishes remain unclear. Genomic resources for pygmy perches would significantly enhance both conservation efforts and evolutionary research. Here, we describe chromosome-level genome assemblies for the Southern Pygmy Perch (Nannoperca australis) and the Yarra Pygmy Perch (N. obscura). We provide evidence for high synteny among percichthyid genomes and for positive selection on size- and growth-related genes in the pygmy perch lineage. The N. australis and N. obscura assemblies are 675.7 and 653.9 Mbp in length, with scaffold N50 values of 26.98 and 26.22 Mbp, respectively. Each assembly comprises 24 pseudo-chromosomes and 890 or 1,476 contigs, with repeat contents of 22.3% and 28.1%. We predicted 26,390 and 24,654 protein-coding genes, and BUSCO completeness scores for Teleostei genes were 96.9% and 96.5%, respectively for N. australis and N. obscura. Their genomes are comparable in quality to the best currently available percichthyid assemblies. These genomic resources will support conservation management efforts and provide a foundation for comparative studies on the adaptive significance of miniaturization.
As a hallmark of avian ecological innovation, powered flight has fundamentally shaped diverse aspects of birds. The energy demand of flight may have mutagenic impacts on genomes, influencing how fast genomes evolve. However, the relationship between flight, metabolism, and evolutionary rates remains relatively underexplored. Leveraging 363 newly available avian genomes from >90% of avian families, we quantified three distinct types of genomic evolutionary rates to capture a broad spectrum of mutational processes. By combining four flight-related traits and three metabolic metrics, we uncovered significant associations between flight style, metabolism, and multiple evolutionary rates. Next, using a causal inference framework, we demonstrated that metabolism accounted for 43.3% of the total effect between flight and evolutionary rate, underscoring its key role. Together, our findings establish a robust connection between flight, metabolism, and evolutionary rates, offering new insights into how key innovations and associated phenotypes shape the tempo of genome evolution.
Human activity is driving a biodiversity crisis marked not only by accelerating species extinctions but also by rapid erosion of genetic and phylogenetic diversity. De-extinction science has emerged in response. Here, we synthesize de-extinction as a conservation workflow that integrates ancient and museum genomics, comparative genome analysis, high-precision genome engineering, stem cell platforms, advanced assisted reproductive technologies (ART), emerging ex-utero gestation systems, and AI-enabled ecological modelling and monitoring. We frame three primary conservation applications: (i) reconstruction of lost ecological functions via engineered de-extinct species, (ii) genetic rescue and de-endangerment through restoration of lost diversity, repair of deleterious alleles, and enhancement of adaptive potential in living species, and (iii) acceleration of enabling technologies, particularly ART and stem cell capabilities, that remove reproductive bottlenecks in threatened taxa. Recent advances in sequencing and assembly now support high-quality genomes from extinct and archival material (e.g., thylacine, mammoth, dodo), enabling identification of functionally relevant variation, much of which resides in regulatory landscapes rather than coding sequence alone. In parallel, next-generation editing systems (base, prime, twinPE and large-fragment integration approaches) are shifting the field from single-variant correction to systematic rewriting of loci and regulatory modules, supported by long-read validation and stringent cell-line quality control. We discuss complementary cellular routes (somatic cells and pluripotent stem cells), the promise of in vitro gametogenesis and synthetic embryo models, and the potential of artificial gestation to overcome surrogate scarcity and interspecies incompatibility. Finally, we highlight rewilding as the decisive endpoint, requiring adaptive management, Indigenous partnership, and high-fidelity AI-assisted monitoring. Taken together, de-extinction is best understood as a technology engine for conservation, one that expands the actionable toolkit for preventing extinctions, restoring resilience, and rebuilding lost biodiversity.
Emergomyces (Ajellomycetaceae, Onygenales) is a genus of dimorphic fungal pathogens that cause severe, opportunistic infections around the world. We used long-read nanopore sequencing to generate de novo genome assemblies followed by comparative analyses for the type-strains of Emergomyces africanus, Emergomyces canadensis, Emergomyces crescens, Emergomyces europaeus, Emergomyces orientalis, Emergomyces pasteurianus, and Emergomyces soli. The average Emergomyces genome was found to be 32,965,087 bp in size (range 28,687,700-35,863,369 bp) with GC%-contents of 42.01-45.72%. Average Nucleotide Identity analysis of the seven type-strain genomes with the publicly available Emergomyces and Blastomyces genomes was performed, showing values of 85-100% similarity between genomes and 85-91% between the twelve species tested. Furthermore, we formally validate the species descriptions for Emergomyces crescens and Emergomyces soli.
Haemaphysalis longicornis is an important tick species and pathogen vector characterized by the co-circulation of triploid parthenogenetic and diploid bisexual strains. However, the evolutionary basis of parthenogenesis in this species is unclear. Here we report reference-quality, haplotype-resolved genome assemblies of the parthenogenetic strain and two reference-quality genomes of the bisexual strains. Comparative genomic analysis revealed high collinearity between the parthenogenetic and bisexual genomes, with a stable chromosomal architecture maintained among the three haplotypes of the parthenogenetic strain. The parthenogenetic H. longicornis genome exhibited a major expansion in cell cycle-related gene families, including the inhibitor of apoptosis protein (IAP) family, but was characterized by a contraction in other gene families. Population resequencing of 179 individuals revealed two distinct subpopulations, with chromosome 7 harbouring high genetic differentiation and several candidate genes probably associated with parthenogenesis. Functional experiments showed that knockdown of the BIRC5 gene, a member of the IAP family, suppressed oviposition in both strains, with the parthenogenetic strain exhibiting milder adverse effects probably due to a stronger transcriptional response. Overall, our results reveal the genomic and evolutionary features associated with polyploid parthenogenesis in H. longicornis.
Marine and coastal fungi experience intense environmental variability, yet the genomic features associated with tolerance to such conditions remain unclear. From 56 fungal isolates collected along the Lailai rocky shore in northern Taiwan, we selected the coastal isolate Annulohypoxylon annulatoides RYS0019 for phenotypic and genomic investigation because of its prevalence and distinctive stress-response profile. Compared with five bark-derived conspecific strains, RYS0019 showed distinct growth and recovery dynamics under salinity, temperature, and UV-associated stress treatments. We generated a high-quality 41.8 Mbp de novo genome assembly with 11,523 predicted proteins and compared it with 15 other Hypoxylaceae genomes. Across Annulohypoxylon genomes, we identified variably sized and dispersed AT-rich isochores that are repeat-enriched and gene-poor. Despite variation in AT content, core gene content and Pfam domain profiles remained broadly conserved. Most AT-rich isochores were embedded within syntenically conserved regions and showed limited positional conservation across species, supporting recurrent, lineage-specific formation or expansion after species divergence. These regions also exhibit several sequence and structural features consistent with scaffold/matrix attachment regions (S/MARs), raising the possibility that they influence higher-order genome organisation or context-dependent regulation. Together, our findings identify repeat-rich genome architecture as a dynamic feature of Annulohypoxylon genome evolution and provide a framework for testing how such regions may contribute to fungal environmental flexibility.