A tumor often consists of multiple cell subpopulations (clones). Current chemo-treatments often target one clone of a tumor. Although the drug kills that clone, other clones overtake it and the tumor reoccurs. Genome sequencing and computational analysis allows to computational dissection of clones from tumors, while singe-cell genome sequencing including RNA-Seq allows to profiling of these clones. This opens a new window for treating a tumor as a system in which clones are evolving. Future cancer systems biology studies should consider a tumor as an evolving system with multiple clones. Therefore, topics discussed in Part 2 of this review include evolutionary dynamics of clonal networks, early-warning signals for formation of fast-growing clones, dissecting tumor heterogeneity, and modeling of clone-clone-stroma interactions for drug resistance. The ultimate goal of the future systems biology analysis is to obtain a whole-system understanding of a tumor and therefore provides a more efficient and personalized management strategies for cancer patients.
Recent tumor genome sequencing confirmed that one tumor often consists of multiple cell subpopulations (clones) which bear different, but related, genetic profiles such as mutation and copy number variation profiles. Thus far, one tumor has been viewed as a whole entity in cancer functional studies. With the advances of genome sequencing and computational analysis, we are able to quantify and computationally dissect clones from tumors, and then conduct clone-based analysis. Emerging technologies such as single-cell genome sequencing and RNA-Seq could profile tumor clones. Thus, we should reconsider how to conduct cancer systems biology studies in the genome sequencing era. We will outline new directions for conducting cancer systems biology by considering that genome sequencing technology can be used for dissecting, quantifying and genetically characterizing clones from tumors. Topics discussed in Part 1 of this review include computationally quantifying of tumor subpopulations; clone-based network modeling, cancer hallmark-based networks and their high-order rewiring principles and the principles of cell survival networks of fast-growing clones.
A model of genome evolution is proposed. Based on three assumptions the evolutionary theory of a genome is formulated. The general law on the direction of genome evolution is given. Both the deterministic classical equation and the stochastic quantum equation are proposed. It is proved that the classical equation can be put in a form of the least action principle and the latter can be used for obtaining the quantum generalization of the evolutionary law. The wave equation and uncertainty relation for the quantum evolution are deduced logically. It is shown that the classical trajectory is a limiting case of the general quantum evolution depicted in the coarse-grained time. The observed smooth/sudden evolution is interpreted by the alternating occurrence of the classical and quantum phases. The speciation event is explained by the quantum transition in quantum phase. Fundamental constants of time dimension, the quantization constant and the evolutionary inertia, are introduced for characterizing the genome evolution. The size of minimum genome is deduced from the quantum uncertainty lower bound. The present work shows the quantum law may be more general than thought, since it plays
Genome sizes have evolved to vary widely, from 250 bases in viroids to 670 billion bases in amoeba. This remarkable variation in genome size is the outcome of complex interactions between various evolutionary factors such as point mutation rate, population size, insertions and deletions, and genome editing mechanisms that may be specific to certain taxonomic lineages. While comparative genomics analyses have uncovered some of the relationships between these diverse evolutionary factors, we still do not understand what drives genome size evolution. Specifically, it is not clear how primordial mutational processes of base substitutions, insertions, and deletions influence genome size evolution in asexual organisms. Here, we use digital evolution to investigate genome size evolution by tracking genome edits and their fitness effects in real time. In agreement with empirical data, we find that mutation rate is inversely correlated with genome size in asexual populations. We show that at low point mutation rate, insertions are significantly more beneficial than deletions, driving genome expansion and acquisition of phenotypic complexity. Conversely, high mutational load experienced at h
In our previous studies, we developed discrete-space Birth, Death and Innovation Models (BDIM) of genome evolution. These models explain the origin of the characteristic Pareto distribution of paralogous gene family sizes in genomes, and model parameters that provide for the evolution of these distributions within a realistic timeframe have been identified. Here we develop the diffusion version of BDIM whose dynamics is described by the Fokker-Plank equation and the stationary solution could be any specified Pareto function. The diffusion models have time-dependent solutions of a special kind, namely, the generalized self-similar solutions, which describe the transition from one stationary distribution of the system to another; this provides for the possibility of examining the temporal dynamics of genome evolution. Analysis of the generalized self-similar solutions of the diffusion BDIM reveals a biphasic curve of genome growth in which the initial, relatively short, self-accelerating phase is followed by a prolonged phase of slow deceleration. In biological terms, this regime of evolution can be tentatively interpreted as a punctuated-equilibrium-like phenomenon such that whereby
Research in quantitative evolutionary genomics and systems biology led to the discovery of several universal regularities connecting genomic and molecular phenomic variables. These universals include the log-normal distribution of the evolutionary rates of orthologous genes; the power law-like distributions of paralogous family size and node degree in various biological networks; the negative correlation between a gene's sequence evolution rate and expression level; and differential scaling of functional classes of genes with genome size. The universals of genome evolution can be accounted for by simple mathematical models similar to those used in statistical physics, such as the birth-death-innovation model. These models do not explicitly incorporate selection, therefore the observed universal regularities do not appear to be shaped by selection but rather are emergent properties of gene ensembles. Although a complete physical theory of evolutionary biology is inconceivable, the universals of genome evolution might qualify as 'laws of evolutionary genomics' in the same sense 'law' is understood in modern physics.
In this article, I put forward the idea that the neoplastic process (NP) has deep evolutionary roots and make specific predictions about the connection between cancer and the formation of the first embryo, which allowed for the evolutionary radiation of metazoans. My main hypothesis is that the NP is at the heart of cellular mechanisms responsible for animal morphogenesis and, given its embryological basis, also at the center of animal evolution. It is thus understood that NP-associated mechanisms are deeply rooted in evolutionary history and tied to the formation of the first animal embryo. In my consideration of these arguments, I expound on how cancer biology is perfectly intertwined with evolutionary biology. I describe essential cellular components of unicellular holozoans that served as a basis for the formation of the neoplastic functional module (NFM) and its subsequent exaptation, which brought forth two great biophysical revolutions within the first embryo. Finally, I examine the role of Physics in the modeling of the NFM and its contribution to morphogenesis to reveal the totipotency of the zygote.
This article frames the relation between biology and physics by characterizing the former as a subdiscipline rather than a special case of the latter. To do this, we posit biological physics as the science of living matter in contrast to classic biophysics, the study of organismal properties by physical techniques. At the scale of the individual cell, living matter is nonunitary, i.e., not composed of aggregated subunits, and has features (e.g., intracellular organizational arrangements and biomolecular condensates) that are unlike any materials of the nonliving world. In transiently or constitutively multicellular forms (social microorganisms, animals, plants), living matter sustains physical processes that are generic (shared with nonliving matter, e.g., subunit communication by molecular diffusion in cellular slime molds), biogeneric (analogous to nonliving matter but realized through cellular activities, e.g., subunit demixing in animal embryos) or nongeneric (pertaining to sui generis materials, e.g., budding of active solids in plants). This "forms of matter" perspective is philosophically situated in the dialectical materialism of Engels and Hessen and the multilevel physica
One way to understand the role history plays on evolutionary trajectories is by giving ancient life a second opportunity to evolve. Our ability to empirically perform such an experiment, however, is limited by current experimental designs. Combining ancestral sequence reconstruction with synthetic biology allows us to resurrect the past within a modern context and has expanded our understanding of protein functionality within a historical context. Experimental evolution, on the other hand, provides us with the ability to study evolution in action, under controlled conditions in the laboratory. Here we describe a novel experimental setup that integrates two disparate fields - ancestral sequence reconstruction and experimental evolution. This allows us to rewind and replay the evolutionary history of ancient biomolecules in the laboratory. We anticipate that our combination will provide a deeper understanding of the underlying roles that contingency and determinism play in shaping evolutionary processes.
Understanding the biological mechanisms of disease is crucial for medicine, and in particular, for drug discovery. AI-powered analysis of genome-scale biological data holds great potential in this regard. The increasing availability of single-cell RNA sequencing data has enabled the development of large foundation models for disease biology. However, existing foundation models only modestly improve over task-specific models in downstream applications. Here, we explored two avenues for improving single-cell foundation models. First, we scaled the pre-training data to a diverse collection of 116 million cells, which is larger than those used by previous models. Second, we leveraged the availability of large-scale biological annotations as a form of supervision during pre-training. We trained the \model family of models comprising six transformer-based state-of-the-art single-cell foundation models with 70 million, 160 million, and 400 million parameters. We vetted our models on several downstream evaluation tasks, including identifying the underlying disease state of held-out donors not seen during training, distinguishing between diseased and healthy cells for disease conditions and
We survey and introduce concepts and tools located at the intersection of information theory and network biology. We show that Shannon's information entropy, compressibility and algorithmic complexity quantify different local and global aspects of synthetic and biological data. We show examples such as the emergence of giant components in Erdos-Renyi random graphs, and the recovery of topological properties from numerical kinetic properties simulating gene expression data. We provide exact theoretical calculations, numerical approximations and error estimations of entropy, algorithmic probability and Kolmogorov complexity for different types of graphs, characterizing their variant and invariant properties. We introduce formal definitions of complexity for both labeled and unlabeled graphs and prove that the Kolmogorov complexity of a labeled graph is a good approximation of its unlabeled Kolmogorov complexity and thus a robust definition of graph complexity.
Darwin introduced the concept of the "living fossil" to describe species belonging to lineages that have experienced little evolutionary change, and suggested that species in more slowly evolving lineages are more prone to extinction (1). Recent studies revealed that some living fossils such as the lungfish are indeed evolving more slowly than other vertebrates (2, 3). The reason for the slower rate of evolution in these lineages remains unclear, but the same observations suggest a possible genome size effect on rates of evolution. Genome size (C-value) in vertebrates varies over 200 fold ranging from pufferfish (0.4 pg) to lungfish (132.8 pg) (4). Variation in genome size and architecture is a fundamental cellular adaptation that remains poorly understood (5). C-value is correlated with several allometric traits such as body size and developmental rates in many, but not all, organisms (6, 7). To date, no consensus exists concerning the mechanisms driving genome size evolution or the effect that genome size has on species traits such as evolutionary rates (8-12). In the following we show that: 1) within the same range of divergence times, genetic diversity decreases as genome size
We test whether artificial intelligence architectural evolution obeys the same statistical laws as biological evolution. Compiling 935 ablation experiments from 161 publications, we show that the distribution of fitness effects (DFE) of architectural modifications follows a heavy-tailed Student's t-distribution with proportions (68% deleterious, 19% neutral, 13% beneficial for major ablations, n=568) that place AI between compact viral genomes and simple eukaryotes. The DFE shape matches D. melanogaster (normalized KS=0.07) and S. cerevisiae (KS=0.09); the elevated beneficial fraction (13% vs. 1-6% in biology) quantifies the advantage of directed over blind search while preserving the distributional form. Architectural origination follows logistic dynamics (R^2=0.994) with punctuated equilibria and adaptive radiation into domain niches. Fourteen architectural traits were independently invented 3-5 times, paralleling biological convergences. These results demonstrate that the statistical structure of evolution is substrate-independent, determined by fitness landscape topology rather than the mechanism of selection.
Background: Transposed elements (TEs) have a substantial impact on mammalian evolution and are involved in numerous genetic diseases. We compared the impact of TEs on the human transcriptome and the mouse transcriptome. Results: We compiled a dataset of all TEs in the human and mouse genomes, identifying 3,932,058 and 3,122,416 TEs, respectively. We than extracted TEs located within human and mouse genes and, surprisingly, we found that 60% of TEs in both human and mouse are located in intronic sequences, even though introns comprise only 24% of the human genome. All TE families in both human and mouse can exonize. TE families that are shared between human and mouse exhibit the same percentage of TE exonization in the two species, but the exonization level of Alu, a primatespecific retroelement, is significantly greater than that of other TEs within the human genome, leading to a higher level of TE exonization in human than in mouse (1,824 exons compared with 506 exons, respectively). We detected a primate-specific mechanism for intron gain, in which Alu insertion into an exon creates a new intron located in the 3' untranslated region (termed 'intronization'). Finally, the insertio
Molecular clock (MC) is a central concept of molecular evolution according to which each gene evolves at a characteristic, near constant rate. Numerous evolutionary studies have demonstrated the validity of MC but also have shown that MC is substantially overdispersed, i.e. lineage-specific deviations of the evolutionary rate of the given gene from the clock greatly exceed the expectation from the sampling error. A fundamental observation of comparative genomics that appears to complement the MC is that the distribution of evolution rates across orthologous genes in pairs of related genomes remains virtually unchanged throughout the evolution of life, from bacteria to mammals. The conservation of this distribution implies that the relative evolution rates of all genes remain nearly constant, or in other words, that evolutionary rates of different genes are strongly correlated within each evolving genome. We hypothesized that this correlation is not a simple consequence of MC but could be better explained by a model we dubbed Universal PaceMaker (UPM) of genome evolution. The UPM model posits that the rate of evolution changes synchronously across genome-wide sets of genes in all ev
We developed a theory showing that under appropriate normalizations and rescalings, temperature response curves show a remarkably regular behavior and follow a general, universal law. The impressive universality of temperature response curves remained hidden due to various curve-fitting models not well-grounded in first principles. In addition, this framework has the potential to explain the origin of different scaling relationships in thermal performance in biology, from molecules to ecosystems. Here, we summarize the background, principles and assumptions, predictions, implications, and possible extensions of this theory.
The impact of climate change on populations will be contingent upon their contemporary adaptive evolution. In this study, we investigated the contemporary evolution of four populations of the cold-water kelp Laminaria digitata by analysing their spatial and temporal genomic variation using ddRAD-sequencing. These populations were sampled from the center to the southern margin of its north-eastern Atlantic distribution at two-time points, spanning at least two generations. Through genome scans for local adaptation at a single time point, we identified candidate loci that showed clinal variation correlated with changes in sea surface temperature (SST) along latitudinal gradients. This finding suggests that SST may drive the adaptive response of these kelp populations, although factors such as species' demographic history should also be considered. Additionally, we performed a simulation approach to distinguish the effect of selection from genetic drift in allele frequency changes over time. This enabled the detection of loci in the southernmost population that exhibited temporal differentiation beyond what would be expected from genetic drift alone: these are candidate loci which cou
In recent years, Whole Genome Sequencing (WGS) evolved from a futuristic-sounding research project to an increasingly affordable technology for determining complete genome sequences of complex organisms, including humans. This prompts a wide range of revolutionary applications, as WGS promises to improve modern healthcare and provide a better understanding of the human genome -- in particular, its relation to diseases and response to treatments. However, this progress raises worrisome privacy and ethical issues, since, besides uniquely identifying its owner, the genome contains a treasure trove of highly personal and sensitive information. In this article, after summarizing recent advances in genomics, we discuss some important privacy issues associated with human genomic information and identify a number of particularly relevant research challenges.
Rearrangements of bacterial chromosomes can be studied mathematically at several levels, most prominently at a local, or sequence level, as well as at a topological level. The biological changes involved locally are inversions, deletions, and transpositions, while topologically they are knotting and catenation. These two modelling approaches share some surprising algebraic features related to braid groups and Coxeter groups. The structural approach that is at the core of algebra has long found applications in sciences such as physics and analytical chemistry, but only in a small number of ways so far in biology. And yet there are examples where an algebraic viewpoint may capture a deeper structure behind biological phenomena. This article discusses a family of biological problems in bacterial genome evolution for which this may be the case, and raises the prospect that the tools developed by algebraists over the last century might provide insight to this area of evolutionary biology. .
Mapping genotypes to phenotypes (G2P) is a fundamental goal in biology. So called PhyloG2P methods are a relatively new set of tools that leverage replicated evolution in phylogenetically independent lineages to identify genomic regions associated with traits of interest. Here, we review recent developments in PhyloG2P methods, focusing on three key areas: methods based on replicated amino acid substitutions, methods detecting changes in evolutionary rates, and methods analysing gene duplication and loss. We discuss how the definition and measurement of traits impacts the utility of these methods, arguing that focusing on simple rather than compound traits will lead to more meaningful genotype-phenotype associations. We advocate for the use of methods that work with continuous traits directly rather than collapsing them to binary representations. We examine the strengths and limitations of different approaches to modeling genetic replication, highlighting the importance of explicit modeling of evolutionary processes. Finally, we outline promising future directions, including the integration of population-level variation, as well as epigenetic and environmental information. No one m