The introduction of models like RFDiffusionAA, AlphaFold3, AlphaProteo, and Chai1 has revolutionized protein structure modeling and interaction prediction, primarily from a binding perspective, focusing on creating ideal lock-and-key models. However, these methods can fall short for enzyme-substrate interactions, where perfect binding models are rare, and induced fit states are more common. To address this, we shift to a functional perspective for enzyme design, where the enzyme function is defined by the reaction it catalyzes. Here, we introduce \textsc{GENzyme}, a \textit{de novo} enzyme design model that takes a catalytic reaction as input and generates the catalytic pocket, full enzyme structure, and enzyme-substrate binding complex. \textsc{GENzyme} is an end-to-end, three-staged model that integrates (1) a catalytic pocket generation and sequence co-design module, (2) a pocket inpainting and enzyme inverse folding module, and (3) a binding and screening module to optimize and predict enzyme-substrate complexes. The entire design process is driven by the catalytic reaction being targeted. This reaction-first approach allows for more accurate and biologically relevant enzyme de
Microbial genomes and metagenomes contain millions of proteins whose enzymatic functions remain unknown, the enzyme dark matter. While deep learning has improved protein function prediction, most methods are black boxes relying on sequence or structural similarity, limiting discovery of novel catalytic activities. The ESMC-6B protein language model and its sparse autoencoder with a 16,384-dimensional codebook of interpretable biological concepts, each annotated by GPT-5, creates a new opportunity: using these features directly as semantic signatures for enzyme function. Here, we show that ESMC-SAE features enable accurate and interpretable enzyme commission (EC) number prediction without task-specific training or GPU-intensive computation. On a balanced benchmark of 4,868 microbial SwissProt enzymes across 161 EC3 subclasses, ESMC-SAE binary features achieve 78.9% top-1 and 88.5% top-5 accuracy, 37.6% higher than 3-mer baselines (57.3%). In leave-one-EC3-class-out evaluation simulating discovery of novel enzyme classes, SAE features recover the EC1 superclass in 47.7% of cases (3.3x random, 14.3%), versus 26.6% for sequence methods. Discriminative features correspond to mechanistic
Identifying enzymes that catalyze target biochemical reactions is a key step in computational enzyme discovery and biocatalyst design. Recent representation-learning methods formulate this problem as enzyme--reaction matching, where paired enzymes and reactions are embedded into a shared space. However, most existing approaches primarily rely on pairwise enzyme--reaction supervision and make limited use of the relationships within reaction sets or enzyme families. This work introduces a multi-alignment contrastive learning framework for biochemical retrieval. The framework jointly models cross-domain compatibility between enzymes and reactions and within-domain relationships induced by functional annotations. In addition, a Gromov--Wasserstein-inspired regularization objective encourages geometric consistency between the learned enzyme and reaction representation spaces. By combining pairwise catalytic supervision with higher-order relational alignment, the model captures both direct enzyme--reaction associations and broader functional organization. We evaluate the approach on enzyme virtual screening and bidirectional enzyme--reaction retrieval tasks. Experiments on EnzymeMap show
As belief around the potential of computational social science grows, fuelled by recent advances in machine learning, data scientists are ostensibly becoming the new experts in education. Scholars engaged in critical studies of education and technology have sought to interrogate the growing datafication of education yet tend not to use computational methods as part of this response. In this paper, we discuss the feasibility and desirability of the use of computational approaches as part of a critical research agenda. Presenting and reflecting upon two examples of projects that use computational methods in education to explore questions of equity and justice, we suggest that such approaches might help expand the capacity of critical researchers to highlight existing inequalities, make visible possible approaches for beginning to address such inequalities, and engage marginalised communities in designing and ultimately deploying these possibilities. Drawing upon work within the fields of Critical Data Studies and Science and Technology Studies, we further reflect on the two cases to discuss the possibilities and challenges of reimagining computational methods for critical research in
The enzyme turnover rate is a fundamental parameter in enzyme kinetics, reflecting the catalytic efficiency of enzymes. However, enzyme turnover rates remain scarce across most organisms due to the high cost and complexity of experimental measurements. To address this gap, we propose a multimodal framework for predicting the enzyme turnover rate by integrating enzyme sequences, substrate structures, and environmental factors. Our model combines a pre-trained language model and a convolutional neural network to extract features from protein sequences, while a graph neural network captures informative representations from substrate molecules. An attention mechanism is incorporated to enhance interactions between enzyme and substrate representations. Furthermore, we leverage symbolic regression via Kolmogorov-Arnold Networks to explicitly learn mathematical formulas that govern the enzyme turnover rate, enabling interpretable and accurate predictions. Extensive experiments demonstrate that our framework outperforms both traditional and state-of-the-art deep learning approaches. This work provides a robust tool for studying enzyme kinetics and holds promise for applications in enzyme e
Enzyme reactions are highly dependent on reaction conditions. To ensure reproducibility of enzyme reaction parameters, experiments need to be carefully designed and kinetic modelling meticulously executed. Furthermore, to enable the judgement of the quality of enzyme reaction parameters, the experimental conditions, the modelling process as well as the raw data need to be reported comprehensively. By taking these steps, enzyme reaction parameters can be open and FAIR (findable, accessible, interoperable, re-usable) as well as repeatable, replicable and reproducible. This review discusses these issues and provides a practical guide to designing initial rate experiments for the determination of enzyme reaction parameters and gives an open, FAIR and re-editable example of the kinetic modelling of an enzyme reaction. Both the guide and example are scripted with Python in Jupyter Notebooks and are publicly available (https://fairdomhub.org/investigations/483). Finally, the prerequisites of automated data analysis and machine learning algorithms are briefly discussed to provide further motivation for the comprehensive, open and FAIR reporting of enzyme reaction parameters.
Genome-scale stoichiometric modeling of metabolism has become a standard systems biology tool for modeling cellular physiology and growth. Extensions of this approach are also emerging as a valuable avenue for predicting, understanding and designing microbial communities. COMETS (Computation Of Microbial Ecosystems in Time and Space) was initially developed as an extension of dynamic flux balance analysis, which incorporates cellular and molecular diffusion, enabling simulations of multiple microbial species in spatially structured environments. Here we describe how to best use and apply the most recent version of this platform, COMETS 2, which incorporates a more accurate biophysical model of microbial biomass expansion upon growth, as well as several new biological simulation modules, including evolutionary dynamics and extracellular enzyme activity. COMETS 2 provides user-friendly Python and MATLAB interfaces compatible with the well-established COBRA models and methods, and comprehensive documentation and tutorials, facilitating the use of COMETS for researchers at all levels of expertise with metabolic simulations. This protocol provides a detailed guideline for installing, te
Expressing plant metabolic pathways in microbial platforms is an efficient, cost-effective solution for producing many desired plant compounds. As eukaryotic organisms, yeasts are often the preferred platform. However, expression of plant enzymes in a yeast frequently leads to failure because the enzymes are poorly adapted to the foreign yeast cellular environment. Here we first summarize current engineering approaches for optimizing performance of plant enzymes in yeast. A critical limitation of these approaches is that they are labor-intensive and must be customized for each individual enzyme, which significantly hinders the establishment of plant pathways in cellular factories. In response to this challenge, we propose the development of a cost-effective computational pipeline to redesign plant enzymes for better adaptation to the yeast cellular milieu. This proposition is underpinned by compelling evidence that plant and yeast enzymes exhibit distinct sequence features that are generalizable across enzyme families. Consequently, we introduce a data-driven machine learning framework designed to extract 'yeastizing' rules from natural protein sequence variations, which can be bro
Chemotaxis of enzymes in response to gradients in the concentration of their substrate has been widely reported in recent experiments, but a basic understanding of the process is still lacking. Here, we develop a microscopic theory for chemotaxis, valid for enzymes and other small molecules. Our theory includes both non-specific interactions between enzyme and substrate, as well as complex formation through specific binding between the enzyme and the substrate. We find that two distinct mechanisms contribute to enzyme chemotaxis: a diffusiophoretic mechanism due to the non-specific interactions, and a new type of mechanism due to binding-induced changes in the diffusion coefficient of the enzyme. The latter chemotactic mechanism points towards lower substrate concentration if the substrate enhances enzyme diffusion, and towards higher substrate concentration if the substrate inhibits enzyme diffusion. For a typical enzyme, attractive phoresis and binding-induced enhanced diffusion will compete against each other. We find that phoresis dominates above a critical substrate concentration, whereas binding-induced enhanced diffusion dominates for low substrate concentration. Our results
Enzymes, with their specific catalyzed reactions, are necessary for all aspects of life, enabling diverse biological processes and adaptations. Predicting enzyme functions is essential for understanding biological pathways, guiding drug development, enhancing bioproduct yields, and facilitating evolutionary studies. Addressing the inherent complexities, we introduce a new approach to annotating enzymes based on their catalyzed reactions. This method provides detailed insights into specific reactions and is adaptable to newly discovered reactions, diverging from traditional classifications by protein family or expert-derived reaction classes. We employ machine learning algorithms to analyze enzyme reaction datasets, delivering a much more refined view on the functionality of enzymes. Our evaluation leverages the largest enzyme-reaction dataset to date, derived from the SwissProt and Rhea databases with entries up to January 8, 2024. We frame the enzyme-reaction prediction as a retrieval problem, aiming to rank enzymes by their catalytic ability for specific reactions. With our model, we can recruit proteins for novel reactions and predict reactions in novel proteins, facilitating en
Quantum technology is an emergent and potentially disruptive discipline, with the ability to affect many human activities. Quantum technologies are dual-use technologies, and as such are of interest to the defence and security industry and military and governmental actors. This report reviews and maps the possible quantum technology military applications, serving as an entry point for international peace and security assessment, ethics research, military and governmental policy, strategy and decision making. Quantum technologies for military applications introduce new capabilities, improving effectiveness and increasing precision, thus leading to `quantum warfare', wherein new military strategies, doctrines, policies and ethics should be established. This report provides a basic overview of quantum technologies under development, also estimating the expected time scale of delivery or the utilisation impact. Particular military applications of quantum technology are described for various warfare domains (e.g. land, air, space, electronic, cyber and underwater warfare and ISTAR -- intelligence, surveillance, target acquisition and reconnaissance), and related issues and challenges ar
Microbial communities are ubiquitous in nature and come in a multitude of forms, ranging from communities dominated by a handful of species to communities containing a wide variety of metabolically distinct organisms. This huge range in diversity is not a curiosity - microbial diversity has been linked to outcomes of substantial ecological and medical importance. However, the mechanisms underlying microbial diversity are still under debate, as simple mathematical models only permit as many species to coexist as there are resources. A plethora of mechanisms have been proposed to explain the origins of microbial diversity, but many of these analyses omit a key property of real microbial ecosystems: the propensity of the microbes themselves to change their growth properties within and across generations. In order to explore the impact of this key property on microbial diversity, we expand upon a recently developed model of microbial diversity in fluctuating environments. We implement changes in growth strategy in two distinct ways. First, we consider the regulation of a cell's enzyme levels within short, ecological times, and second we consider evolutionary changes driven by mutations
Many enzymes appear to diffuse faster in the presence of substrate and to drift either up or down a concentration gradient of their substrate. Observations of these phenomena, termed enhanced enzyme diffusion (EED) and enzyme chemotaxis, respectively, lead to a novel view of enzymes as active matter. Enzyme chemotaxis and EED may be important in biology, and they could have practical applications in biotechnology and nanotechnology. They also are of considerable biophysical interest; indeed, their physical mechanisms are still quite uncertain. This review provides an analytic summary of experimental studies of these phenomena and of the mechanisms that have been proposed to explain them, and offers a perspective of future directions for the field.
The metabolic state of a cell, comprising fluxes, metabolite concentrations and enzyme levels, is shaped by a compromise between metabolic benefit and enzyme cost. This hypothesis and its consequences can be studied by computational models and using a theory of metabolic value. In optimal metabolic states, any increase of an enzyme level must improve the metabolic performance to justify its own cost, so each active enzyme must contribute to the cell's benefit by producing valuable products. This principle of value production leads to variation rules that relate metabolic fluxes and reaction elasticities to enzyme costs. Metabolic value theory provides a language to describe this. It postulates a balance of local values, which I derive here from concepts of metabolic control theory. Economic state variables, called economic potentials and loads, describe how metabolites, reactions, and enzymes contribute to metabolic performance. Economic potentials describe the indirect value of metabolite production, while economic loads describe the indirect value of metabolite concentrations. These economic variables, and others, are linked by local balance equations. These laws for optimal meta
High throughput sequencing (HTS)-based technology enables identifying and quantifying non-culturable microbial organisms in all environments. Microbial sequences have enhanced our understanding of the human microbiome, the soil and plant environment, and the marine environment. All molecular microbial data pose statistical challenges due to contamination sequences from reagents, batch effects, unequal sampling, and undetected taxa. Technical biases and heteroscedasticity have the strongest effects, but different strains across subjects and environments also make direct differential abundance testing unwieldy. We provide an introduction to a few statistical tools that can overcome some of these difficulties and demonstrate those tools on an example. We show how standard statistical methods, such as simple hierarchical mixture and topic models, can facilitate inferences on latent microbial communities. We also review some nonparametric Bayesian approaches that combine visualization and uncertainty quantification. The intersection of molecular microbial biology and statistics is an exciting new venue. Finally, we list some of the important open problems that would benefit from more ca
Filament formation by non-cytoskeletal enzymes has been known for decades, yet only relatively recently has its wide-spread role in enzyme regulation and biology come to be appreciated. This comprehensive review summarizes what is known for each enzyme confirmed to form filamentous structures in vitro, and for the many that are known only to form large self-assemblies within cells. For some enzymes, studies describing both the in vitro filamentous structures and cellular self-assembly formation are also known and described. Special attention is paid to the detailed structures of each type of enzyme filament, as well as the roles the structures play in enzyme regulation and in biology. Where it is known or hypothesized, the advantages conferred by enzyme filamentation are reviewed. Finally, the similarities, differences, and comparison to the SgrAI system are also highlighted.
Blockchain is an emerging digital technology allowing ubiquitous financial transactions among distributed untrusted parties, without the need of intermediaries such as banks. This article examines the impact of blockchain technology in agriculture and food supply chain, presents existing ongoing projects and initiatives, and discusses overall implications, challenges and potential, with a critical view over the maturity of these projects. Our findings indicate that blockchain is a promising technology towards a transparent supply chain of food, with many ongoing initiatives in various food products and food-related issues, but many barriers and challenges still exist, which hinder its wider popularity among farmers and systems. These challenges involve technical aspects, education, policies and regulatory frameworks.
Recent fluorescence spectroscopy measurements of the turnover time distribution of single-enzyme turnover kinetics of $β$-galactosidase provide evidence of Michaelis-Menten kinetics at low substrate concentration. However, at high substrate concentrations, the dimensionless variance of the turnover time distribution shows systematic deviations from the Michaelis-Menten prediction. This difference is attributed to conformational fluctuations in both the enzyme and the enzyme-substrate complex and to the possibility of both parallel and off-pathway kinetics. Here, we use the chemical master equation to model the kinetics of a single fluctuating enzyme that can yield a product through either parallel or off-pathway mechanisms. An exact expression is obtained for the turnover time distribution from which the mean turnover time and randomness parameters are calculated. The parallel and off-pathway mechanisms yield strikingly different dependences of the mean turnover time and the randomness parameter on the substrate concentration. In the parallel mechanism, the distinct contributions of enzyme and enzyme-substrate fluctuations are clearly discerned from the variation of the randomness
The interactions among the constituent members of a microbial community play a major role in determining the overall behavior of the community and the abundance levels of its members. These interactions can be modeled using a network whose nodes represent microbial taxa and edges represent pairwise interactions. A microbial network is a weighted graph that is constructed from a sample-taxa count matrix, and can be used to model co-occurrences and/or interactions of the constituent members of a microbial community. The nodes in this graph represent microbial taxa and the edges represent pairwise associations amongst these taxa. A microbial network is typically constructed from a sample-taxa count matrix that is obtained by sequencing multiple biological samples and identifying taxa counts. From large-scale microbiome studies, it is evident that microbial community compositions and interactions are impacted by environmental and/or host factors. Thus, it is not unreasonable to expect that a sample-taxa matrix generated as part of a large study involving multiple environmental or clinical parameters can be associated with more than one microbial network. However, to our knowledge, micr
Statistical analysis of distributions of occurrence frequencies of short words in 108 microbial complete genomes reveals the existence of a set of universal "root-sequence lengths" shared by all microbial genomes. These lengths and their universality give powerful clues to the way microbial genomes are grown. We show that the observed genomic properties are explained by a model for genome growth in which primitive genomes grew mainly by maximally stochastic duplications of short segments from an initial length of about 200 nucleotides (nt) to a length of about one million nt typical of microbial genomes. The relevance of the result of this study to the nature of simultaneous random growth and information acquisition by genomes, to the so-called RNA world in which life evolved before the rise of proteins and enzymes and to several other topics are discussed.