共找到 20 条结果
Deamination has historically been important for authenticating ancient biomolecules. However, expanding paleogenomic datasets indicate that damage patterns are more influenced by burial hydrology and microstructural context than by molecular age or ancestry. Fossils interact with their environments differently: some form closed, water-restricted compartments that preserve minimally damaged endogenous biomolecules, whereas others serve as open molecular reservoirs in which infiltrated environmental biomolecules undergo extensive deamination from repeated water exposure. Reliance on deamination alone can therefore suppress endogenous signals and complicate the interpretation of exogenous sequences. By introducing the molecular sedimentation model for fossil biomolecules, this Perspective outlines a source tracing framework that integrates fossil microstructure, ecological reference sets, and species-specific fragments to enable more reliable molecular inference across diverse depositional environments.
Geometric graph neural networks (GNNs) that respect E(3) symmetries have achieved strong performance on small molecule modeling, but they face scalability and expressiveness challenges when applied to large biomolecules such as RNA and proteins. These systems require models that can simultaneously capture fine-grained atomic interactions, long-range dependencies across spatially distant components, and biologically relevant hierarchical structure, such as atoms forming residues, which in turn form higher-order domains. Existing geometric GNNs, which typically operate exclusively in either Euclidean or Spherical Harmonics space, are limited in their ability to capture both the fine-scale atomic details and the long-range, symmetry-aware dependencies required for modeling the multi-scale structure of large biomolecules. We introduce DualEquiNet, a Dual-Space Hierarchical Equivariant Network that constructs complementary representations in both Euclidean and Spherical Harmonics spaces to capture local geometry and global symmetry-aware features. DualEquiNet employs bidirectional cross-space message passing and a novel Cross-Space Interaction Pooling mechanism to hierarchically aggrega
Label-free single-molecule detection is essential for studying biomolecules in their native state, yet having materials with intrinsic molecular specificity to identify a broad range of molecules without complex functionalization remains challenging. We present a method that utilizes emissions from selectively activated defects at the aqueous hexagonal boron nitride (hBN) interface to detect and identify biomolecules, including lipids, nucleotides, and amino acids. Using spectrally-resolved single-molecule localization microscopy combined with machine learning, we harvest spatial, spectral and temporal data of individual events to uncover optical fingerprints of biomolecules. This approach allows us to probe fine chemical differences, as small as the single deprotonation of an amino acid's side chain and detect dynamics of biomolecules at the interface with exceptional detail. As a proof-of-concept, we identified five different amino acids at the single-molecule level with high accuracy. Our findings shed light on hBN-biomolecule interactions and highlight the potential of hBN for label-free single-molecule identification.
First-line antiproliferatives for non-small cell lung cancer (NSCLC) have a relatively high failure rate due to high intrinsic resistance rates and acquired resistance rates to therapy. 57% patients are diagnosed in late-stage disease due to the tendency of early-stage NSCLC to be asymptomatic. For patients first diagnosed with metastatic disease the 5-year survival rate is approximately 5%. To help accelerate the development of novel therapeutics and computer-based tools for optimizing individual therapy, we have collated data from 11 different clinical trials in NSCLC and developed a semi-mechanistic, clinical model of NSCLC growth and pharmacodynamics relative to the various therapeutics represented in the study. In this study, we have produced extremely precise estimates of clinical parameters fundamental to cancer modeling such as the rate of acquired resistance to various pharmaceuticals, the relationship between drug concentration and rate of cancer cell death, as well as the fine temporal dynamics of anti-VEGF therapy. In the simulation sets documented in this study, we have used the model to make meaningful descriptions of efficacy gain in making bevacizumab-antiproliferat
Therapeutic development is a costly and high-risk endeavor that is often plagued by high failure rates. To address this, we introduce TxGemma, a suite of efficient, generalist large language models (LLMs) capable of therapeutic property prediction as well as interactive reasoning and explainability. Unlike task-specific models, TxGemma synthesizes information from diverse sources, enabling broad application across the therapeutic development pipeline. The suite includes 2B, 9B, and 27B parameter models, fine-tuned from Gemma-2 on a comprehensive dataset of small molecules, proteins, nucleic acids, diseases, and cell lines. Across 66 therapeutic development tasks, TxGemma achieved superior or comparable performance to the state-of-the-art generalist model on 64 (superior on 45), and against state-of-the-art specialist models on 50 (superior on 26). Fine-tuning TxGemma models on therapeutic downstream tasks, such as clinical trial adverse event prediction, requires less training data than fine-tuning base LLMs, making TxGemma suitable for data-limited applications. Beyond these predictive capabilities, TxGemma features conversational models that bridge the gap between general LLMs an
Macromolecular crowding affects biophysical processes as diverse as diffusion, gene expression, cell growth, and senescence. Yet, there is no comprehensive understanding of how crowding affects reactions, particularly multivalent binding. Herein, we use scaled particle theory and develop a molecular simulation method to investigate the binding of monovalent to divalent biomolecules. We find that crowding can increase or reduce cooperativity--the extent to which the binding of a second molecule is enhanced after binding a first molecule--by orders of magnitude, depending on the sizes of the involved molecular complexes. Cooperativity generally increases when a divalent molecule swells and then shrinks upon binding two ligands. Our calculations also reveal that, in some cases, crowding enables binding that does not occur otherwise. As an immunological example, we consider Immunoglobulin G-antigen binding and show that crowding enhances its cooperativity in bulk but reduces it when an Immunoglobulin G binds antigens on a surface.
Diffusion probabilistic models have made their way into a number of high-profile applications since their inception. In particular, there has been a wave of research into using diffusion models in the prediction and design of biomolecular structures and sequences. Their growing ubiquity makes it imperative for researchers in these fields to understand them. This paper serves as a general overview for the theory behind these models and the current state of research. We first introduce diffusion models and discuss common motifs used when applying them to biomolecules. We then present the significant outcomes achieved through the application of these models in generative and predictive tasks. This survey aims to provide readers with a comprehensive understanding of the increasingly critical role of diffusion models.
Developing therapeutics is a lengthy and expensive process that requires the satisfaction of many different criteria, and AI models capable of expediting the process would be invaluable. However, the majority of current AI approaches address only a narrowly defined set of tasks, often circumscribed within a particular domain. To bridge this gap, we introduce Tx-LLM, a generalist large language model (LLM) fine-tuned from PaLM-2 which encodes knowledge about diverse therapeutic modalities. Tx-LLM is trained using a collection of 709 datasets that target 66 tasks spanning various stages of the drug discovery pipeline. Using a single set of weights, Tx-LLM simultaneously processes a wide variety of chemical or biological entities(small molecules, proteins, nucleic acids, cell lines, diseases) interleaved with free-text, allowing it to predict a broad range of associated properties, achieving competitive with state-of-the-art (SOTA) performance on 43 out of 66 tasks and exceeding SOTA on 22. Among these, Tx-LLM is particularly powerful and exceeds best-in-class performance on average for tasks combining molecular SMILES representations with text such as cell line names or disease names
The understanding of protein structure, folding, and interaction with other proteins remains one of the grand challenges of modern biology. Tremendous progress has been made thanks to X-ray- or electron-based techniques that have provided atomic configurations of proteins, and their solvation shell. These techniques though require a large number of similar molecules to provide an average view, and lack detailed compositional information that might play a major role in the biochemical activity of these macromolecules. Based on its intrinsic performance and recent impact in materials science, atom probe tomography (APT) has been touted as a potential novel tool to analyse biological materials, including proteins. However, analysis of biomolecules in their native, hydrated state by APT have not yet been routinely achieved, and the technique's true capabilities remain to be demonstrated. Here, we present and discuss systematic analyses of individual amino-acids in frozen aqueous solutions on two different nanoporous metal supports across a wide range of analysis conditions. Using a ratio of the molecular ions of water as a descriptor for the conditions of electrostatic field, we study
The integration of biomolecular modeling with natural language (BL) has emerged as a promising interdisciplinary area at the intersection of artificial intelligence, chemistry and biology. This approach leverages the rich, multifaceted descriptions of biomolecules contained within textual data sources to enhance our fundamental understanding and enable downstream computational tasks such as biomolecule property prediction. The fusion of the nuanced narratives expressed through natural language with the structural and functional specifics of biomolecules described via various molecular modeling techniques opens new avenues for comprehensively representing and analyzing biomolecules. By incorporating the contextual language data that surrounds biomolecules into their modeling, BL aims to capture a holistic view encompassing both the symbolic qualities conveyed through language as well as quantitative structural characteristics. In this review, we provide an extensive analysis of recent advancements achieved through cross modeling of biomolecules and natural language. (1) We begin by outlining the technical representations of biomolecules employed, including sequences, 2D graphs, and
How life started on Earth is an unsolved mystery. There are various hypotheses for the location ranging from outer space to the seafloor, subseafloor or potentially deeper. Here, we applied extensive ab initio molecular dynamics (AIMD) simulations to study chemical reactions between NH$_3$, H$_2$O, H$_2$, and CO at pressures (P) and temperatures (T) approximating the conditions of Earth's upper mantle (i.e. 10-13 GPa, 1000-1400 K). Contrary to the previous assumptions that larger organic molecules might readily disintegrate in aqueous solutions at extreme P-T conditions, we found that many organic compounds formed without any catalysts and persisted in C-H-O-N fluids under these extreme conditions, including glycine, ribose, urea, and uracil-like molecules. Particularly, our free energy calculations showed that the C-N bond is thermodynamically stable at 10 GPa and 1400 K. Moreover, while the pyranose (six-membered-ring) form of ribose is more stable than the furanose (five-membered-ring) form at ambient conditions, we observed the predominant formation of the five-membered-ring form of ribose at extreme conditions, which is consistent with the exclusive incorporation of $β$-D-ribo
Structural transition induced by a local conformational change in biomolecules is formulated based on the generalized Langevin theory for the structural fluctuation of a molecule in solution, and the linear response theory, derived by Kim and Hirata in 2012. A chemical/mechanical change introduced at a moiety of biomolecules, such as an amino acid substitution or a structural change of a chromophore by the photo-excitation, is considered as a perturbation, and the rest of the protein as the reference system. The linear-response equation consists of two parts: a mechanical/chemical perturbation introduced at the moiety, and the variance-covariance matrix of the reference system that works as a response function. The physical meaning of the equation is transparent: the force exerted by atoms in the moiety induces the displacement in an atom of protein, which propagates through the variance-covariance matrix to cause a global conformational change in the molecule. A few examples of possible application of the theory, including those in industry, are suggested.
Molecular simulations are essential tools in computational chemistry, enabling the prediction and understanding of molecular interactions and thermodynamic properties of biomolecules. However, traditional force fields face significant challenges in accurately representing novel molecules and complex chemical environments due to the labor-intensive process of manually setting optimization parameters and the high computational cost of quantum mechanical calculations. To overcome these difficulties, we fine-tuned a high-accuracy DPA-2 pre-trained model and applied it to optimize force field parameters on-the-fly, significantly reducing computational costs. Our method combines this fine-tuned DPA-2 model with a node-embedding-based similarity metric, allowing seamless augmentation to new chemical species without manual intervention. We applied this process to the TYK2 inhibitor and PTP1B systems and demonstrated its effectiveness through the improvement of free energy perturbation calculation results. This advancement contributes valuable insights and tools for the computational chemistry community.
Generators of space-time dynamics in bioimaging have become essential to build ground truth datasets for image processing algorithm evaluation such as biomolecule detectors and trackers, as well as to generate training datasets for deep learning algorithms. In this contribution, we leverage a stochastic model, called birth-death-move (BDM) point process, in order to generate joint dynamics of biomolecules in cells. This approach is very flexible and allows us to model a system of particles in motion, possibly in interaction, that can each possibly switch from a motion regime (e.g. Brownian) to another (e.g. a directed motion), along with the appearance over time of new trajectories and their death after some lifetime, all of these features possibly depending on the current spatial configuration of all existing particles. We explain how to specify all characteristics of a BDM model, with many practical examples that are relevant for bioimaging applications. Based on real fluorescence microscopy datasets, we finally calibrate our model to mimic the joint dynamics of Langerin and Rab11 proteins near the plasma membrane. We show that the resulting synthetic sequences exhibit comparable
Label-free optical absorption microscopy techniques continue to evolve as promising tools for label-free histopathological imaging of cells and tissues. However, critical challenges relating to specificity and contrast, as compared to current gold-standard methods continue to hamper adoption. This work introduces Photon Absorption Remote Sensing (PARS), a new absorption microscope modality, which simultaneously captures the dominant de-excitation processes following an absorption event. In PARS, radiative (auto-fluorescence) and non-radiative (photothermal and photoacoustic) relaxation processes are collected simultaneously, providing enhanced specificity to a range of biomolecules. As an example, a multiwavelength PARS system featuring UV (266 nm) and visible (532 nm) excitation is applied to imaging human skin, and murine brain tissue samples. It is shown that PARS can directly characterize, differentiate, and unmix, clinically relevant biomolecules inside complex tissues samples using established statistical processing methods. Gaussian mixture models (GMM) are used to characterize clinically relevant biomolecules (e.g., white, and gray matter) based on their PARS signals, while
Digit therapeutics are novel software devices that clinicians may utilize in delivering quality mental health care and ensuring positive outcomes. However, uptake of digital therapeutics and clinically tested software-based programs remains low. This article presents possible reasons for attrition and low engagement in clinical studies investigating digital therapeutics, analyses of studies in which engagement was high, and design constructs that may encourage user engagement. The aim is to shed light on the importance of real-world attrition data of digital therapeutics, and important characteristics of medical devices that have positively influenced user engagement. The findings presented in this article will be useful to relevant stakeholders and medical device experts tasked with addressing the gap between software medical design and user engagement present in digital therapeutic clinical trials.
Therapeutics machine learning is an emerging field with incredible opportunities for innovatiaon and impact. However, advancement in this field requires formulation of meaningful learning tasks and careful curation of datasets. Here, we introduce Therapeutics Data Commons (TDC), the first unifying platform to systematically access and evaluate machine learning across the entire range of therapeutics. To date, TDC includes 66 AI-ready datasets spread across 22 learning tasks and spanning the discovery and development of safe and effective medicines. TDC also provides an ecosystem of tools and community resources, including 33 data functions and types of meaningful data splits, 23 strategies for systematic model evaluation, 17 molecule generation oracles, and 29 public leaderboards. All resources are integrated and accessible via an open Python library. We carry out extensive experiments on selected datasets, demonstrating that even the strongest algorithms fall short of solving key therapeutics challenges, including real dataset distributional shifts, multi-scale modeling of heterogeneous data, and robust generalization to novel data points. We envision that TDC can facilitate algor
We demonstrate the fabrication of sharp nanopillars of high aspect ratio onto specialized atomic force microscopy (AFM) microcantilevers and their use for high-speed AFM of DNA and nucleoproteins in liquid. The fabrication technique uses localized charged-particle-induced deposition with either a focused beam of helium ions or electrons in a helium ion microscope (HIM) or scanning electron microscope (SEM). This approach enables customized growth onto delicate substrates with nanometer-scale placement precision and in-situ imaging of the final tip structures using the HIM or SEM. Tip radii of <10 nm are obtained and the underlying microcantilever remains intact. Instead of the more commonly used organic precursors employed for bio-AFM applications, we use an organometallic precursor (tungsten hexacarbonyl) resulting in tungsten-containing tips. Transmission electron microscopy reveals a thin layer of carbon on the tips. Consequently, the interaction of the new tips with biological specimens is likely very similar to that of standard carbonaceous tips, with the added benefit of robustness. A further advantage of the organometallic tips is that compared to carbonaceous tips they b
Biomolecules are the prime information processing elements of living matter. Most of these inanimate systems are polymers that compute their structures and dynamics using as input seemingly random character strings of their sequence, following which they coalesce and perform integrated cellular functions. In large computational systems with a finite interaction-codes, the appearance of conflicting goals is inevitable. Simple conflicting forces can lead to quite complex structures and behaviors, leading to the concept of "frustration" in condensed matter. We present here some basic ideas about frustration in biomolecules and how the frustration concept leads to a better appreciation of many aspects of the architecture of biomolecules, and how structure connects to function. These ideas are simultaneously both seductively simple and perilously subtle to grasp completely. The energy landscape theory of protein folding provides a framework for quantifying frustration in large systems and has been implemented at many levels of description. We first review the notion of frustration from the areas of abstract logic and its uses in simple condensed matter systems. We discuss then how the f
We give a theoretical treatment of the interaction of electronic excitations (excitons) in biomolecules and quantum dots with the surrounding polar solvent. Significant quantum decoherence occurs due to the interaction of the electric dipole moment of the solute with the fluctuating electric dipole moments of the individual molecules in the solvent. We introduce spin boson models which could be used to describe the effects of decoherence on the quantum dynamics of biomolecules which undergo light-induced conformational change and on biomolecules or quantum dots which are coupled by Forster resonant energy transfer.