Nucleic acids such as mRNA have emerged as a promising therapeutic modality with the capability of addressing a wide range of diseases. Lipid nanoparticles (LNPs) as a delivery platform for nucleic acids were used in the COVID-19 vaccines and have received much attention. While modern manufacturing processes which involve rapidly mixing an organic stream containing the lipids with an aqueous stream containing the nucleic acids are conceptually straightforward, detailed understanding of LNP formation and structure is still limited and scale-up can be challenging. Mathematical and computational methods are a promising avenue for deepening scientific understanding of the LNP formation process and facilitating improved process development and control. This article describes strategies for the mechanistic modeling of LNP formation, starting with strategies to estimate and predict important physicochemical properties of the various species such as diffusivities and solubilities. Subsequently, a framework is outlined for constructing mechanistic models of reactor- and particle-scale processes. Insights gained from the various models are mapped back to product quality attributes and proces
Nucleic acids are increasingly recognized as therapeutic targets beyond conventional protein-centered drug discovery, yet accurate and efficient docking of small molecules to nucleic acid structures remains challenging. Physics-based docking methods often show limited accuracy and efficiency, whereas deep learning approaches are constrained by the scarcity of experimentally resolved nucleic acid-ligand complexes. Here, we present NucleoDock, a deep learning framework for nucleic acid-small molecule docking. To address data scarcity, NucleoDock combines physics-guided large-scale pretraining on millions of docking-generated synthetic complexes with fine-tuning on curated experimental co-crystal structures. It further integrates sequence- and structure-informed nucleotide representations with atomistic three-dimensional features to capture both biological context and binding-site geometry. A mixture density network-based geometric scoring head is used to model conditional interaction-distance distributions for pose ranking. On an external benchmark of 125 nucleic acid-ligand complexes, NucleoDock achieved a top-1 success rate of 56 percent at an RMSD cutoff of 2.0 Angstrom, outperfor
An accurate prediction of protein-nucleic acid binding affinity is vital for deciphering genomic processes, yet existing approaches often struggle in reconciling high accuracy with interpretability and computational efficiency. In this study, we introduce commutative algebra prediction (CAP), which couples persistent Stanley-Reisner theory with advanced sequence embedding for predicting protein-nucleic acid binding affinities. CAP encodes proteins through transformer-learned embeddings that retain long-range evolutionary context and represents DNA and RNA with $\textit{k}$-mer algebra embeddings derived from persistent facet ideals, which capture fine-scale nucleotide geometry. We demonstrate that CAP surpasses the SVSBI protein-nucleic acid benchmark and, in a further test, maintains reasonable performance on newly curated protein-RNA and protein-nucleic acid datasets. Leveraging only primary sequences, CAP generalizes to any protein-nucleic acid pair with minimal preprocessing, enabling genome-scale analyses without 3D structural data and promising faster virtual screening for drug discovery and protein engineering.
The interaction between proteins and nucleic acids is crucial for processes that sustain cellular function, including DNA maintenance and the regulation of gene expression and translation. Amino acid mutations in protein-nucleic acid complexes often lead to vital diseases. Experimental techniques have their own specific limitations in predicting mutational effects in protein-nucleic acid complexes. In this study, we compiled a large dataset of 1951 mutations including both protein-DNA and protein-RNA complexes and integrated structural and sequential features to build a deep learning-based regression model named DeepPNI. This model estimates mutation-induced binding free energy changes in protein-nucleic acid complexes. The structural features are encoded via edge-aware RGCN and the sequential features are extracted using protein language model ESM-2. We have achieved a high average Pearson correlation coefficient (PCC) of 0.76 in the large dataset via five-fold cross-validation. Consistent performance across individual dataset of protein-DNA, protein-RNA complexes, and different experimental temperature split dataset make the model generalizable. Our model showed good performance
Understanding the flexibility of protein-nucleic acid complexes, often characterized by atomic B-factors, is essential for elucidating their structure, dynamics, and functions, such as reactivity and allosteric pathways. Traditional models such as Gaussian Network Models (GNM) and Elastic Network Models (ENM) often fall short in capturing multiscale interactions, especially in large or complex biomolecular systems. In this work, we apply the Persistent Sheaf Laplacian (PSL) framework for the B-factor prediction of protein-nucleic acid complexes. The PSL model integrates multiscale analysis, algebraic topology, combinatoric Laplacians, and sheaf theory for data representation. It reveals topological invariants in its harmonic spectra and captures the homotopic shape evolution of data with its non-harmonic spectra. Its localization enables accurate B-factor predictions. We benchmark our method on three diverse datasets, including protein-RNA and nucleic-acid-only structures, and demonstrate that PSL consistently outperforms existing models such as GNM and multiscale FRI (mFRI), achieving up to a 21% improvement in Pearson correlation coefficient for B-factor prediction. These results
The conformational dynamics of single-stranded nucleic acids are fundamental for nucleic acid folding and function. However, their elementary chain dynamics have been difficult to resolve experimentally. Here we employ a combination of single-molecule Förster resonance energy transfer, nanosecond fluorescence correlation spectroscopy, fluorescence lifetime analysis, and nanophotonic enhancement to determine the conformational ensembles and rapid chain dynamics of short single-stranded nucleic acids in solution. To interpret the experimental results in terms of end-to-end distance dynamics, we utilize the hierarchical chain growth approach, simple polymer models, and refinement with Bayesian inference of ensembles to generate structural ensembles that closely align with the experimental data. The resulting chain reconfiguration times are exceedingly rapid, in the 10-ns range. Solvent viscosity-dependent measurements indicate that these dynamics of single-stranded nucleic acids exhibit negligible internal friction and are thus dominated by solvent friction. Our results provide a detailed view of the conformational distributions and rapid dynamics of single-stranded nucleic acids.
Understanding how protein mutations affect protein-nucleic acid binding is critical for unraveling disease mechanisms and advancing therapies. Current experimental approaches are laborious, and computational methods remain limited in accuracy. To address this challenge, we propose a novel topological machine learning model (TopoML) combining persistent Laplacian (from topological data analysis) with multi-perspective features: physicochemical properties, topological structures, and protein Transformer-derived sequence embeddings. This integrative framework captures robust representations of protein-nucleic acid binding interactions. To validate the proposed method, we employ two datasets, a protein-DNA dataset with 596 single-point amino acid mutations, and a protein-RNA dataset with 710 single-point amino acid mutations. We show that the proposed TopoML model outperforms state-of-the-art methods in predicting mutation-induced binding affinity changes for protein-DNA and protein-RNA complexes.
Nucleic acids theoretically possess a Szilard engine function that can convert the energy associated with the Shannon entropy of molecules for which they have coded recognition, into the useful work of geometric reconfiguration of the nucleic acid molecule. This function is logically reversible because its mechanism is literally and physically constructed out of the information necessary to reduce the Shannon entropy of such molecules, which means that this information exists on both sides of the theoretical engine, and because information is retained in the geometric degrees of freedom of the nucleic acid molecule, a quantum gate is formed through which multi-state nucleic acid qubits can interact. Entangled biophotons emitted as a consequence of symmetry breaking nucleic acid Szilard engine (NASE) function can be used to coordinate relative positioning of different nucleic acid locations, both within and between cells, thus providing the potential for quantum coherence of an entire biological system. Theoretical implications of understanding biological systems as such "quantum adaptive systems" include the potential for multi-agent based quantum computing, and a better understand
Nucleic acids can form diverse non-canonical structures, such as G-quadruplexes (G4s) and i-motifs (iMs), which are critical in biological processes and disease pathways. This study presents an innovative probe design strategy based on groove size differences, leading to the development of BT-Cy-1, a supramolecular cyanine probe optimized by fine-tuning dimer "thickness". BT-Cy-1 demonstrated high sensitivity in detecting structural transitions and variations in G4s and iMs, even in complex environments with excess dsDNA. Applied to clinical blood samples, it revealed significant differences in RNA G4 and iM levels between liver cancer patients and healthy individuals, marking the first report of altered iM levels in clinical samples. This work highlights a novel approach for precise nucleic acid structural profiling, offering insights into their biological significance and potential in disease diagnostics.
The transformer architecture has revolutionized bioinformatics and driven progress in the understanding and prediction of the properties of biomolecules. To date, most biosequence transformers have been trained on single-omic data - either proteins or nucleic acids - and have seen incredible success in downstream tasks in each domain, with particularly noteworthy breakthroughs in protein structural modeling. However, single-omic pretraining limits the ability of these models to capture cross-modal interactions. Here we present OmniBioTE, the largest open-source multi-omic model trained on over 250 billion tokens of mixed protein and nucleic acid data. We show that despite only being trained on unlabeled sequence data, OmniBioTE learns joint representations mapping genes to their corresponding protein sequences. We further demonstrate that OmniBioTE achieves state-of-the-art results predicting the change in Gibbs free energy ({ΔG}) of the binding interaction between a given nucleic acid and protein. Remarkably, we show that multi-omic biosequence transformers emergently learn useful structural information without any a priori structural training, allowing us to predict which protein
Nucleic acid-based drugs like aptamers have recently demonstrated great therapeutic potential. However, experimental platforms for aptamer screening are costly, and the scarcity of labeled data presents a challenge for supervised methods to learn protein-aptamer binding. To this end, we develop an unsupervised learning approach based on the predicted pairwise contact map between a protein and a nucleic acid and demonstrate its effectiveness in protein-aptamer binding prediction. Our model is based on FAFormer, a novel equivariant transformer architecture that seamlessly integrates frame averaging (FA) within each transformer block. This integration allows our model to infuse geometric information into node features while preserving the spatial semantics of coordinates, leading to greater expressive power than standard FA models. Our results show that FAFormer outperforms existing equivariant models in contact map prediction across three protein complex datasets, with over 10% relative improvement. Moreover, we curate five real-world protein-aptamer interaction datasets and show that the contact map predicted by FAFormer serves as a strong binding indicator for aptamer screening.
Lipid nanoparticles (LNPs) are highly effective carriers for gene therapies, including mRNA and siRNA delivery, due to their ability to transport nucleic acids across biological membranes, low cytotoxicity, improved pharmacokinetics, and scalability. A typical approach to formulate LNPs is to establish a quantitative structure-activity relationship (QSAR) between their compositions and in vitro/in vivo activities which allows for the prediction of activity based on molecular structure. However, developing QSAR for LNPs can be challenging due to the complexity of multi-component formulations, interactions with biological membranes, and stability in physiological environments. To address these challenges, we developed a machine learning framework to predict the activity and cell viability of LNPs for nucleic acid delivery. We curated data from 6,398 LNP formulations in the literature, applied nine featurization techniques to extract chemical information, and trained five machine learning models for binary and multiclass classification. Our binary models achieved over 90% accuracy, while the multiclass models reached over 95% accuracy. Our results demonstrated that molecular descripto
The precise quantification of nucleic acids is pivotal in molecular biology, underscored by the rising prominence of nucleic acid amplification tests (NAAT) in diagnosing infectious diseases and conducting genomic studies. This review examines recent advancements in digital Polymerase Chain Reaction (dPCR) and digital Loop-mediated Isothermal Amplification (dLAMP), which surpass the limitations of traditional NAAT by offering absolute quantification and enhanced sensitivity. In this review, we summarize the compelling advancements of dNNAT in addressing pressing public health issues, especially during the COVID-19 pandemic. Further, we explore the transformative role of artificial intelligence (AI) in enhancing dNAAT image analysis, which not only improves efficiency and accuracy but also addresses traditional constraints related to cost, complexity, and data interpretation. In encompassing the state-of-the-art (SOTA) development and potential of both software and hardware, the all-encompassing Point-of-Care Testing (POCT) systems cast new light on benefits including higher throughput, label-free detection, and expanded multiplex analyses. While acknowledging the enhancement of AI-
To understand and engineer biological and artificial nucleic acid systems, algorithms are employed for prediction of secondary structures at thermodynamic equilibrium. Dynamic programming algorithms are used to compute the most favoured, or Minimum Free Energy (MFE), structure, and the Partition Function (PF), a tool for assigning a probability to any structure. However, in some situations, such as when there are large numbers of strands, or pseudoknoted systems, NP-hardness results show that such algorithms are unlikely, but only for MFE. Curiously, algorithmic hardness results were not shown for PF, leaving two open questions on the complexity of PF for multiple strands and single strands with pseudoknots. The challenge is that while the MFE problem cares only about one, or a few structures, PF is a summation over the entire secondary structure space, giving theorists the vibe that computing PF should not only be as hard as MFE, but should be even harder. We answer both questions. First, we show that computing PF is #P-hard for systems with an unbounded number of strands, answering a question of Condon Hajiaghayi, and Thachuk [DNA27]. Second, for even a single strand, but allowin
Protein-nucleic acid interactions play a very important role in a variety of biological activities. Accurate identification of nucleic acid-binding residues is a critical step in understanding the interaction mechanisms. Although many computationally based methods have been developed to predict nucleic acid-binding residues, challenges remain. In this study, a fast and accurate sequence-based method, called ESM-NBR, is proposed. In ESM-NBR, we first use the large protein language model ESM2 to extract discriminative biological properties feature representation from protein primary sequences; then, a multi-task deep learning model composed of stacked bidirectional long short-term memory (BiLSTM) and multi-layer perceptron (MLP) networks is employed to explore common and private information of DNA- and RNA-binding residues with ESM2 feature as input. Experimental results on benchmark data sets demonstrate that the prediction performance of ESM2 feature representation comprehensively outperforms evolutionary information-based hidden Markov model (HMM) features. Meanwhile, the ESM-NBR obtains the MCC values for DNA-binding residues prediction of 0.427 and 0.391 on two independent test
What constitutes a habitable planet is a frontier to be explored and requires pushing the boundaries of our terracentric viewpoint for what we deem to be a habitable environment. Despite Venus' 700 K surface temperature being too hot for any plausible solvent and most organic covalent chemistry, Venus' cloud-filled atmosphere layers at 48 to 60 km above the surface hold the main requirements for life: suitable temperatures for covalent bonds; an energy source (sunlight); and a liquid solvent. Yet, the Venus clouds are widely thought to be incapable of supporting life because the droplets are composed of concentrated liquid sulfuric acid-an aggressive solvent that is assumed to rapidly destroy most biochemicals of life on Earth. Recent work, however, demonstrates that a rich organic chemistry can evolve from simple precursor molecules seeded into concentrated sulfuric acid, a result that is corroborated by domain knowledge in industry that such chemistry leads to complex molecules, including aromatics. We aim to expand the set of molecules known to be stable in concentrated sulfuric acid. Here, we show that nucleic acid bases adenine, cytosine, guanine, thymine, and uracil, as well
Halogen bonding (X-bonding) has attracted notable attention among noncovalent interactions. This highly directional attraction between a halogen atom and an electron donor has been exploited in knowledge-based drug design. A great deal of information has been gathered about X-bonds in protein-ligand complexes, as opposed to nucleic acid complexes. Here we provide a thorough analysis of nucleic acid complexes containing either halogenated building blocks or halogenated ligands. We analyzed close contacts between halogens and electron-rich moieties. The phosphate backbone oxygen is clearly the most common halogen acceptor. We identified 21 X-bonds within known structures of nucleic acid complexes. A vast majority of the X-bonds is formed by halogenated nucleobases, such as bromouridine, and feature excellent geometries. Noncovalent ligands have been found to form only interactions with suboptimal interaction geometries. Hence, the first X-bonded nucleic acid binder remains to be discovered.
Protein-nucleic acid complexes are important for many cellular processes including the most essential function such as transcription and translation. For many protein-nucleic acid complexes, flexibility of both macromolecules has been shown to be critical for specificity and/or function. Flexibility-rigidity index (FRI) has been proposed as an accurate and efficient approach for protein flexibility analysis. In this work, we introduce FRI for the flexibility analysis of protein-nucleic acid complexes. We demonstrate that a multiscale strategy, which incorporates multiple kernels to capture various length scales in biomolecular collective motions, is able to significantly improve the state of art in the flexibility analysis of protein-nucleic acid complexes. We take the advantage of the high accuracy and ${\cal O}(N)$ computational complexity of our multiscale FRI method to investigate the flexibility of large ribosomal subunits, which is difficult to analyze by alternative approaches. An anisotropic FRI approach, which involves localized Hessian matrices, is utilized to study the translocation dynamics in an RNA polymerase.
The structural flexibility of nucleic acids plays a key role in many fundamental life processes, such as gene replication and expression, DNA-protein recognition, and gene regulation. To obtain a thorough understanding of nucleic acid flexibility, extensive studies have been performed using various experimental methods and theoretical models. In this review, we will introduce the progress that has been made in understanding the flexibility of nucleic acids including DNAs and RNAs, and will emphasize the experimental findings and the effects of salt, temperature, and sequence. Finally, we will discuss the major unanswered questions in understanding the flexibility of nucleic acids.
Simple methods to detect biomolecules including specific nucleic acid sequences have received renewed attention since the Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) virus pandemic. Notably, biomolecule detection that uses some form of signal amplification will have some form of amplification-related error, which in the polymerase chain reaction involves mispriming and subsequent signal amplification in the no template control, ultimately providing a limit of detection. To demonstrate the feasibility of the detection of a DNA target sequence without molecular or chemical signal amplification that avoids amplification errors, a gold nanoparticle aggregation assay was developed and tested. Two primers bracketing a 94 base pair target sequence from SARS-CoV-2 were conjugated to 10 nm diameter gold nanoparticles by the salt aging method, with conjugation and primer-target hybridization confirmed by agarose gel electrophoresis and nanospectrophotometry. Upon mixing of both conjugated nanoparticles with target, a surface plasmon resonance shift of 6 nm was observed, and lower electrophoretic mobility of a band containing both DNA fluorescence and gold absorption signals.