Recent advances in machine learning (ML) have transformed protein science, enabling engineering and de novo design of artificial proteins with novel structures and functions. However, experimental analysis of key design features, such as oligomerization, folding, ligand binding, and dynamic conformational changes, remains critical. Here, we outline how mass spectrometry (MS) complements protein design through its ability to corroborate a wide range of design objectives. Furthermore, engineered proteins have become valuable tools for exploring the use of MS in detecting structural features, charge effects, and weak interactions by serving as testbeds for method development. Integrating ML and native MS thus creates a feedback loop: new designs challenge analytical techniques, while improved methods provide richer data to guide and improve future predictions. This synergy is vital for expanding the capabilities of protein engineering, including toward applications in synthetic biology and artificial protocell development.
The cytochrome P450 enzymes catalyze the hydroxylation of organic substrates by dioxygen. The high-potential reactive intermediate in cytochrome P450 catalysis, compound I (CI), has the capacity to deliver oxidizing equivalents (holes) to the side chains of tryptophan, tyrosine, and cysteine amino acids. Successful P450 catalysis requires that CI reacts more rapidly with a substrate than with these redox-active residues. The kinetics of hole transfer to tryptophan, tyrosine, and cysteine residues in four different P450 enzymes have been modeled using X-ray crystal structure coordinates and the semiclassical theory of electron transfer. Monte Carlo sampling of reaction driving forces has been used to account for uncertainties in the formal potentials of redox-active groups. The kinetics simulations suggest that the mean survival lifetimes of holes on the hemes range from ~100 ns to ~100 μs. Although hole transfer to the enzyme surface through redox-active amino acid reduces substrate oxidation efficiency, it can protect the enzyme from damage when reaction with substrate fails.
Blood can clot into anomalous, fibrinolysis-resistant forms that arise from prothrombotic seeding areas, including damaged cellular debris and membrane-derived surfaces, giving rise to what we have termed fibrinaloid microclot complexes (colloquially: microclots). Their proteolytic resistance is due in part to the fact that they are amyloid in nature, and they can also entrap inhibitors of proteolysis. They consist of a variety of proteins besides the expected fibrin, and are highly enriched for other amyloidogenic proteins (in contrast to normal clots, whose proteome largely reflects the soluble plasma proteome). They also contain DNA in the form of neutrophil extracellular traps (NETs). Importantly, fibrinaloid microclot complexes are heterogeneous structures comprising multiple phenotypic forms, including those that nucleate and grow on cellular debris such as damaged membranes, microparticles, and immune-derived material. We consider that these debris-associated complexes can act as catalytic scaffolds that recruit fibrin(ogen) and inflammatory molecules, thereby amplifying amyloidogenic transformation and prothrombotic activity. Fibrinaloid microclot complexes have been reported in a widening range of chronic inflammatory and thrombo-inflammatory diseases in which they have been sought, and are highly enriched for amyloidogenic proteins. Additionally, the thrombi extracted from ischaemic stroke also contain proteins in an amyloid form. One mechanism that explains how such macroclots can form and block arteries larger than any leading to them is that this occurs via the accretion of microclots that already contain amyloid. We here show that these microclots exhibit a classical 'apple-green' birefringence when stained with the dye Congo red. It is now important to determine whether inhibiting amyloid-forming clot transitions has therapeutic value.
Infrared (IR) nanoscopy represents a collection of imaging and spectroscopy techniques capable of resolving IR absorption on the nanometer scale. Chemical specificity is leveraged from vibrational spectroscopy, while light-matter interactions are detected by observing perturbations in the optical near field with an atomic force microscopy probe. Therefore, imaging is wavelength independent and has a spatial resolution on the nanometer scale, well beyond the classical diffraction limit. In this perspective, we outline the recent biological applications of scattering type scanning near-field optical microscopy and nanoscale Fourier-transform IR spectroscopy. These techniques are uniquely suited to resolving subcellular ultrastructure from a variety of cell types, as well as studying biological processes such as metabolic activity on the single-cell level. Furthermore, this review describes recent technical advances in IR nanoscopy, and emerging machine learning supported approaches to sampling, signal enhancement, and data processing. This emphasizes that label-free IR nanoscopy holds significant potential for ongoing and future biological applications.
During the past decade, emerging studies using electrochemistry and nanoscale imaging have demonstrated that partial exocytotic release is prevailing in neuroendocrine cell models. However, due to complicated structure and culture process, few studies have been carried out using neurons, especially human neurons. Here, dopamine (DA) release from individual vesicles and DA content stored within vesicles were quantified from induced pluripotent stem cell-derived DA neurons with electrochemical techniques. The results indicate that around 61% of the total vesicular DA content is released from these neurons during exocytosis. The vesicular content quantified in DA neurons is significantly higher than that in undifferentiated neural progenitor cells, owing to the increased appearance of dense-core vesicles that are able to store more DA molecules than the clear vesicles. When the neurons are differentiated with BAY-K8644, which stimulates neuronal maturation as well as DA release, the release fraction rises to 91%. The use of BAY-K8644 can be considered as chronic stimulation and leads to similar effects on exocytosis as repetitive stimulation, which triggers short-term plasticity. This study demonstrates partial release in DA transmission in human neurons and provides a link between neuronal maturation and the formation of plasticity. Furthermore, this work suggests that the fraction of release in exocytosis at human neurons may be a factor in determining plasticity.
Real-time, in situ monitoring of neurochemical dynamics in intact neural circuits is critical for elucidating brain function. Recent innovations in micro- and nanoelectrode engineering have markedly advanced our ability to detect neurotransmitter and neuromodulator release with high spatiotemporal resolution, while the application of machine learning (ML) has facilitated the development of next-generation electrodes and enhanced signal processing capabilities. Here, we outline a vision for the potential directions for electrode interface design and the deepening integration of ML in in situ neurochemical sensing, illustrating how breakthroughs over the past decade have illuminated these opportunities.
Machine learning (ML) has revolutionised the field of structure-based drug design (SBDD) in recent years. During the training stage, ML techniques typically analyse large amounts of experimentally determined data to create predictive models in order to inform the drug discovery process. Deep learning (DL) is a subfield of ML, that relies on multiple layers of a neural network to extract significantly more complex patterns from experimental data, and has recently become a popular choice in SBDD. This review provides a thorough summary of the recent DL trends in SBDD with a particular focus on de novo drug design, binding site prediction, and binding affinity prediction of small molecules.
暂无摘要(点击查看详情)
Human mitochondrial Complex I is one of the largest multi-subunit membrane protein megacomplexes, which plays a critical role in oxidative phosphorylation and ATP production. It is also involved in many neurodegenerative diseases. However, studying its structure and the mechanisms underlying proton translocation remains challenging due to the hydrophobic nature of its transmembrane parts. In this structural bioinformatic study, we used the QTY code to reduce the hydrophobicity of megacomplex I, while preserving its structure and function. We carried out the structural bioinformatics analysis of 20 key enzymes in the integral membrane parts. We compare their native structure, experimentally determined using Cryo-electron microscopy (CryoEM), with their water-soluble QTY analogs predicted using AlphaFold 3. Leveraging AlphaFold 3's advanced capabilities in predicting protein-protein complex interactions, we further explore whether the QTY-code integral membrane proteins maintain their protein-protein interactions necessary to form the functional megacomplex. Our structural bioinformatics analysis not only demonstrates the feasibility of engineering water-soluble integral membrane proteins using the QTY code, but also highlights the potential to use the water-soluble membrane protein QTY analogs as soluble antigens for discovery of therapeutic monoclonal antibodies, thus offering promising implications for the treatment of various neurodegenerative diseases.
RNA molecules play many functional and regulatory roles in cells, and hence, have gained considerable traction in recent times as therapeutic interventions. Within drug discovery, structure-based approaches have successfully identified potent and selective small-molecule modulators of pharmaceutically relevant protein targets. Here, we embrace the perspective of computational chemists who use these traditional approaches, and we discuss the challenges of extending these methods to target RNA molecules. In particular, we focus on recognition between RNA and small-molecule binders, on selectivity, and on the expected properties of RNA ligands.
Integrins are critical transmembrane receptors that connect the extracellular matrix (ECM) to the intracellular cytoskeleton, playing a central role in mechanotransduction - the process by which cells convert mechanical stimuli into biochemical signals. The dynamic assembly and disassembly of integrin-mediated adhesions enable cells to adapt continuously to changing mechanical cues, regulating essential processes such as adhesion, migration, and proliferation. In this review, we explore the molecular clutch model as a framework for understanding the dynamics of integrin - ECM interactions, emphasizing the critical importance of force loading rate. We discuss how force loading rate bridges internal actomyosin-generated forces and ECM mechanical properties like stiffness and ligand density, determining whether sufficient force is transmitted to mechanosensitive proteins such as talin. This force transmission leads to talin unfolding and activation of downstream signalling pathways, ultimately influencing cellular responses. We also examine recent advances in single-molecule DNA tension sensors that have enabled direct measurements of integrin loading rates, refining the range to approximately 0.5-4 pN/s. These findings deepen our understanding of force-mediated mechanotransduction and underscore the need for improved sensor designs to overcome current limitations.
Single-molecule methods offer powerful insights into DNA-protein interactions at the individual DNA molecule level. We developed an automated, high-throughput nanofluidic imaging platform to characterize DNA-protein complexes in solution. The platform uses a nanofluidic chip with 10 sets of nanochannels where thousands of DNA molecules can be simultaneously analyzed in different conditions. Using this approach, we investigate Rok, a multifunctional Bacillus subtilis protein involved in genome organization and transcription regulation. Our findings confirm the DNA-condensing activity of Rok, likely attributed to its ability to bridge distant DNA segments. Additionally, Rok promotes the hybridization of 12 base complementary single-stranded DNA overhangs, suggesting a potential role in homology search during recombination. Rok also displays sequence-selective binding, preferentially associating with adenine and thymine-rich (AT-rich) DNA regions. To explore the structural features of Rok underlying these activities and test our nanofluidic system further, we compare wild-type Rok with two variants: ∆Rok, lacking the neutral part of the internal linker, and sRok, a naturally occurring variant without the linker. This comparison highlights the role of the linker in hybridization, i.e., interaction with single-stranded DNA. Together, these findings enhance our understanding of Rok-mediated DNA dynamics and establish single-molecule nanofluidics as a powerful tool for high-throughput studies of DNA-protein interactions.
Single-stranded nucleic acid (ssNA) binding proteins must both stably protect ssNA transiently exposed during replication and other NA transactions, and also rapidly reorganize and dissociate to allow further NA processing. How these seemingly opposing functions can coexist has been recently elucidated by optical tweezers (OT) experiments that isolate and manipulate single long ssNA molecules to measure conformation in real time. The effective length of an ssNA substrate held at fixed tension is altered upon protein binding, enabling quantification of both the structure and kinetics of protein-NA interactions. When proteins exhibit multiple binding states, however, OT measurements may produce difficult to analyze signals including non-monotonic response to free protein concentration and convolution of multiple fundamental rates. In this review we compare single-molecule experiments with three proteins of vastly different structure and origin that exhibit similar ssNA interactions. These results are consistent with a general model in which protein oligomers containing multiple binding interfaces switch conformations to adjust protein:NA stoichiometry. These characteristics allow a finite number of proteins to protect long ssNA regions by maximizing protein-ssNA contacts while also providing a pathway with reduced energetic barriers to reorganization and eventual protein displacement when these ssNA regions are diminished.
The molecular mechanism of olfaction, namely, how we smell with limited olfactory receptors to recognize exceedingly diverse and large numbers of scents remains unknown despite the recent advances in chemistry, chemical, structural, and molecular biology. Olfactory receptors are notoriously difficult to study because they are fully embedded in the cell membrane. After decades of efforts and significant funding, there are only three olfactory receptor structures known. To understand olfaction, we carried out the structural bioinformatic study of six human olfactory receptors including OR51E1, OR51E2, OR52cs, OR1A1, OR1A2, TAAR9, and their AlphaFold3 predicted water-soluble QTY variants with odorants. We applied the QTY code to replace leucine (L) with glutamine (Q), isoleucine (I) and valine (V) with threonine (T), and phenylalanine (F) with tyrosine (Y) only in the transmembrane helices. Therefore, these QTY variants become water-soluble. We also present the superimposed structures of native olfactory receptors and their water-soluble QTY variants. The superimposed structures show remarkable similarity with RMSDs between 0.441 and 1.275 Å despite significant changes to the protein sequence of the transmembrane domains (43.03%-50.31%). We also show the differences in hydrophobicity surfaces between the native olfactory receptors and their QTY variants. Furthermore, we also used AlphaFold3 and molecular dynamics to study the odorant octanoate with OR1A2 and spermidine with TAAR9. Our bioinformatics studies provide insight into the differences between the hydrophobic helices and hydrophilic helices, and will likely further stimulate designs of water-soluble integral transmembrane proteins and other aggregated proteins.
Per-Arnt-Sim (PAS) domain kinase (PASK) is a conserved metabolic sensor that modulates the activation of critical proteins involved in liver metabolism and fitness. However, despite its key role in mastering the metabolic regulation, the molecular mechanism of PASK's activity is ongoing research, and structural information of this important protein is scarce. To investigate this, we integrated structural bioinformatics with state-of-the-art modeling and molecular simulation techniques. Our goals were to address (1) how many regulatory PAS domains PASK is likely to have, (2) how those domains modulate the kinase activity, and (3) how those interactions could be controlled by small molecules. Our results indicated the existence of three N-terminal PAS domains. Solvent mapping and fragment docking identified a consensus set of 'druggable hot spots' within all domains, as well as at domain-domain interfaces. Those 'hot spots' could be modulated with chemically diverse small molecular probes, which may serve as a starting point for rationally designed therapeutics modulating these specific sites. Our results identified a plausible mechanism of autoinhibition of kinase activity, suggesting that all three putative PAS domains may be required. Future work will focus on validation of the predicted PASK models and development of small-molecule inhibitors of PASK by targeting its 'druggable hot spots'.
Viruses are highly dynamic macromolecular assemblies. They undergo large-scale changes in structure and organization at nearly every stage of their infectious cycles from virion assembly to maturation, receptor docking, cell entry, uncoating and genome delivery. Understanding structural transformations and dynamics across the virus infectious cycle is an expansive area for research that that can also provide insight into mechanisms for blocking infection, replication, and transmission. Additionally, the processes viruses carry out serve as excellent model systems for analogous cellular processes, but in more accessible form. Capturing and analyzing these dynamic events poses a major challenge for many structural biological approaches due to the size and complexity of the assemblies and the heterogeneity and transience of the functional states that are populated. Here we examine the process of protein-mediated membrane fusion, which is carried out by specialized machinery on enveloped virus surfaces leading to delivery of the viral genome. Application of two complementary methods, cryo-electron tomography and structural mass spectrometry enable dynamic intermediate states in intact fusion systems to be imaged and probed, providing a new understanding of the mechanisms and machinery that drive this fundamental biological process.
Neurodegenerative disorders, such as Alzheimer's and Parkinson's diseases, are associated with the formation of amyloid fibrils. The DNAJB6b (JB6) chaperone greatly inhibits the disease-related self-assembly of amyloid peptides in an ATP-independent manner. The molecular basis of this process is, however, not understood. Here, we studied the low complexity linker between the N- and C-terminal domains of JB6 as an isolated 110 amino acid residue construct, to get a better understanding of the role of the composition of the intact protein. We investigate the structure and aggregation behaviour of the linker and its anti-amyloid activity in comparison with the full-length chaperone. We find that the linker contains ca. 45% α-helix and 20% β-sheet and is in itself an amyloid-like peptide that self-assembles into different structures, which are bigger than those formed by the intact chaperone, including fibrils. The isolated linker protects against fibril formation of Aβ42 as well as α-synuclein, but is less potent than the intact chaperone. Based on our results, we propose a possible mechanism behind JB6 and linker amyloid suppression relating to their self-assembly behaviour. In the intact protein, the domains serve to solubilize the linker such that the solution concentration of exposed linker is high enough to sustain its high potency against amyloid formation.
DNA unzipping by nanopore translocation has implications in diverse contexts, from polymer physics to single-molecule manipulation to DNA-enzyme interactions in biological systems. Here we use molecular dynamics simulations and a coarse-grained model of DNA to address the nanopore unzipping of DNA filaments that are knotted. This previously unaddressed problem is motivated by the fact that DNA knots inevitably occur in isolated equilibrated filaments and in vivo. We study how different types of tight knots in the DNA segment just outside the pore impact unzipping at different driving forces. We establish three main results. First, knots do not significantly affect the unzipping process at low forces. However, knotted DNAs unzip more slowly and heterogeneously than unknotted ones at high forces. Finally, we observe that the microscopic origin of the hindrance typically involves two concurrent causes: the topological friction of the DNA chain sliding along its knotted contour and the additional friction originating from the entanglement with the newly unzipped DNA. The results reveal a previously unsuspected complexity of the interplay of DNA topology and unzipping, which should be relevant for interpreting nanopore-based single-molecule unzipping experiments and improving the modeling of DNA transactions in vivo.
Metabolism is at the core of all functions of living cells as it provides Gibbs free energy and building blocks for synthesis of macromolecules, which are necessary for structures, growth, and proliferation. Metabolism is a complex network composed of thousands of reactions catalyzed by enzymes involving many co-factors and metabolites. Traditionally it has been difficult to study metabolism as a whole network and most traditional efforts were therefore focused on specific metabolic pathways, enzymes, and metabolites. By using engineering principles of mathematical modeling to analyze and study metabolism, as well as engineer it, that is, design and build, new metabolic features, it is possible to gain many new fundamental insights as well as applications in biotechnology. Here, we present the history and basic principles of engineering metabolism, as well as the newest developments in the field. We are using examples of applications in: (1) production of protein pharmaceuticals and chemicals; (2) basic studies of metabolism; and (3) impacting health care. We will end by discussing how engineering metabolism can benefit from advances in artificial intelligence (AI)-based models.