Carbohydrates, essential biological building blocks, exhibit functional mechanisms tied to their intricate stereochemistry. Subtle stereochemical differences, such as those between the anomers maltose and cellobiose, lead to distinct properties due to their differing glycosidic bonds; the former is digestible by humans, while the latter is not. This underscores the importance of precise structural determination of individual carbohydrate molecules for deeper functional insights. However, their structural complexity and conformational flexibility, combined with the high spatial resolution needed, have hindered direct imaging of carbohydrate stereochemistry. Here, we employ non-contact atomic force microscopy integrated with a data-efficient, multi-fidelity structure search approach accelerated by machine learning integration to determine the precise 3D atomic coordinates of two carbohydrate anomers. We observe that glycosidic bond stereochemistry regulates on-surface chiral selection in carbohydrate self-assemblies. The reconstructed models, validated against experimental data, provide reliable atomic-scale structural evidence, uncovering the origin of on-surface chirality from carb
Molecular representation learning (MRL) is a powerful tool for bridging the gap between machine learning and chemical sciences, as it converts molecules into numerical representations while preserving their chemical features. These encoded representations serve as a foundation for various downstream biochemical studies, including property prediction and drug design. MRL has had great success with proteins and general biomolecule datasets. Yet, in the growing sub-field of glycoscience (the study of carbohydrates, where longer carbohydrates are also called glycans), MRL methods have been barely explored. This under-exploration can be primarily attributed to the limited availability of comprehensive and well-curated carbohydrate-specific datasets and a lack of Machine learning (ML) pipelines specifically tailored to meet the unique problems presented by carbohydrate data. Since interpreting and annotating carbohydrate-specific data is generally more complicated than protein data, domain experts are usually required to get involved. The existing MRL methods, predominately optimized for proteins and small biomolecules, also cannot be directly used in carbohydrate applications without sp
This work looks at the relationship between the index of hydrogen deficiency (IHD) and the enthalpy of combustion of "sugar propellant," as well as the performance of carbohydrates with similar IHD values. The study used eight different carbohydrate sources as fuels in the propellant, combined with potassium nitrate as an oxidizer in a 35:65 ratio. The IHD of the sugars ranged from 0 (polyols) to 2 (disaccharides). Different propellant mixtures (carbohydrate-KN) were tested using calorimeters and chemical analysis. The results support the hypothesis that IHD is associated with the enthalpy of combustion of sugar propellant, with polyol reactions showing the highest enthalpy change. Moreover, carbohydrates with a higher molar mass and an IHD of 2 exhibit better performance than those with an IHD of 1.
Carbohydrates, vital components of biological systems, are well-known for their structural diversity. Nuclear Magnetic Resonance (NMR) spectroscopy plays a crucial role in understanding their intricate molecular arrangements and is essential in assessing and verifying the molecular structure of organic molecules. An important part of this process is to predict the NMR chemical shift from the molecular structure. This work introduces a novel approach that leverages E(3) equivariant graph neural networks to predict carbohydrate NMR spectra. Notably, our model achieves a substantial reduction in mean absolute error, up to threefold, compared to traditional models that rely solely on two-dimensional molecular structure. Even with limited data, the model excels, highlighting its robustness and generalization capabilities. The implications are far-reaching and go beyond an advanced understanding of carbohydrate structures and spectral interpretation. For example, it could accelerate research in pharmaceutical applications, biochemistry, and structural biology, offering a faster and more reliable analysis of molecular structures. Furthermore, our approach is a key step towards a new data-
Carbohydrates such as the trisaccharide motif LeX are key constituents of cell surfaces. Despite intense research, the interactions between carbohydrates of apposing cells or membranes are not well understood. In this article, we investigate carbohydrate-carbohydrate interactions in membrane adhesion as well as in solution with extensive atomistic molecular dynamics simulations that exceed the simulation times of previous studies by orders of magnitude. For LeX, we obtain association constants of soluble carbohydrates, adhesion energies of lipid-anchored carbohydrates, and maximally sustained forces of carbohydrate complexes in membrane adhesion that are in good agreement with experimental results in the literature. Our simulations thus appear to provide a realistic, detailed picture of LeX-LeX interactions in solution and during membrane adhesion. In this picture, the LeX-LeX interactions are fuzzy, i.e. LeX pairs interact in a large variety of short-lived, bound conformations. For the synthetic tetrasaccharide Lac 2, which is composed of two lactose units, we observe similarly fuzzy interactions and obtain association constants of both soluble and lipid-anchored variants that are
The Martini 3 force field is a full re-parametrization of the Martini coarse-grained model for biomolecular simulations. Due to the improved interaction balance it allows for more accurate description of condensed phase systems. In the present work we develop a consistent strategy to parametrize carbohydrate molecules accurately within the framework of Martini 3. In particular, we develop a canonical mapping scheme that decomposes arbitrarily large carbohydrates into a limited number of fragments. Bead types for these fragments have been assigned by matching physicochemical properties of mono- and disaccharides. In addition, guidelines for assigning bonds, angles, and dihedrals are developed. These guidelines enable a more accurate description of carbohydrate conformations than in the Martini 2 force field. We show that models obtained with this approach are able to accurately reproduce osmotic pressures of carbohydrate water solutions. Furthermore, we provide evidence that the model differentiates correctly the solubility of the poly-glucoses dextran (water soluble) and cellulose (water insoluble, but soluble in ionic-liquids). Finally, we demonstrate that the new building blocks
To substitute petroleum-based materials with bio-based alternatives, microbial fermentation combined with inexpensive biomass is suggested. In this study Saccharina latissima hydrolysate, candy-factory waste, and digestate from full-scale biogas plant were explored as substrates for lactic acid production. The lactic acid bacteria Enterococcus faecium, Lactobacillus plantarum, and Pediococcus pentosaceus were tested as starter cultures. Sugars released from seaweed hydrolysate and candy-waste were successfully utilized by the studied bacterial strains. Additionally, seaweed hydrolysate and digestate served as nutrient supplements supporting microbial fermentation. According to the highest achieved relative lactic acid production, a scaled-up co-fermentation of candy-waste and digestate was performed. Lactic acid reached a concentration of 65.65 g/L, with 61.69% relative lactic acid production, and 1.37 g/L/hour productivity. The findings indicate that lactic acid can be successfully produced from low-cost industrial residues.
The possibility of carbohydrate separation in BEH HILIC (Ethylene Bridged Hybride, Hydrophilic Interaction Liquid Chromatography) column was studied by ultra-performance liquid chromatography (UPLC) with evaporative light scattering detector (ELSD) and mobile phase containing amine compounds as modifiers. The chromatography conditions and ELSD parameters were optimized to separate five typical carbohydrates and applied to analysis of four infant milk powders. The linear ranges of carbohydrate determination were 20-300mg/L for fructose and glucose, 20-250mg/L for sucrose and lactose, and 35-180mg/L for fructo-oligosaccharide. The LODs were 16.4mg/L for fructose and glucose, 17.3mg/L for sucrose, 20.0mg/L for lactose, and 46.7mg/L for fructo-oligosaccharide. Relative standard deviations (RSDs) ranged between 3.45-4.23%, 1.46-4.17%, 4.14-5.60%, 1.39-4.09%, and 2.49-3.61% for fructose, glucose, sucrose, lactose, and fructo-oilgosaccharide, respectively and recoveries ranged between 95.0 and 105.4%
In the present study, different photoperiods and nutritional conditions were applied to a mixed wastewater-borne cyanobacterial culture in order to enhance the intracellular accumulation of polyhydroxybutyrates (PHBs) and carbohydrates. Two different experimental set-ups were used. In the first, the culture was permanently exposed to illumination, while in the second; it was submitted to light/dark alternation (12h cycles). In both cases, two different nutritional regimes were also evaluated, N-limitation and P-limitation. Results showed that the highest PHB concentration (104 mg L-1) was achieved under P limited conditions and permanent illumination, whereas the highest carbohydrate concentration (838 mg L-1) was obtained under N limited condition and light/dark alternation. With regard to bioplastics and biofuel generation, this study demonstrates that the accumulation of PHBs (bioplastics) and carbohydrates (potential biofuel substrate) is favored in wastewater-borne cyanobacteria under conditions where nutrients are limited.
To avoid serious diabetic complications, people with type 1 diabetes must keep their blood glucose levels (BGLs) as close to normal as possible. Insulin dosages and carbohydrate consumption are important considerations in managing BGLs. Since the 1960s, models have been developed to forecast blood glucose levels based on the history of BGLs, insulin dosages, carbohydrate intake, and other physiological and lifestyle factors. Such predictions can be used to alert people of impending unsafe BGLs or to control insulin flow in an artificial pancreas. In past work, we have introduced an LSTM-based approach to blood glucose level prediction aimed at "what if" scenarios, in which people could enter foods they might eat or insulin amounts they might take and then see the effect on future BGLs. In this work, we invert the "what-if" scenario and introduce a similar architecture based on chaining two LSTMs that can be trained to make either insulin or carbohydrate recommendations aimed at reaching a desired BG level in the future. Leveraging a recent state-of-the-art model for time series forecasting, we then derive a novel architecture for the same recommendation task, in which the two LSTM
Time-resolved transient absorption spectroscopy has been used to study nanosecond and sub-microsecond electron dynamics in aqueous anatase nanoparticles in the presence of hole scavengers: chemisorbed polyols and carbohydrates. These polyhydroxylated compounds are rapidly oxidized by the holes; 50-60% of these holes are scavenged within the duration of 355 nm excitation laser pulse. The scavenging efficiency rapidly increases with the number of anchoring hydroxyl groups and varies considerably as a function of the carbohydrate structure. A specific binding site for the polyols and carbohydrates is suggested that involves an octahedral Ti atom chelated by the poly-OH ligand. This mode of binding accounts for the depletion of undercoordinated Ti atoms observed in the XANES spectra of polyol coated nanoparticles. We suggest that these binding sites trap a substantial fraction of holes before the latter descend to surface traps and/or recombine with free electrons. The resulting oxygen hole center rapidly loses a CH proton to the environment, yielding a metastable C-centered radical.
A comprehensive study of the effects of carbohydrate doping on the superconductivity of MgB2 has been conducted. In accordance with the dual reaction model, more carbon substitution is achieved at lower sintering temperature. As the sintering temperature is lowered, lattice disorder is increased. Disorder is an important factor determining the transition temperature for the samples studied in this work, as evidenced from the correlations among the lattice strain, the resistivity and the transition temperature. It is further shown that the increased critical current density in the high field region can be understood by a recently-proposed percolation model [Phys. Rev. Lett. 90 (2003) 247002]. For the critical current density analysis, the upper critical field is estimated from a correlation that has been reported in a recent review article [Supercond. Sci. Technol. 20 (2007) R47], where a sharp increase in the upper critical field by doping is mainly due to an increase in lattice disorder or impurity scattering. On the other hand, it is shown that the observed reduction in self-field critical current density is related to the reduction in the pinning force density by carbohydrate do
With the relatively high critical temperature (Tc) of 39 K1 and the high critical current density (Jc) of > 100000 A/cm2 in moderate fields, magnesium diboride (MgB2) superconductors could offer the promise of important large-scale and electronic device applications to be operated at 20 K. A significant enhancement in the electromagnetic properties of MgB2 has been achieved through doping with various form of carbon (C). However, doping effect has been limited by the agglomeration of nano-sized dopants and the poor reactivity of C containing dopants with MgB2. Un-reacted dopants result in a reduction of superconductor volume. In this work, we demonstrate the advantages of carbohydrate doping over other dopants, resulting in an increase of in-field Jc by more than one order of magnitude without any degradation of self-field Jc. As there are numerous carbohydrates readily available this finding has significant ramifications not only for the fabrication of MgB2 but also for many C based compounds and composites.
Predicting a patient's physiological trajectory under a planned treatment sequence is a prospective interventional problem, not standard time-series extrapolation. We study this problem in glucose management, where insulin and carbohydrate records are policy-dependent: future drivers are coupled to patient state, behavior, and clinical decision rules, so observational forecasting accuracy alone does not guarantee correct responses to planned interventions. We introduce Interventional Flow Matching (IFM), a continuous-time generative framework for physiologically constrained prospective forecasting. IFM conditions a flow-matching velocity field on patient history and planned future drivers in a bounded latent glucose space. Rather than embedding strict mechanistic glucose--insulin ODE equations or enforcing causality through rollout-based simulations, IFM uses a solver-free regularization: it penalizes the Jacobian of the instantaneous velocity field with respect to smoothed treatment drivers. This imposes signed, dose-bounded local sensitivities directly on the learned dynamics: insulin lowers glucose, carbohydrates raise it, and both responses remain within plausible ranges. On a
Glucose forecasting algorithms are an important aspect of glycemic control management in type 1 diabetes. So far, the research community has developed numerous algorithms and models for forecasting. However, it is well-recognized that the lack of standardized model performance evaluation benchmarks makes fair comparison difficult and hinders further innovation, and thus benchmark standardization is in urgent need. Furthermore, many published glucose forecasting algorithms are limited to CGM data alone, ignoring other multimodal signals such as insulin dosing and carbohydrate intake. Here, we introduce MetaboNet-Bench, a benchmark for multimodal glucose forecasting for patients with type 1 diabetes that provides an extensible open-source evaluation framework for comparison of glucose forecasting algorithms that leverage glucose, insulin, and carbohydrate data. We then demonstrate its utility by benchmarking several recently published glucose forecasting models and a custom multimodal time-series model, representing different model architectures. The results show that the benefit of adding data modalities is conditioned on the complexity of the model and that incorporating more clini
Current closed-loop insulin delivery algorithms need to be informed of carbohydrate intake disturbances. This can be a burden on people using these systems. Pramlintide is a hormone that delays gastric emptying, which enables insulin kinetics to align with the kinetics of carbohydrate absorption. Integrating pramlintide into an automated insulin delivery system can be helpful in reducing the postprandial glucose excursion and may be helpful in enabling fully-closed loop whereby meals do not need to be announced. We present an AI-enabled dual-hormone model predictive control (MPC) algorithm that delivers insulin and pramlintide without requiring meal announcements that uses a neural network to automatically detect and deliver meal insulin. The MPC algorithm includes a new pramlintide pharmacokinetics and pharmacodynamics model that was identified using data collected from people with type 1 diabetes undergoing a meal challenge. Using a simulator, we evaluated the performance of various pramlintide delivery methods and controller models, as well as the baseline insulin-only scenario. Meals were automatically dosed using a neural network meal detection and dosing (MDD) algorithm. The
Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first publicly available natural language meal description nutrition benchmark. NutriBench consists of 11,857 meal descriptions generated from real-world global dietary intake data. The data is human-verified and annotated with macro-nutrient labels, including carbohydrates, proteins, fats, and calories. We conduct an extensive evaluation of NutriBench on the task of carbohydrate estimation, testing twelve leading Large Language Models (LLMs), including GPT-4o, Llama3.1, Qwen2, Gemma2, and OpenBioLLM models, using standard, Chain-of-Thought and Retrieval-Augmented Generation strategies. Additionally, we present a study involving professional nutritionists, finding that LLMs can provide comparable but significantly faster estimates. Finally, we perform a real-world risk assessment by simulating the effect of carbohydrate predictions on the blood glucose levels of individuals with diabetes. Our work highlights the opportunities and challenges of using LLMs for nutrition estimation, demonstrating their potential to aid
A continuing frustration for origin of life scientists is that abiotic and, by extension, pre-biotic attempts to develop self-sustaining, evolving molecular systems tend to produce more dead-end substances than macromolecular products with the necessary potential for biostructure and function -- the so-called `tar problem'. Nevertheless primordial life somehow emerged despite that presumed handicap. A~resolution of this problem is important in emergence-of-life science because it would provide valuable guidance in choosing subsequent paths of investigation, such as identifying pre-biotic patterns on Mars. To study the problem we set up a simple non-equilibrium flow dynamical model for the coupled temperature and mass dynamics of the decomposition of a polymeric carbohydrate adsorbed on a mineral surface, with incident stochastic thermal fluctuations. Results show that the model system behaves as a reciprocating thermochemical oscillator. The output fluctuation distribution is bimodal, with a right-weighted component that guarantees a bias towards detachment and desorption of monomeric species such as ribose, even while tar is formed concomitantly. This fluctuating thermochemical re
To overcome antimalarial drug resistance, carbohydrate derivatives as selective PfHT1 inhibitor have been suggested in recent experimental work with orthosteric and allosteric dual binding pockets. Inspired by this promising therapeutic strategy, herein, molecular dynamics simulations are performed to investigate the molecular determinants of co-administration on orthosteric and allosteric inhibitors targeting PfHT1. Our binding free energy analysis capture the essential trend of inhibitor binding affinity to protein from published experimental IC50 data in three sets of distinct characteristics. In particular, we rank the contribution of key residues as binding sites which categorized into three groups based on linker length, size of tail group, and sugar moiety of inhibitors. The pivotal roles of these key residues are further validated by mutant analysis where mutated to nonpolar alanine leading to reduced affinities to different degrees. The exception was fructose derivative, which exhibited a significant enhanced affinity to mutation on orthosteric sites due to strong changed binding poses. This study may provide useful information for optimized design of precision medicine to
High quality real world datasets are essential for advancing data driven approaches in type 1 diabetes (T1D) management, including personalized therapy design, digital twin systems, and glucose prediction models. However, progress in this area has been limited by the scarcity of publicly available datasets that offer detailed and comprehensive patient data. To address this gap, we present AZT1D, a dataset containing data collected from 25 individuals with T1D on automated insulin delivery (AID) systems. AZT1D includes continuous glucose monitoring (CGM) data, insulin pump and insulin administration data, carbohydrate intake, and device mode (regular, sleep, and exercise) obtained over 6 to 8 weeks for each patient. Notably, the dataset provides granular details on bolus insulin delivery (i.e., total dose, bolus type, correction specific amounts) features that are rarely found in existing datasets. By offering rich, naturalistic data, AZT1D supports a wide range of artificial intelligence and machine learning applications aimed at improving clinical decision making and individualized care in T1D.