Sleep is a fundamental biological process with profound implications for physical and mental health, yet our understanding of its complex patterns and their relationships to a broad spectrum of diseases remains limited. While polysomnography (PSG), the gold standard for sleep analysis, captures rich multimodal physiological data, analyzing these measurements has been challenging due to limited flexibility across recording environments, poor generalizability across cohorts, and difficulty in leveraging information from multiple signals simultaneously. To address this gap, we curated over 585,000 hours of high-quality sleep recordings from approximately 65,000 participants across multiple cohorts and developed SleepFM, a multimodal sleep foundation model trained with a novel contrastive learning approach, designed to accommodate any PSG montage. SleepFM produces informative sleep embeddings that enable predictions of future diseases. We systematically demonstrate that SleepFM embeddings can predict 130 future diseases, as modeled by Phecodes, with C-Index and AUROC of at least 0.75 on held-out participants (Bonferroni-corrected p < 0.01). This includes accurate predictions for death (C-Index: 0.84 [95% CI: 0.81-0.87]), heart failure (C-Index: 0.80 [95% CI: 0.77-0.83]), chronic kidney disease (C-Index: 0.79 [95% CI: 0.77-0.81]), dementia (C-Index: 0.85 [95% CI: 0.82-0.87]), stroke (C-Index: 0.78 [95% CI: 0.76-0.81]), atrial fibrillation (C-Index: 0.78 [95% CI: 0.75-0.81]), and myocardial infarction (C-Index: 0.81 [95% CI: 0.78-0.84]). The model's generalizability was further validated through strong performance on the Sleep Heart Health Study (SHHS), a dataset unseen during pre-training. Additionally, SleepFM demonstrates strong performance on traditional sleep analysis tasks, achieving competitive results in both sleep staging (mean F1 scores: 0.70-0.78) and sleep apnea diagnosis (AUROC: 0.90-0.94). Beyond these standard applications, our analysis reveals that specific sleep stages and physiological signals carry distinct predictive power for different diseases. This work demonstrates how foundation models can leverage sleep polysomnography data to uncover the extensive relationship between sleep physiology and future disease risk.
Vitamin D status has been found to be inversely associated with risk of respiratory tract infections (RTIs). Although vitamin D status varies by ethnicity, the relationship between serum 25-hydroxyvitamin D (25[OH]D) and RTIs in United Kingdom ethnic groups remains unclear. This study aimed to investigate the association between serum 25(OH)D status and hospitalization for RTI in United Kingdom adults. An unmatched case-control study was conducted using data from United Kingdom Biobank, which includes 500k adults with serum 25(OH)D status and hospital episodes from linked records. Survival analyses and binary logistic regression models were used to explore the association between serum 25(OH)D and RTIs. Of the 36,258 participants included in the analysis, 34% were White, 28% Asian, 19% Black, 11% other, and 7% of mixed ethnicity. The RTI rate was 8.5% (median time to RTI, 14.8 y). Higher serum 25(OH)D (each +10 nmol/L increase) was significantly associated with a 4% lower hazard ratio (HR) for RTI hospitalization [HR: 0.96, 95% confidence interval (CI), 0.94, 0.99]. When stratifying for serum 25(OH)D, compared to those with ≥75 nmol/L (reference), those with <15 nmol/L had a higher HR for RTI hospitalization (HR: 1.33, 95% CI: 1.05, 1.67). Categories 15 to 24 nmol/L, 24 to 49 nmol/L, and 50 to 74 nmol/L were not statistically significant. Logistic regression models supported the above findings. Inclusion of an interaction term for 25(OH)D × ethnicity was trialed in the survival analysis, but the interaction term was not statistically significant. Serum 25(OH)D status <15 nmol/L is associated with 33% higher HR for RTI hospitalization among United Kingdom adults, compared with ≥75 nmol/L. Furthermore, studies are warranted to validate these findings and explore the mechanisms underlying the association between vitamin D status and RTIs in different ethnic groups.
Exhaust gas recirculation (EGR) technology creates opportunities for chemical interactions between transportation fuels and NOx. In this study, 2-ethylfuran (2EF), a representative component of furan-based biofuels, was selected to systematically investigate its reaction kinetic characteristics with those of NO2. Rate constants over a wide temperature range of 298-2400 K were calculated using multistructural canonical variational transition-state theory (MS-VTST) combined with a multidimensional tunneling correction method. The results show that the reaction rate exhibits a strong temperature dependence across the entire temperature range. This dependence is determined not only by the structure of fuel molecules but also by whether O-atom or N-atom attack occurs. In the low-temperature range (298-800 K), the combined effects of multistructural torsional anharmonicity, variational effects, and tunneling effects do not outweigh the influence of energy barrier height on reaction rates, causing the order of the rate constants to still follow the energy barrier height trend. In both NO2-addition and H-abstraction mechanisms, the reactions involving N-atom attack on 2EF are kinetically dominant. Notably, the tunneling effect plays a significant role in H-abstraction reactions in the low temperature range (T ⩽ 500K). This study identifies key kinetic factors governing the reaction of furan-based biofuels with NO2, thereby providing crucial theoretical support for refining the interaction mechanism between oxygenated furan biofuels and NOX under EGR conditions.
Multimodal foundation models demonstrate remarkable understanding ability to perform accurate molecular representation, including SMILES sequence-text and molecular graph-text representations. Molecular images represent another molecular modality that provides a more intuitive way to capture the spatial structural relationships of molecules. Existing models suffer from integrating image-text modalities for molecular property prediction, largely due to (1) the scarcity of high-quality image-text paired datasets, as existing molecular databases predominantly contain structural representations (SMILES, graphs) with limited descriptive text tailored for visual molecular understanding, and (2) the fundamental semantic gap between molecular images and textual descriptions, where complex visual features must be precisely aligned with domain chemical terminology and functional descriptions. In this study, we present a molecular image-text foundation model, named ITMol, pretrained on 500k molecular image-text pairs. Specifically, ITMol addresses the scarcity of high-quality image-text paired datasets by constructing a comprehensive dataset through collection from chemical databases and automated generation using MolT5 to create high-quality image-text pairs, and tackles the fundamental semantic gap between molecular images and textual descriptions by implementing a sophisticated cross-attention mechanism designed specifically for molecular modalities, coupled with three complementary self-supervised learning strategies. We demonstrate the high performance of ITMol in molecular property prediction across 8 benchmark datasets, achieving an ROC-AUC score of 0.885 on the BBBP dataset. ITMol shows high R@5 scores on a dataset of 100k molecules in the cross-modal retrieval task. We further illustrate interpretability of ITMol using key chemical structures and functional groups. ITMol provides an effective framework for integrating 2D molecular images and molecular textual descriptions. The results show that image-text multimodal pretraining can improve molecular representation learning for property prediction and retrieval, offering a useful direction for multimodal molecular modeling.
This paper analyzes microstructural layout and electrical behavior of silicon nanowire-based Schottky diodes, for use as wide-domain temperature sensors. The employed nanostructured three-dimensional substrates provide larger contact areas and enable higher Schottky barrier heights, ultimately leading to a better operable temperature range. Two metal deposition techniques (Radio Frequency sputtering and Electron-beam evaporation) are used to fabricate experimental Schottky diode samples. Scanning electron microscopy, X-ray diffraction, and diffuse reflectance investigations are carried out in order to determine nanowire distribution and the influence of subsequent metal deposition. The analyses evince the formation of a slightly inhomogeneous contact. The findings are validated by a thorough electrical characterization over a wide temperature domain. Inhomogeneity models are used in order to determine the main device parameters and the bias regions where they can be used as precise temperature sensors. The sputtered sample exhibits the best sensitivity, between 1 and 1.4 mV/K, while excellent linearity (R2 > 99.5%) is obtained for Electron-beam evaporated devices. Both types of silicon nanowire-based Schottky diode sensors have 100-500K operable ranges, much larger than planar counterparts.
We evaluated the data requirement for modern AI tools to outperform simpler models in predicting short-term mortality in over 500 000 patients with hemodialysis-dependent kidney failure. We compared logistic regression, boosting, and transformers using increasingly complex feature sets (from last-visit data to full trajectories). Performance was measured using the area under the ROC curve (AUC-ROC) and the Precision-Recall curve (AUC-PR) across training data sizes ranging from 500 to 490 197 samples. Using features with temporal information is beneficial across all models. On the full dataset, Transformers (AUC-ROC = 0.8568) and boosting (AUC-ROC = 0.8598) perform similarly. Transformers require large datasets to outperform simpler models like boosting, limiting their usefulness in smaller datasets, even on datasets as big as 500K. Modern AI tools require substantial data to justify their computational cost over simpler approaches. However, a more complex feature set seems to be beneficial across all models.
NAT2 is an important pharmacogene which encodes the N-acetyltransferase 2 enzyme that is involved in the metabolism of multiple medications, and variants in this gene can affect patient response to these medications. CPIC has published a clinical guideline for prescribing hydralazine using NAT2 genotypes. Just prior to the guideline, updated NAT2 star allele numbering and definitions were released, differing somewhat from the historical nomenclature. Clinical pharmacogenomic testing panels often test for the most common star alleles, so knowledge of the most common updated NAT2 star alleles is critical for the implementation of the CPIC NAT2/hydralazine guideline. We first determine NAT2 diplotype frequencies from UK Biobank (UKBB) 200k phased genomes, then analyzed allele, diplotype, and phenotype population frequencies from the All of Us Research program, PennMedicine BioBank (PMBB) and UKBB 500k datasets. We found that analyzing NAT2 diplotypes from phased data provides critical information for algorithms designed to predict diplotypes from unphased data. We observed that NAT2*5 , *6 , and *4 were the most common star alleles in that order, and the top 11 most frequent NAT2 star alleles were the same across all biobanks. However, differences in star allele frequencies across biogeographical populations were observed. The largest difference led to a higher frequency of NAT2 poor metabolizer phenotypes as compared to rapid and intermediate metabolizer phenotypes in all global populations except in the EAS population, where NAT2 poor metabolizers were in the minority.
Cell metabolomics, including lipidomics, presents several challenges regarding analyzing limited cell populations and distinguishing cellular metabolites from background signals originated from a stimuli or after a treatment. To address this, we have developed a novel workflow for untargeted cell lipidomics analysis. To study the impact of varying input cell numbers on the outcomes of untargeted cell lipidomics analysis, CD3+ cells isolated from a healthy donor at 6 different cell counts (50k, 100k, 250k, 500k, 750k, and 1M) were analyzed by liquid chromatography coupled with quadrupole-time-of-flight mass spectrometry (LC-QTOF-MS) in positive and negative electrospray ionization (ESI+ and ESI-, respectively) modes. After data quality assurance (QA), Spearman correlation analyses were carried out to select chemical signals derived from cells (ρ ≥ 0.7, p-value < 0.05). Then, this methodology was applied to human microvascular dermal endothelial cells (HMVEC-d), where a cell number calibration curve including 4 cell counts (25k, 50k, 75k, and 100k) was incorporated alongside the experimental samples to enable the analysis of cell-derived chemical signals. Here, the lipid response of HMVEC-d after contact with sera from patients at baseline and during the acute stage of anaphylaxis triggered by three different mechanisms was explored. For the CD3+ model, we found that although 1087 chemical signals (k) passed the QA, samples did not cluster according to their cell count when taking all signals into account. After correlation analyses, the widest cell count interval considered for correlation analyses (50k-to-1M; k = 70) showed clear clustering by cell number. The principal component analysis (PCA) models for ESI+ showed that for this cell count interval, the first component explained over 90% of the variance among samples. After applying the same methodology to HMVEC-d, we found k = 157 and k = 278 correlated chemical signals for ESI+ and ESI- in the cell curve (25k-100k). Statistical analysis identified 193 chemical signals that significantly (p-value < 0.05 and p-adjusted value < 0.2) differed between the acute and baseline stages of anaphylaxis. Without this correlation approach, 67 additional chemical signals would have been selected as significant. From the 193 chemical signals, 75 unique lipids were annotated, mainly including fatty acids, acyl carnitines, glycerophospholipids, and sphingolipids, all increased in the acute phase. These changes were associated with sphingolipid and glycosphingolipid metabolism, and ceramide and phospholipid signaling pathways. This workflow for cell lipidomics analysis allows the selection of lipids derived from the intracellular content regardless external sources, supporting specific intracellular metabolism profiling.
Extreme Multi-Label Text Classification (XMTC) is a crucial task in natural language processing, aiming to assign the most relevant subset of labels to an input text from an extremely large label set. Existing deep learning models often neglect the correlation information among labels when addressing XMTC, resulting in limited prediction performance. This paper proposes the Label Correlation Enhancement Network (LCENet), a pluggable modular architecture capable of effectively capturing and utilizing inter-label correlation knowledge without modifying the original model structure. The LCENet module introduces a bottleneck layer and residual connection mechanism, transforming raw label predictions into correlation-enhanced outputs. The bottleneck design reduces parameter complexity from [Formula: see text] to O(RL), effectively addressing computational feasibility under extreme-scale label spaces. We integrate the LCENet module into various mainstream deep XMTC models-including CNN-, BERT-, and RNN-based architectures-and conduct comprehensive evaluations on three benchmark datasets: EUR-Lex, AmazonCat-13K, and Wiki-500K. Experimental results demonstrate that LCENet significantly improves the performance of baseline models, achieving consistent gains across multiple metrics such as Precision@k and nDCG@k, with a maximum increase of 5.22 percentage points in P@1, while also accelerating model convergence. Ablation studies further verify the effectiveness of key components, including the bottleneck layer, residual connections, and nonlinear activation functions. Training curve analysis shows that LCENet provides stronger learning signals through label correlation constraints, alleviating the early-stage stagnation observed in some baseline models (Fig. 5) and reducing the required convergence steps by nearly half. This study presents a practical and effective enhancement framework for deep multi-label learning, offering excellent scalability and practical applicability. The core idea of LCENet can be extended to other structured prediction tasks such as multimodal learning and sequence labeling.
Two novel quaternary oxyarsenides, Eu8Zn2As6O and Eu14Zn5As12O, were synthesized through metal flux reactions, and their crystal structures were established by single-crystal X-ray diffraction methods. Eu8Zn2As6O crystallizes in the orthorhombic space group Pbca, featuring polyanionic ribbons composed of corner-shared triangular [ZnAs3] units, running along the [100] direction. The structure of Eu14Zn5As12O crystallizes in the monoclinic space group P2/m and its anionic substructure can be described as an infinite "ribbonlike" chain comprised of [ZnAs3] trigonal-planar units, although the structural complexity here is greater and also amplified by disorder on multiple crystallographic positions. In both structures, the O2- anion occupies an octahedral void with six neighboring Eu2+ cations. Formal electron counting, electronic structure calculations, and transport properties reveal the charge-balanced semiconducting nature of these heteroanionic Zintl phases. High-temperature thermoelectric transport properties measurements on Eu14Zn5As12O reveal relatively high resistivity (ρ500K = 8 Ω·cm) and Seebeck coefficient values (S500K = 220 μV K-1), along with a low concentration and mobility of holes as the dominant charge-carriers (n500K = 8.0 × 1017 cm-3, μ500K = 6.4 cm2/V s). Magnetic studies indicate the presence of divalent Eu2+ species in Eu14Zn5As12O and complex magnetic ordering, with two transitions observed at T1 = 21.6 K and T2 = 9 K.
p53 is the most important tumor suppressor in humans as well as the most frequently mutated gene found in human cancers with ~50% of all human tumors bearing p53 missense mutations that leave p53 inactive. Restoring the p53 activity proved to lead to tumor regression even in advanced tumors in mouse models- and thus, is among the most attractive potential strategies for novel cancer therapy. Full-length p53 (fl-p53) consists of 393 residues and multiple domains; some folded and some disordered. Using crystal structures of folded domains and integrative molecular modelling techniques for disordered domains, we generated the first wild-type fl-p53 tetramer model bound to DNA. When solvated, the system size nears 500K atoms challenging extensive sampling. Using Anton2 supercomputer for microsecond-timescale simulations in explicit solvent and the rigorous Markov state model (MSM) framework, we elucidated the conformational landscape of wild-type p53 as well as two of the p53 hot-spot cancer mutants, Y220C and G245S, in a physiological DNA-bound, full-length tetramer context. In the simulated timescale, DNA-bound fl-p53 tetramer bent DNA and formed a compact complex with interactions between the N-terminal and DNA-binding domains (DBDs), and the C-terminal domains (CTDs) with DNA. WT fl-p53 tetramer also sampled a unique quaternary DBD organization not accessed by the cancer mutants. Free energy landscapes indicated differential dynamics for inner and outer p53 DBDs due to the dimer-dimer interface. The dynamics of the druggable L1/S3 pocket is also closely monitored. Ultimately the MSMs identified an underexplored loop 6 (L6) cryptic pocket and captured the effect of p53 tetramerization and cancer mutations.
Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.
Nanocrystal-based light-emitting diodes (LEDs) are a promising technology for the next generation of flexible and large-area displays, offering high brightness, tunable narrow emission, and high display contrast. Although internal quantum efficiency (IQE) has reached unity, the external quantum efficiency (EQE) of LEDs utilizing spherical quantum dots (QDs) is limited by low light outcoupling efficiency (ηout). A promising approach to improve ηout is using horizontal alignment of nanocrystals' transition dipole moments, such as aligned quantum rods (QRs), which provide directional light emission. Though the IQE for QRs has recently improved significantly, creating efficient and bright green-emitting QRs (515-560 nm) remains a big challenge, which is essential for full-color display applications. In this study, a uniform and highly bright green-emitting CdSe/ZnxCd1-xS QRs of gradient shell structure are synthesized with minimized shell thickness and reduced Zn content, coupled with shorter organic ligands to reduce the energy barriers and enhance carriers injection. The electron leakage current at the interface between the polymeric hole transport layer (HTL) and QRs is the primary factor limiting the QRLED's performance. HTL with a higher energy offset is employed to prevent electron leakage at the organic/inorganic interface. Furthermore, is developed a bilayer HTL that enhances hole injections while minimizing electron leakage, thereby improving charge balance. The resulting QRLEDs demonstrate a record-high efficiency, with an EQE of 24%, current efficiency (CE) of 89 cd A-1, and maximum brightness (Lmax) exceeding 500k cd m- 2. Additionally, they exhibited an extended operational T50 lifetime of over 22k h at 100 cd m- 2, making them well-suited for high-color-gamut display and lighting applications.
A polymorphic variant in the ataxia telangiectasia-mutated ( ATM ) gene, rs56009889, was recently associated with an increased risk of lung cancer. We studied the role of this variant in the etiology of other cancers. Data from three population-based case-control studies of colon, breast, and lung cancer were used. Participants in these studies (4517 cases, 3383 controls) underwent a genome-wide association study using 500K Illumina OncoArray. The frequency of the AG/AA genotypes differed between Ashkenazi (4.6%) and Sephardi (0.2%) Jews ( P  < 0.001). AG/AA frequency was significantly higher in Ashkenazi lung cancer (11.9%) than in controls (2.8%) [adjusted odds ratio (OR) = 5.4]. Females had a higher risk than males (OR = 12.8 versus 3.5). The adjusted OR for colorectal cancer was 1.40 [95% confidence interval (CI) = 1.01-2.0, P  = 0.045] and for breast cancer was 1.43 (95% CI = 1.01-2.04, P  = 0.046). Never-smokers variant carriers were at higher risk of lung and colon, but not breast, cancer. Cases with the AG/AA genotype had lower mean age at diagnosis, but this difference was significant only for breast cancer (-3.2 years, P  = 0.007). No associations were observed with overall survival. Among the breast cancer subjects, the OR for having triple-negative tumors was 0.45 for AG/AA versus GG genotype (95% CI = 0.2-0.9, P  = 0.02). We confirm the strong association between ATM rs56009889 and lung cancer risk in Ashkenazi Jews and report a mild association with the risk of breast cancer and colorectal cancer.
Ensuring data integrity in cloud-edge environments is critical for IoT ecosystems but is challenged by dynamic data and resource constraints. This paper proposes a certificateless auditing scheme harmonizing cloud security with edge efficiency. By integrating online/offline cryptography and sparse Merkle trees, our approach achieves (1) significant user-side computation reduction via offline or edge-side tag generation, (2) [Formula: see text] dynamic update complexity versus traditional [Formula: see text] approaches, and (3) 75% communication overhead savings through pre-download mechanism. The scheme eliminates certificate management and mitigates Key Generation Centre (KGC) risks via decentralized trust mechanisms. Security proofs demonstrate resilience against KGC collusion and tag forgery under the Inv-CDH assumption. Experiments show our scheme audits faster than prior schemes, supporting 500k+ operations at sub-second latency. This work bridges scalability and real-time demands for smart cities and Industry 4.0 while enabling future extensions in ML-optimized caching and blockchain trust models.
In this paper, we present a symbolic dataset, named FIE-500k, for the second-kind Fredholm integral equations. Our approach systematically generates a dataset for these types of integral equations, enabling various applications in language models, such as modeling and symbolic solving of integral equations. Each record in the dataset includes the key terms of a second-kind Fredholm integral equation. To ensure the generation of valid mathematical expressions, we employ context-free grammars. The basis functions used in these grammars are carefully selected to cover a wide range of function types. The proposed dataset comprises 500,000 records, refined to ensure balance across different function types. A link to download the dataset, along with the code used for its generation, is provided in this paper. Researchers and practitioners can freely download the dataset, generate additional samples, and further enhance the proposed method.
Objective: Studies show a high prevalence of vitamin D deficiency in Greece and Cyprus despite an abundance of sunlight. We investigate the vitamin D status of Greeks and Cypriots living in the UK, where sunlight availability is more limited. Design: Cross-sectional study of serum 25-hydroxyvitamin D (25(OH)D) using the UK Biobank cohort. Setting: The UK Biobank is a study of over 500K UK dwelling participants, with baseline measurements from 2006-2010. Participants: A sample of 325 Greek/Cypriot and 4158 British/Irish participants (aged 40-69 years). Results: The Greeks/Cypriots had statistically significantly lower median serum 25-hydroxyvitamin D (25(OH)D) (40.3 nmol/L) compared to the British/Irish (47.6 nmol/L). Eleven percent of British/Irish and 22.8% of Greeks/Cypriots had serum 25(OH)D < 25 nmol/L. Being exposed to summer sunlight for >30 min/d, as well as having a blood draw in summer or autumn, was statistically significantly associated with lower odds of 25 (OH))D < 50 nmol/L. Living in Scotland, having a winter blood draw, and not using a vitamin D-containing supplement were associated with increased odds of 25(OH)D < 50 nmol/L. Ethnicity was not a predictor of 25(OH)D < 50 nmol/L after confounder adjustment (Greek/Cypriot OR = 1.18 (95% CI 0.85, 1.63; British/Irish OR = 1.0). Conclusions: UK dwelling Greeks/Cypriots have a higher prevalence of vitamin D deficiency (<25 nmol/L) compared to the British/Irish population, but evidence from the literature is mixed as to whether they have a higher prevalence than when living in their country of origin. Public health interventions are required to improve 25(OH)D status in UK ethnic minority groups.
The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries ( www.nemad.org ). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R2) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.
Back pain (BP) is a complex heritable trait with an estimated heritability of 40% to 60%. Less than half of this can be explained by known genetic variants identified in genome-wide association studies. We applied a powerful multi-trait and gene-based approach to association analysis of BP to identify novel genes associated with BP. Using phenotypes and imputed genotypes from the UK Biobank 500k dataset, we generated a multi-trait phenotype by combining 3 BP-related phenotypes: chronic BP, dorsalgia, and intervertebral disk disorders. We performed gene-based association analysis for 3 BP-related phenotypes and multi-trait phenotype. Conditional analysis was applied to account for the effects of genetic variants outside the gene. Finally, we replicated significantly associated genes using the FinnGen database. We identified 32 genes associated with BP and replicated 16 of them. Thirteen genes were detected using the multi-trait phenotype. Seven of the detected genes, MIPOL1, PTPRC, RHOA, MAML3, JADE2, MLLT10, and RERG, were not previously reported. Several new genes are known to be associated with traits genetically correlated with BP or to be involved in pathways associated with BP. Using new powerful methods of association analysis, we identified 7 novel genes associated with BP. Our results provide new insights into the genetics of back pain.
Understanding the complex relationships among genetic variations, brain structural anatomy, and functional alterations is a fundamental yet challenging task in neuroimaging genetics. In this study, we employ a causal mediation framework under structural modeling to systematically investigate the mediating role of interand intra-network structural connectivity (SC) in linking whole-genome single nucleotide polymorphisms (SNPs) to brain functional connectivity (FC) across both resting-state and task-based conditions during neurodevelopment. Utilizing baseline and follow-up SC and FC network traits along with ∼500k SNPs along the genome from 11,666 unique subjects under the Adolescent Brain Cognitive Development (ABCD) study, we first conduct genome-wide association studies (GWAS) to identify candidate SNPs associated with structural and functional network traits. Subsequently, mediation analyses reveal key genetic exposures that directly influence brain functional networks and indirectly impact FC through SC network mediators. These results provide deeper insights into how genetic variations shape brain structural and functional network organizations, along with revealing the influence of brain anatomical topologies on functional fingerprints. This work enhances the understanding of the causal effect pathways among genetic factors and large-scale brain structural and functional networks, advancing our understanding of the genetic underpinnings of neurodevelopmental processes.