Inborn errors of metabolism (IEMs) are rare, heterogeneous disorders traditionally diagnosed through genetic testing, enzyme assays, and metabolite measurements. However, these tools often do not fully explain phenotypic variability, organ involvement, disease progression, or treatment response. Clinical proteomics provides a complementary functional layer by capturing changes in protein abundance, proteoforms, post-translational modifications (PTM), and biological pathways, offering insights beyond genotype- and metabolite-based approaches. This review examines the role of high-resolution mass spectrometry and computational proteomics in biomarker discovery and clinical decision-making for IEMs. It focuses on their contribution to diagnosis, variant interpretation, patient stratification, and treatment monitoring. Disease-specific applications are discussed, with the strongest evidence in lysosomal storage disorders, mitochondrial diseases, congenital disorders of glycosylation, and selected neurodegenerative or renal metabolic conditions. The literature search was performed in PubMed, Scopus, Web of Science, and Google Scholar, covering peer-reviewed articles available up to 2026, with emphasis on methodological advances and translational applications in clinical proteomics for IEMs. Proteomics will not replace established diagnostic tools, but it can help address clinically actionable questions in selected contexts. Translation into clinical practice will require standardized workflows, multicenter validation, clinically anchored endpoints, and integration with other omics approaches.
With the advancement of proteomics technologies, an increasing number of studies have begun to examine semaglutide-associated protein expression changes and pathway alterations across different biological contexts. However, existing evidence remains fragmented across different disease backgrounds, sample types, and research platforms, lacking systematic integration. This review searched PubMed, Embase, and Web of Science for relevant studies published as of April 2026, including population-based, animal, and in vitro model studies that implemented semaglutide interventions and reported proteomic results. A total of 16 studies were ultimately included, comprising 4 population-based studies and 12 animal and in vitro model studies. The included studies examined both circulating samples and tissue-level specimens, including serum, plasma, adipose tissue, myocardium, aorta, lung, kidney, hippocampus, and skeletal muscle. Common proteomics platforms utilized included SomaScan, TMT-LC-MS/MS, DIA proteomics, phosphorylation proteomics, and mitochondrial proteomics. Multiple studies identified the PPAR signaling pathway, oxidative phosphorylation, and fatty acid metabolism as frequently occurring pathways, while ECM remodeling, complement/inflammatory pathways, and mTORC1 signaling were also observed in some studies. These results suggest that semaglutide is associated with proteomic changes across metabolic, cardiovascular, hepatic, pulmonary, renal, and nervous systems, with recurring signals involving fatty acid metabolism, mitochondrial function, inflammatory protein networks, and extracellular matrix-related pathways. Overall, proteomic evidence provides a useful molecular framework for describing semaglutide-associated biological responses across multiple tissues. However, existing studies still have limitations such as small sample sizes, high subject heterogeneity, significant differences in proteomic platforms, and limited population studies. Therefore, future research will require larger sample sizes, standardized designs, and multi-omics integration studies to further determine which proteomic signatures are reproducible, biologically meaningful, and clinically translatable.
The Metaproteomics Initiative was officially launched in 2021 to strengthen collaboration, promote knowledge exchange, and support and lead standardization efforts within the growing metaproteomics community. Over the past 5 years, the Initiative has developed into a structured, global network of researchers. It has launched community-driven benchmark studies, helped shape emerging metadata and reporting standards, developed practical guidance and training materials, organized international symposia, and fostered connections across the microbiome research landscape ( https://metaproteomics.org/ ). We outline the Initiative's organization, activities, achievements, and ongoing efforts, and reflect on how sustained, community-led coordination has shaped the development of metaproteomics as a field. We further position the Grand Metaproteome Challenges as a next step toward coordinated, community-scale biological research, aimed at advancing functional microbiome studies across clinical, industrial, and environmental application domains, and invite engagement from the wider microbiome and omics communities. Video Abstract.
Venomous Lepidoptera constitute an underrecognized yet medically significant group of toxin-producing arthropods that employ contact-mediated defensive envenomation through specialized integumentary structures such as setae, spines, and scoli. Unlike actively stinging arthropods, these insects deliver venom passively upon contact, eliciting a diverse spectrum of clinical manifestations collectively termed lepidopterism. Clinical outcomes range from localized pain and dermatitis to severe systemic effects, including hemorrhagic syndromes, complement activation, and chronic inflammatory disorders. Recent advances in proteomic and transcriptomic technologies have transformed our understanding of lepidopteran venoms, revealing unexpectedly complex toxin repertoires comprising serine proteases, phospholipases, pore-forming proteins, disulfide-rich peptides, neuroactive RF-amide peptides, and immune-modulating components. These findings have provided new insights into the molecular basis of toxicity, host-pathogen interactions, and the evolutionary diversification of venom systems within Lepidoptera. This review synthesizes current knowledge on the morphology of venom-delivery structures, venom composition, mechanisms of action, and associated clinical manifestations, while highlighting medically important taxa, particularly species of the genus Lonomia. The successful development of antivenom against Lonomia envenomation underscores the translational relevance of lepidopteran toxin research and its potential for therapeutic innovation. By integrating molecular, clinical, and evolutionary perspectives, this review repositions venomous Lepidoptera as a legitimate and important component of arthropod toxinology. Furthermore, it identifies critical methodological limitations and key knowledge gaps, providing a framework for future investigations aimed at advancing our understanding of toxin biology, immunopathology, and the development of novel biomedical applications.
To identify preconception clinical and multi-omics factors associated with conception and early pregnancy loss in women with unexplained recurrent pregnancy loss (URPL). In this prospective cohort study, 149 women with URPL selected from 420 outpatients based on guideline-recommended criteria were enrolled between November 2024 and May 2025 and followed-up for 12 months. Preconception fasting plasma was analyzed for clinical biomarkers, untargeted metabolomics, and data-independent acquisition proteomics. Outcomes included conception, ongoing pregnancy beyond 12 weeks and early pregnancy loss before 12 weeks. Multivariable logistic regression was used to evaluate the associations of clinical and multi-omics factors with reproductive outcomes, adjusting for maternal age, body mass index, number of prior losses, and use of assisted reproductive technology. Of 149 women, 99 conceived (66.4%) during the follow-up. By 12 weeks of gestation, 67 (67.7%) had ongoing pregnancies and 32 (32.3%) experienced early pregnancy loss. Higher testosterone was associated with lower probability of conception (adjusted odds ratio [aOR] 0.50, 95% confidence interval [CI] 0.28-0.89, p = 0.019). Women who conceived showed higher levels of progesterone-related metabolites, including 17-hydroxyprogesterone (fold change, FC = 3.89), pregnanediol 3-O-glucuronide (FC = 2.37), and pregnanetriol 3α-O-β-D-glucuronide (FC = 2.39). Among women who conceived, higher prolactin was associated with higher odds of early pregnancy loss (aOR 1.09, 95% CI 1.01-1.18, p = 0.036), and anti-phosphatidylserine/prothrombin showed a borderline association (aOR 1.07, 95% CI 1.00-1.14, p = 0.052). Early pregnancy loss was characterized by lower bile acid-related metabolites and higher caffeine and methylxanthine metabolites. Proteomic analysis showed enrichment of bile acid biosynthetic process. After adjustment, higher 7α,12α-dihydroxy-3-oxocholest-4-en-27-oic acid was associated with lower odds of early pregnancy loss (aOR 0.64, 95% CI 0.44-0.95, p = 0.025). Higher testosterone was associated with lower odds of conception. Among women who conceived, early pregnancy loss was associated with higher prolactin, borderline higher anti-PS/PT, lower bile acid-related metabolites, and higher caffeine-related metabolites, with 7α,12α-dihydroxy-3-oxocholest-4-en-27-oic acid identified as an exploratory metabolomic candidate. These findings suggest that potential androgen-related biology, prolactin, non-criteria antiphospholipid antibodies, and bile acid metabolism may be associated with reproductive outcomes in URPL, warranting validation in independent cohorts.
This study aimed to identify plasma biomarkers associated with systemic sclerosis-related interstitial lung disease (SSc-ILD) and progressive pulmonary fibrosis (SSc-PPF) using an unbiased proteomic approach, and to assess their predictive value across independent cohorts. Plasma from 202 patients with SSc was analysed, including a discovery cohort (n = 45) and a replication cohort (n = 157) from 3 centres (Paris, Bordeaux, Leeds). Patients were stratified by ILD status (SSc without ILD [SSc-NoILD] vs SSc-ILD) and ILD progression (SSc with nonprogressive pulmonary fibrosis [SSc-NoPPF] vs SSc-PPF), with SSc-PPF defined according to INBUILD criteria. Mass spectrometry-based proteomics identified candidate biomarkers in the discovery cohort. Top candidates were evaluated by enzyme-linked immunosorbent assay (ELISA) in the replication cohort. Differential expression, pathway enrichment, and predictive performance analyses were performed. A total of 1507 proteins were quantified, of which 44 and 51 were associated with SSc-ILD and SSc-PPF, respectively. Five candidates progressed to ELISA validation, among which surfactant protein D (SP-D) and matrix metalloproteinase 8 (MMP-8) were successfully replicated. In the combined cohort, SP-D was elevated in SSc-ILD vs SSc-NoILD (17.36 ± 11.99 vs 9.49 ± 7.92 ng/mL; P < .0001). MMP-8 was higher in SSc-PPF vs SSc-NoPPF (17.67 ± 12.45 vs 10.77 ± 10.92 ng/mL; P < .001). Its prognostic value for PPF was further replicated in another cohort of 113 patients from Tokyo and Milan. In combined analyses, MMP-8 independently predicted SSc-PPF and improved performance when added to clinical variables. SP-D and MMP-8 emerge as candidate biomarkers for diagnosing SSc-ILD and predicting SSc-PPF, respectively. Their validation across multicentric cohorts supports their relevance for clinical risk stratification.
Liquid biopsy has become a revolutionary method for the early detection of cancer as a non-invasive technology that can assess circulating tumor material in biofluids. Liquid biopsy allows dynamic monitoring of tumor evolution, genetic changes and treatment responses, which is different from traditional tissue biopsy which offers a static and potentially narrow view of tumor biology. This mini-review will summarize the rapidly evolving future of circulating biomarkers (circulating tumor cells (CTCs), circulating tumor DNA (ctDNA), microRNAs, proteins, exosomes and epigenetic fingerprints), with their potential for early multi-cancer biomarkers, and their integration into early detection. The analytical sensitivity of liquid biopsy has expanded dramatically through technologies such as next-generation sequencing (NGS), digital PCR, and advanced proteomics. In addition, data collection has been enhanced through machine learning for increased predictive performance and the identification of new biomarkers. This review discusses clinical performance between liquid and tissue biopsy and the value of combined biomarker methods to improve accuracy for the detection of early disease. Finally, future directions will be presented to identify new methods of integration, improved costs and the establishment of early detection programs at the population scale.
Stroke is a major complication of atrial fibrillation (AF), and risk prediction using the congestive heart failure, hypertension, age, diabetes, stroke, vascular disease, and sex category score (CHA₂DS₂-VASc) remains limited by residual heterogeneity. We aimed to identify plasma proteins associated with post-AF stroke and evaluate whether a protein score provides incremental predictive information beyond CHA₂DS₂-VASc. We analyzed 709 AF participants from the UK Biobank Pharma Proteomics Project, with 76 incident strokes. Stroke-related proteins were identified using multivariable Cox regression, least absolute shrinkage and selection operator (LASSO) Cox regression, and machine-learning approaches. A five-protein score was constructed, and its incremental value beyond CHA₂DS₂-VASc was assessed by discrimination, calibration, reclassification, clinical net benefit, and 1000-bootstrap internal validation. Mendelian randomization served as supportive genetic evidence. Five core proteins were identified: epidermal growth factor receptor (EGFR), V-type proton ATPase subunit D (ATP6V1D), neurotrophin 4 (NTF4), amnionless (AMN), and discoidin, CUB and LCCL domain-containing protein 2 (DCBLD2). Adding the five-protein score to CHA₂DS₂-VASc improved discrimination, increasing the concordance index from 0.681 to 0.768. The combined model had 3-, 5-, and 8-year receiver operating characteristic areas of 0.799, 0.775, and 0.806, respectively, and showed favorable five-year prediction error and calibration, with a Brier score of 0.0347, calibration intercept of -0.068, and calibration slope of 0.968. The score improved continuous net reclassification improvement (0.459; P = 0.028). Mendelian randomization provided supportive genetic evidence for AMN, EGFR, and DCBLD2. The five-protein score provided incremental predictive information beyond CHA₂DS₂-VASc for post-AF stroke risk assessment.
To explore temporal patterns of serially measured cardiovascular-related biomarkers in patients with chronic heart failure (CHF), with the objective of identifying biomarker pathophysiological trajectories associated with clinical outcomes. These explorative analyses can generate hypotheses regarding underlying pathophysiological processes and the potential role of multi-biomarker approaches in chronic HF. The BioMEMS-study involved 334 patients with moderate to severe chronic HF in NYHA class III who received either standard of care or remote hemodynamic monitoring. Serial blood samples were collected at baseline, 3, 6, and 12 months, and biomarker levels were assessed using the Olink Cardiovascular-III panel. Joint modelling analyses were performed, integrating longitudinal biomarker trajectories and risk of the composite endpoint of all-cause mortality or HF hospitalization. In multivariable-adjusted models, 15 biomarkers were consistently and significantly associated with the composite endpoint after adjustment for confounders and multiple testing. MMP-2, ST2, IGFBP-1, IGFBP-7, and NT-proBNP exhibited the strongest associations with the composite endpoint, with hazard ratios (95% CI) of 2.72 (1.87-4.03), 2.71 (2.00-3.80), 2.70 (1.81-4.27), 2.48 (1.81-3.45), and 2.28 (1.72-3.03), respectively. These findings were robust across sensitivity analyses correcting for clinical confounders and treatment groups. Temporal trajectories revealed higher biomarker levels in patients who experienced incident events, with biomarkers showing corresponding changes in levels preceding events. Serial biomarker measurements could provide additional insights in the pathophysiology of worsening HF. Specific biomarkers reflecting myocardial stress, cardiac remodelling and fibrosis show changes in levels over time as worsening HF approaches, which highlights possible involvement of these pathophysiological pathways.
Candida glabrata (currently classified as Nakaseomyces glabratus) is an opportunistic fungal pathogen notable for its intrinsic antifungal tolerance and ability to persist in host environments. Although strain CBS138 has served as the principal model for genetic and functional studies, accumulating evidence indicates substantial intraspecies diversity that may shape virulence, immune interactions and stress adaptation. In particular, the widely used clinical isolate BG2 differs from CBS138 in genome structure, adhesin regulation and macrophage survival, yet the extent to which these differences are reflected at the fungal cell surface remains unknown. Here, we present a comparative characterization of the surface-exposed proteomes (surfaceomes) of CBS138 and BG2 across three biologically relevant growth conditions: YPD-grown yeast-like cells, RPMI-cultured planktonic aggregates and RPMI-formed biofilms. Using trypsin shaving combined with LC-MS/MS, we identified pronounced strain- and condition-dependent differences in surface protein composition, encompassing adhesins, yapsin proteases and selected moonlighting proteins. Whereas CBS138 showed greater representation of adhesion- and interaction-related surface proteins, BG2 preferentially displayed proteins associated with cell-wall architecture and remodelling, consistent with distinct surface-mediated adaptive strategies. Transmission electron microscopy revealed condition-dependent differences in cell-wall thickness in both strains, with BG2 displaying a broader range of values and the highest thickness under biofilm conditions, providing structural context for variation in protease accessibility and surface-protein detectability. Collectively, our findings highlight substantial surfaceome plasticity in C. glabrata and underscore the importance of considering intraspecies diversity when interpreting host-pathogen interactions and fungal virulence pathways.
This study focuses on alliin, a bioactive compound derived from garlic, and investigates its underlying mechanisms through a proteomics-based approach. In HepG2 cells, alliin significantly reduced intracellular total cholesterol and triglyceride levels. Tandem mass tag-based quantitative proteomic analysis identified 67 differentially expressed proteins (53 upregulated, 14 downregulated). Bioinformatics enrichment revealed these proteins were primarily involved in cholesterol metabolism, the PPAR signaling pathway, and the PI3K/Akt pathway. Notably, alliin significantly upregulated low-density lipoprotein receptor (LDLR) expression while downregulating proprotein convertase subtilisin/kexin type 9 (PCSK9). Western blot validation confirmed that alliin modulates the PCSK9/LDLR pathway. In conclusion, alliin influences intracellular cholesterol metabolism in HepG2 cells, potentially through regulation of the PCSK9/LDLR pathway. These findings provide mechanistic insights at the cellular level and suggest that alliin may serve as a candidate compound for further investigation in the context of lipid metabolism-related disorders.
Amyotrophic lateral sclerosis (ALS) is considered a highly complex, heterogeneous, fatal disease with a high unmet medical need that affects multiple pathophysiological pathways and has no known singular cause. Oxidative stress, however, is implicated as a central player in the progression of ALS and other neurodegenerative diseases. To date, only two FDA-approved drugs, Edaravone, an antioxidant, and Riluzole, an antiglutamatergic, have been widely used clinically, albeit with modest effects on the clinical course of ALS disease progression. Additionally, preclinical studies of both drugs in ALS mouse models have not shown any significant survival benefit, although some abatement in disease progression is observed and translated from preclinical animal models to clinical human studies. Using a trifunctional boron-based drug design strategy, we have synthesized a pyrazole small molecule called Borsantrazole (BSZ) and evaluated BSZ in cell-based experiments, acute and chronic toxicity models, and the SOD1-G37R ALS mouse model. We have also evaluated untargeted global proteomic and phosphoproteomic changes induced by BSZ. Borsantrazole (a small molecule that can selectively target oxidative stress) with favorable CNS drug like properties (low ER values of 0.9 and 1.0 at 1 and 10 µm, respectively, suggesting lower Pgp efflux liability and the LogD at pH 7.4 was within the range of 1.2 to 3.1, suggesting BSZ is lipophilic at pH 7.4.) shows no signs of treatment associated toxicity (acute or chronic (120 days)), significantly increases survival (whole animals, 15.1 days (p = 0.0326); males, 16.5 days; females, 13.7 days), rescued weight loss (27.1% control, 18.3% BSZ, p < 0.0001), delays symptom onset (whole animals, 24.9 days (p = 0.0011); males, 23.3 days; females, 26.4 days), delays disease onset (whole animals, 25.5 days (p = 0.0001); males, 21.1 days; females, 29.8 days) and affects global proteomic changes in the SOD1-G37R mouse model of ALS. In total, 51 proteins were found to be significantly differentially expressed (p < 0.05), including Ca3, Gan, Cplx2, Lrp4, Sqstm1, and 29 phosphorylation sites were differentially expressed and considered statistically significant, including T317 and T72 for neurofilaments light and heavy chain, respectively. Herein, we report that BSZ has demonstrated a favorable safety profile and compelling proof-of-concept efficacy in a ALS mouse model and has the potential to become a disease-modifying ALS therapeutic following further clinical development. Within a broader perspective of treatments for neurodegenerative diseases, BSZ offers a new paradigm for trifunctional small molecule targeting of oxidative stress that can mitigate neuronal deterioration and serve as a potential treatment.
Currently, there are no blood biomarkers available for the early diagnosis of laryngeal squamous cell carcinoma (LSCC). The objective of this study was to search for potential plasma biomarkers for the diagnosis of LSCC. Plasma samples were taken from patients with LSCC and healthy controls for proteomic analysis. An enzyme-linked immunosorbent assay (ELISA) was employed to measure the expression levels of differentially expressed proteins in plasma. The expression of differential protein in LSCC cells was downregulated. Subsequently, the function of LSCC cells was evaluated using wound healing, transwell migration, incorporation of 5-ethynyl-2'-deoxyuridine (EdU), and colony formation assays. Analysis of Plasma IGFBP2 Levels in Relation to Clinical Data and Evaluation of IGFBP2 as a Diagnostic and Prognostic Biomarker Using TCGA Database. A total of 16 differentially expressed proteins were identified in the plasma from patients with LSCC and healthy controls. Significant differences in protein expression patterns were observed between the LSCC patient group and the healthy control group. The expression of insulin-like growth factor binding protein 2 (IGFBP2) in plasma of LSCC patients was significantly higher than that of healthy controls. After suppression of IGFBP2 expression, the migration, invasion, and proliferation abilities of LSCC cells were reduced and the expression of signal transduction and signal transducer and activator of transcription 3 (STAT3) was negatively regulated. The analysis of IGFBP2 expression in the TCGA-LSCC data set showed a significant upregulation in tumor tissue compared to adjacent normal tissue. Furthermore, no significant differences in plasma IGFBP2 levels were observed with regard to gender, smoking status, history of malignancy, or tumor subsite. A positive correlation was observed between T stage and IGFBP2 levels. Kaplan-Meier survival analysis revealed that higher IGFBP2 expression levels were significantly associated with worse overall survival (OS) and progression-free survival (PFS). Univariable Cox regression analysis identified high IGFBP2 expression as a significant predictor of poor prognosis (hazard ratio [HR] = 3.52, 95% confidence interval [CI]: 1.84-6.75, p < 0.001). Furthermore, multivariable Cox regression analysis confirmed the independent prognostic value of IGFBP2 expression, with a hazard ratio (HR) of 3.32 (95% CI: 1.63-6.78, p < 0.001). IGFBP2 may play a role in the occurrence and development of LSCC as an oncogene and may be used as a potential plasma biomarker for the diagnosis of LSCC. IGFBP2 may be related to STAT3 in mechanism.
Breast cancer (BC) is a highly complex and heterogeneous malignancy and the most prevalent cancer among women worldwide. The diagnosis, prognosis, and the treatment of BC pose significant challenges that are responsible for their limited therapeutic efficacy. Omics-based technologies have gained substantial attention in BC diagnosis through molecular profiling and diverse clinical analytics. The integration of metabolomics, proteomics, transcriptomics, and genomics provides a multidimensional approach to personalized BC diagnosis and treatment through high-throughput molecular profiling. Moreover, the emergence of artificial intelligence (AI) has also supported more accurate and early diagnosis of BC through multimodal integration of diverse datasets. The integration of advanced deep learning (DL) and machine learning (ML) has been extensively exploited for tumor grading, histopathological classification, molecular profiling, diagnostic imaging, and prognostic prediction. This review aims to summarize recent developments in AI-driven multi-omics approaches for the discovery of BC biomarkers. We have also highlighted the integration of omics-based data like metabolomics, proteomics, transcriptomics, and genomics with key AI techniques, including ML and DL, that play a crucial role in the inclusion of multi-omics in cancer and biomarker discovery. We have further discussed AI-based BC screening and diagnostic approaches, as well as the contribution of AI models for patient stratification, biomarker discovery, and prediction of therapeutic response. Additionally, key limitations and challenges, including data heterogeneity, high computational complexity, and model interpretability, have also been highlighted in the present review. Conclusively, we have also outlined future perspectives on the integration of AI and multi-omics to revolutionize precision clinical medicine and improve clinical outcomes in BC theranostics.
Acute pancreatitis (AP) lacks reliable early biomarkers for predicting progression to severe acute pancreatitis (SAP). While macrophages are central to AP pathogenesis, the specific functional role of trypsinogen-activated macrophage-derived exosomes (Mφ-EXOs) remains inadequately defined. This study investigates whether these Mφ-EXOs and their cargo proteins propagate the inflammatory cascade and serve as early diagnostic assessor for AP severity. Macrophages were stimulated with trypsinogen in vitro, and the derived exosomes were isolated and characterized. Macrophage polarization and cytokine secretion profiles were evaluated following exosome treatment. Label-free quantitative proteomics was utilized to identify differentially expressed exosomal proteins. In vivo effects were validated using a cerulein-induced murine AP model. Clinically, plasma exosomal Rap1a levels were quantified for assessing AP severity, and diagnostic performance was assessed via receiver operating characteristic (ROC) analysis. Macrophage internalization of trypsinogen induced a pro-inflammatory M1 phenotype. Trypsinogen-activated Mφ-EXOs promoted M1 polarization and enhanced IL-1β and TNF-α secretion in recipient macrophages. Proteomics identified Rap1a as significantly upregulated in these exosomes; its knockdown attenuated exosome-mediated M1 polarization. In vivo, administration of these exosomes exacerbated pancreatic and pulmonary injury. Clinically, plasma exosomal Rap1a was markedly elevated in SAP patients compared to those with non-severe AP. A predictive model combining exosomal Rap1a with serum calcium, white blood cell count, and IL-1β demonstrated high diagnostic accuracy. Trypsinogen-activated macrophages contribute to systemic inflammation in AP via the release of Rap1a-enriched exosomes. Plasma exosomal Rap1a demonstrates potential clinical utility as an early, mechanistically-grounded biomarker for predicting SAP severity.
ObjectiveTo review the application of crowdsourcing and machine learning contests in Parkinson's disease (PD) research, identify best practices for successful implementation, and highlight future opportunities.MethodsThis paper analyzes the landscape of crowdsourcing in PD research through a literature survey and a comparative case study of two major machine learning contests: the MJFF Freezing of Gait (FOG) Challenge and the AMP PD Proteomics Challenge. We also describe a taxonomy of crowdsourcing projects and a framework of success characteristics for machine learning contest design.ResultsThe analysis of previous crowdsourcing and machine learning contests revealed that contest success is highly dependent on specific design factors. The FOG challenge, which addressed a "solvable but not yet solved" problem with a suitable scoring metric, successfully produced a high-performing algorithm with real-world clinical value. In contrast, the Proteomics challenge did not yield biologically meaningful results, as winning models bypassed the core proteomic data, highlighting issues of data signal and metric selection. The review also identified underutilized crowdsourcing approaches in PD research, including gamification and community-based open-source development.ConclusionsMachine learning contests offer a powerful, open-science-aligned method to address complex problems in PD. Success requires careful design, particularly a solvable problem and an appropriate scoring metric. There is significant potential to expand the use of diverse crowdsourcing techniques to accelerate progress in PD research and clinical care. This review explores the use of crowdsourcing techniques in the Parkinson's disease (PD) research ecosystem, with a focus on machine learning contests. We highlight best practices in designing crowdsourcing programs for successful research outcomes and provide an overview of opportunities for the PD research community.
High-throughput shotgun proteomics is often hindered by the incompatibility of detergents with mass spectrometry (MS), making sample preparation a critical bottleneck for accuracy and robustness. This challenge is amplified in muscle proteomics, where the high dynamic range of protein abundance requires highly efficient and reproducible workflows to capture low-abundance proteins. To address this, we performed a systematic benchmarking of five preparation methods-stacking-gel (SG), tube-gel (TG), solid-phase extraction (SPE), filter-aided sample preparation (FASP), and suspension traps (S-TRAP)-using the sarcoplasmic fraction of pig muscle. The protein extracts obtained were subjected to label-free semi-quantitative proteomic analysis using high-performance nano-liquid chromatography coupled to tandem MS. Qualitative and quantitative results were compared using bioinformatics and biostatistics tools. Our study identified 530 proteins with significant variations across methods. S-TRAP provided the highest identification depth, capturing the broadest proteome coverage. Conversely, the TG method demonstrated superior quantitative reproducibility, essential for detecting subtle physiological changes. For translational research, such as meat quality science or muscle-related clinical models, our findings imply that S-TRAP is the preferred choice for discovery-phase proteomics (biomarker hunting), while TG or SG should be prioritized for high-precision validation studies.
Autoimmune encephalitis (AE) represents a heterogeneous group of disorders characterized by immune-mediated attacks on neuronal antigens within the central nervous system (CNS), leading to diverse clinical manifestations and posing significant challenges in diagnosis and treatment. Recent advances in multi-omics technologies-including genomics, transcriptomics, proteomics, metabolomics, and immunomics-are beginning to provide unprecedented opportunities to unravel the complex immunopathophysiology of AE at a systems level. This review critically examines the application of integrated multi-omics approaches in AE research, with particular emphasis on immunological mechanisms underlying disease pathogenesis. We focus on: (i) the identification of potential diagnostic and prognostic biomarkers with rigorous validation frameworks; (ii) the elucidation of molecular heterogeneity underlying clinical variability, including B-cell clonal expansion, small-cohort evidence for T follicular helper (Tfh) expansion with inferred T follicular regulatory (Tfr) imbalance, and cytotoxic CD8+ T-cell-mediated neuronal injury in paraneoplastic/intracellular-antigen contexts; (iii) the conceptual development of precision medicine strategies tailored to individual immunophenotypic profiles; and (iv) the critical appraisal of current evidence quality and methodological limitations. By synthesizing current findings, this article aims to establish a comprehensive immunological framework that supports the future advancement of personalized therapeutic interventions. The integration of multi-omics data offers a promising conceptual lens through which to understand the intricate immune-neuronal interactions driving AE, though substantial challenges in data standardization, clinical validation, and implementation remain to be addressed.
Differential diagnosis of pneumonia, tuberculosis, and lung cancer is highly challenging due to overlapping clinical presentations and high comorbidity rates. To overcome the limitations of traditional diagnostic methods, this study developed and validated an explainable machine learning model using proteomic data for multi-label classification of these three diseases. Bronchoalveolar lavage fluid proteomic data were collected from 358 patients with confirmed lung diseases at the Shandong Provincial Public Health Clinical Center. We constructed a voting ensemble model integrating XGBoost, Random Forest, and Gradient Boosting algorithms based on clinical features, global proteomic statistical features, differentially expressed derived features, and disease-specific biomarker scores. Considering insufficient cancer samples and risk of missed diagnoses, a 1.5-fold weighting strategy was applied to the cancer class to enhance sensitivity. The model was evaluated using 5-fold stratified cross-validation and an independent external cohort of 110 cases. Feature contributions were interpreted using the Shapley Additive exPlanations (SHAP) method and biological significance of key proteins was determined using Gene Ontology functional enrichment analysis. The AUC values were 0.912±0.031 for tuberculosis detection and 0.813±0.037 for cancer detection. Thus, the ensemble model demonstrated excellent performance, significantly outperforming seven baseline models including logistic regression and support vector machines. SHAP analysis identified key protein biomarkers (Cancer: P02775, P61626; Tuberculosis: P0DOX2, P55259). Furthermore, the model achieved an overall accuracy of 86% with the independent external validation cohort. This study established an explainable, multi-label classification model based on proteomics that can provide a valuable reference for the differential diagnosis of complex lung diseases, especially those with comorbidities. The model shows good performance and interpretability, suggesting potential for precise diagnosis of lung diseases.
The foetal testes produce the androgens necessary to masculinise the developing embryo and support the maturation of germ cells, that will eventually develop into sperm, thus ensuring future reproductive capacity. The testes develop from the bi-potential gonads in a highly orchestrated process resulting in the differentiation of a complex tissue with multiple cellular lineages. While recent transcriptomic and chromatin-based analyses of human foetal testes have provided an unprecedented level of insight into signalling pathways activated during this process, proteomic studies of the human foetal gonads remain limited. Proteins are active molecules and post-translational modification (PTM) of proteins influences protein activity, stability and localisation. Studies have shown that PTMs regulate critical proteins in testis development, and their disruptions are implicated in congenital disorders including differences of sex development (DSD), in which sex development is atypical. Despite this, the role and regulation of protein PTM during human testis development remains poorly understood due to limited access to human foetal gonadal tissue, a paucity of large-scale proteomics studies, and a lack of robust of human gonad in vitro models. This review aims to provide a comprehensive analysis of validated PTMs affecting proteins critical for testicular development. We discuss PTMs with evidence for a role in normal testis development, and highlight those disrupted in DSD. We review emerging techniques, including proteomic technologies and organ modelling systems that may advance our understanding of PTMs in foetal testis development. We discuss challenges that have restricted the application of these technologies and how overcoming these will significantly improve our understanding of testis development and disease, diagnostics and patient outcomes. We searched PubMed and the University of Melbourne library for peer-reviewed English-language studies using keywords such as phosphorylation, SUMOylation, acetylation, ubiquitination alongside each protein of interest. PTM sites in proteins involved in testis development were identified using the PhosphoSitePlus database focusing those confirmed in in vitro or animal model studies. ClinVar and the Human Gene Mutation Database were used to identify patient variants that may disrupt PTM sites. Our review finds that proteins required for human foetal testis development are subject to extensive PTM. Several PTM sites and PTM-mediated pathways [e.g. MAPK (mitogen-activated protein kinase) pathway] are disrupted in patients with DSD or related conditions. While recent advances in proteomics technologies hold considerable promise, their application to human foetal gonads has been constrained by technical, ethical, and logistical challenges. Encouragingly, emerging high-sensitivity and low-input technologies, alongside stem cell-based approaches, offer viable pathways to overcoming these barriers. The relationship between gene regulation, protein expression, and cellular outcome is inherently non-linear, shaped by additional regulatory layers-most notably PTMs. The contribution of PTMs to human testis development in both typical and atypical contexts is a major knowledge gap. Addressing this gap has broad clinical and biological relevance: it may help improve genetic diagnosis or shed light on how proteins or pathways critical for testis development respond to environmental signals-an increasingly pressing question as declining global fertility rates bring testicular function under greater scrutiny. N/A.