To leverage the evaluation of the environmental impact of the Direct AP-HP/Lorah e-referral service offered in Paris (France) hospitals to create recommendations for evaluating the impact of digital health services. We review the tools and methods currently available to measure the carbon footprint and electricity consumption of digital services in the context of Life Cycle Assessment (LCA) for a comprehensive evaluation. We use the recent deployment of a telemedicine communication service at a major French hospital as a case study to understand the practical implications of conducting an impact study. Three-deployment scenarii are considered: current usage, double usage and maximum capacity. The bulk of the carbon footprint of the Direct AP-HP/Lorah service is due to servers vs network and user terminals in all scenarios considered. Computing hardware production impacts was instrumental in the overall impact assessment, as embodied impact represent 45% of Carbon footprint and the most of Metallic resource depletion. Recommendations for further studies notably include adequate anticipation of service usage and data collection. The environmental impact of the new telemedicine service could be assessed in sufficient level of details to provide decision makers with an adequate comparison of the service with alternative email communication. The recommendations derived from this use case should facilitate adequate impact data collection for future studies.
To assess prevalence, characteristics, and institutional predictors of clinical informatics (CI) education and student organizations in US allopathic (MD) and osteopathic (DO) medical schools. We reviewed 222 US medical schools from the 2024 Medical School Admission Requirements (MSAR) and American Association of Colleges of Osteopathic Medicine (AACOM) databases. Using predefined criteria, we identified CI-related courses and student groups and abstracted institutional characteristics including affiliated CI fellowships. Bivariate and multivariable logistic regression identified predictors. Of 222 schools, 30.2% offered at least one CI course and 23.0% had a student group. In bivariate analyses, MD programs and institutions with CI fellowships were significantly more likely to offer courses (both P < 0.001). In multivariable analyses, MD program type was the strongest predictor (adjusted odds ratio [aOR]=6.58, 95% confidence interval [CI] 2.22-22.41), followed by CI fellowship presence (aOR = 3.36), private school status (aOR = 2.08), and class size (aOR = 1.01). All 51 student groups were at MD institutions, and urban setting was associated with group presence (P = 0.034). The association with CI fellowships suggests a relationship between graduate and undergraduate medical education institutions. The association with MD programs could be influenced by differing curricular demands and flexibility. The association with urban settings may reflect the role of local innovation ecosystems. CI educational opportunities vary across US medical schools, concentrated at MD programs and institutions with GME-level infrastructure. These findings establish a baseline for CI opportunities and highlight the need to understand whether institutional differences translate to measurable competency gaps.
Dr. Kevin B. Johnson delivered this address on May 16, 2026, at the Commencement Ceremony of The D. Bradley McWilliams School of Biomedical Informatics, UTHealth Houston, to the graduating class of 2026. The address uses the concept of "The Big Mo" (compounding momentum) as a frame for understanding the current inflection point in AI and medicine. Drawing on his own career arc from paper-based clinical practice at Johns Hopkins through early adoption of health informatics to the present era of AI in healthcare, Johnson argues that the fears graduates hold about technological obsolescence and institutional instability are real but misdirected. He reframes both: biomedical informatics professionals are not targets of AI but its essential architects, and the external environment has always been uncertain for those doing important work. His charge to graduates is singular: stay on the wave.
Clarify disciplinary foundations and internal structure of biomedical informatics. We analyze BMI's emergence at disciplinary intersections and map its internal structure across 4 domains: theory and practice of knowledge discovery, knowledge representation and reasoning, knowledge architecture, and knowledge-driven transformation. We compare BMI with mathematics, computer science, biostatistics, and biomedical engineering, and illustrate emergent characteristics through a precision medicine example. BMI's distinctive contribution-elucidating the structure of biomedical knowledge and developing methods to discover, preserve, and make knowledge actionable-requires strength across all 4 domains. BMI developed these domains pragmatically: building systems, extracting principles, and formalizing theories. The discipline must now complement empirical approaches with rigorous theoretical work: assessing adequacy of existing theories, identifying gaps, and orchestrating collaborative development. BMI creates emergent capabilities across disciplines. As biomedicine becomes increasingly complex, BMI must strengthen its theoretical foundations while demonstrating transformative potential of knowledge spanning biological scales and time.
Despite the significant potential for Clinical Decision Support Systems (CDSSs) to improve care processes and health outcomes, several barriers hinder their widespread implementation in healthcare. While numerous systematic reviews have summarized potential barriers and facilitators for CDSS implementation, a comprehensive framework to guide and evaluate the implementation of CDSSs in healthcare is lacking. This overview of reviews, aims to establish a framework-GUIDE-CDSS-aimed at guiding and evaluating implementation of CDSSs in healthcare. An overview of systematic and scoping reviews was conducted by searching 6 databases. Systematic reviews or scoping reviews that used qualitative research methods to described implementation determinants for CDSSs in the healthcare domain were included. The AMSTAR 2 tool was used to assess the methodological quality. Results were collated into the GUIDE-CDSS framework. This framework describes implementation determinants and elements within those determinants found to impact implementation of CDSSs in healthcare. Twenty-three reviews were included in the analysis. All reviews had at least 2 critical weaknesses, showing a limited methodological quality of included reviews. Eight determinants and 38 elements for implementation of CDSSs in healthcare and were described in the GUIDE-CDSS framework: perceived relevance, perceived effect, trustworthiness, ease of use, workflow, training and skills, resources, and implementation strategy. This overview provides a comprehensive synthesis of the determinants influencing the implementation of CDSSs in healthcare, collated in the GUIDE-CDSS framework. The findings underscore that for successful CDSS development, implementation and evaluation is multifactorial. This study was registered in PROSPERO (No. CRD42024512455).
The potential of "big data" in health research remains largely untapped, particularly concerning real-world data sources such as administrative health data and electronic medical records. While healthcare insurance claims data have been essential for assessing medication safety and effectiveness within indicated patient populations, exploring broader drug-outcome associations could uncover significant insights. The REWARD (REal-World Evidence and Research of Drug performance) framework was established to perform large-scale analytics on medication benefits beyond their original indications. Employing an "all-by-all" approach, it investigates all medication-outcome pairs using standardized vocabularies from the OMOP common data model and implements causal inference methods, including self-controlled cohort and active comparator new-user designs. Negative control calibration and large-scale propensity scores are used to control for systematic bias in the study designs. Our framework enables the identification of new benefits associated with thousands of medications across millions of patients. In particular, REWARD facilitates insights into disease mechanisms to guide the development of novel interventions, identifies opportunities for drug repurposing, and informs potential additional indications for drugs in clinical development. REWARD has led to various publications with the goal of informing new drug development. REWARD is distinguished by its open-source implementation, use of standardized OMOP CDM vocabularies, and integration of best-practice pharmacoepidemiologic methods. Results are best interpreted as hypothesis-generating signals, with the active comparator new-user design providing higher causal rigor for prioritized drug-outcome pairs. The REWARD framework demonstrates how real-world evidence can be harnessed to address unmet medical needs, particularly for diseases lacking effective approved treatments. By making the REWARD analytic package open-source and accessible, we promote an open scientific approach, while maintaining best practices in pharmacoepidemiology.
The 2025 American College of Medical Informatics (ACMI) Symposium, themed Charting the Future of ACMI, provided a forum for fellows to reflect on the 2020-2025 ACMI Strategic Plan and define future priorities. Discussions centered on: (1) mobilizing ACMI's expertise to guide the responsible use of artificial intelligence (AI) in medical informatics, (2) advancing informatics education and mentorship across career stages, and (3) strengthening local, national, and global interdisciplinary partnerships. Proposed actions include establishing an AI Task Force, developing an online Education Resource Library, expanding structured mentorship, and cultivating strategic partnerships with relevant informatics, engineering, and clinical organizations. Fellows also identified cross-cutting challenges, including financial pressures, competition with alternative convening spaces, and the need to engage emerging leaders. By addressing these challenges while advancing its strategic imperatives, ACMI aims to serve as a trusted compass for the informatics community.
This study compares the contents of two data standards; the Observational Medical Outcomes Partnership (OMOP) and Fast Healthcare Interoperability Resources (FHIR), highlighting their strength and weaknesses and serve as an initial step toward understanding how each standard supports secondary data analysis. Participant electronic health record data in both OMOP and FHIR formats from the All of Us Research Program (AoURP) were compared, including codeable event volume, healthcare encounters, and person timelines. A phenotype-based assessment was also conducted using Type-II Diabetes Mellitus (T2DM). Among 29 512 participants identified with overlapping FHIR and OMOP data, Median codeable event counts were comparable between FHIR and OMOP within the Measurement (OMOP = 846; FHIR = 832), Drug (OMOP = 92; FHIR = 90), and Observation (OMOP = 65; FHIR = 100) domains, but were higher in OMOP within the Condition (OMOP = 258; FHIR = 11) and Procedure (OMOP = 72; FHIR = 4) domains. Within the T2DM cohort, OMOP contained more data, except for medications. Very few participants had encounters in FHIR (1.5%) relative to OMOP (97.0%). On average, only 15.9% of visit dates overlapped in both standards, with most visit dates occurring only in OMOP (65.3%) or only FHIR (18.4%). Data in OMOP showed a higher volume of observation and procedure codeable events, reported encounters, and T2DM symptoms, complications, and comorbidities. FHIR data was able to capture data across multiple providers and health systems. Based on AoURP data, both standards were shown to support healthcare data capture, although OMOP shows greater utility for research purposes as its extract, transform, and load process enables more flexible data capture relative to extracting data from FHIR payloads. FHIR, however, captures patient-level data from beyond the health system and can thus be used to supplement OMOP.
To enhance the accuracy, interpretability, and robustness of large language models (LLMs) in medical question answering (MedQA). We designed a multi-agent peer-reviewed reasoning method in which multiple LLM agents independently generate chain-of-thought (CoT) reasoning with candidate answers, then act as peer reviewers to evaluate each other's reasoning for factual correctness and logical soundness. The highest-rated reasoning chain is selected to produce the final answer. Experiments were conducted with 5 state-of-the-art LLMs (Llama-3.1-8B, Qwen2.5-7B, Phi-4, DeepSeek-LLM-7B, and GPT-oss-20B) on 3 benchmark datasets: HeadQA, MedQA-USMLE, and PubMedQA. Performance was compared against single-model CoT reasoning and CoT-based majority voting. Peer-reviewed reasoning consistently outperformed both baselines. The best model combination achieved an average accuracy of 0.820 across datasets, exceeding the strongest single model (0.777) and majority voting ensembles (up to 0.789). The method also scaled effectively with more participating models, while peer assessments reliably distinguished high- from low-quality reasoning chains. The proposed multi-agent peer-reviewed reasoning method enables LLMs to act as both solvers and evaluators, yielding superior performance in MedQA. By emphasizing reasoning quality rather than answer agreement alone, this approach improves accuracy, interpretability, and robustness, offering a promising direction for trustworthy biomedical AI systems.
Data modernization (DM) seeks to transform data and information systems in public health (PH) to support core activities. This study characterizes incumbent PH workers who perform informatics and data-centric roles that are critical to DM activities. Responses from the 2024 Public Health Workforce Interests and Needs Survey (PH WINS) on the US governmental PH workforce were analyzed. Out of 77 PH job classifications, 9 were mapped to informatics/data-centric roles. Demographics, work characteristics, and contribution to informatics and other PH program areas of these roles were examined. Based on weighted responses, informatics and data-centric roles comprise approximately 8% (N = 4786) of the PH workforce. This included PH informatics specialists (<1%), epidemiologists (3.9%), data analytics/related roles (1.7%), and information technology/computer science (IT/CS) workers (2.1%). Informatics roles are performed by PH informatics specialists as well as other workers, like epidemiologists, and PH informatics specialists support a broad range of PH activities including surveillance. Achieving the goals of DM and the national Public Health Data Strategy will require developing and supporting a variety of informatics and data-centric roles in PH agencies.
The analysis of care trajectories derived from electronic health records and claims data has become increasingly common in biomedical informatics. This has enabled large-scale studies of care processes, yet widely used binary code representations result in high-dimensional, sparse data that fail to capture semantic relationships between medical concepts. Learning dense vector representations (embeddings) has emerged as a promising approach to address these limitations. We aimed to construct and share joint embeddings for the International Classification of Diseases (ICD-10) and the Anatomical Therapeutic Chemical (ATC) classification system, providing reusable semantic representations of diagnoses and treatments from real-world claims data. Using claims records from 1.5 million patients, we defined code co-occurrences within temporal windows and constructed a Positive Pointwise Mutual Information (PPMI) matrix spanning ICD-10 and ATC codes. Singular Value Decomposition (SVD) was applied to derive a low-dimensional embedding space. Evaluation combined UMAP visualization, nearest-neighbor retrieval, and a code-level classification task based on ICD chapters and ATC classes. The embeddings reflected the hierarchical organization of ICD-10 and ATC and revealed associations across coding systems, including clinically relevant diagnosis-treatment relationships. The classification task achieved mean AUCs of 0.93 for ICD-10 and 0.90 for ATC, indicating strong grouping of semantically related codes. The embeddings provide a reusable, code-level semantic representation that can support code retrieval, reduce manual code grouping, and be aggregated into patient-level features without training a task-specific model. We release the first openly available joint ICD-10-ATC embedding space derived from real-world claims data, providing a reusable resource for biomedical informatics research.
Long COVID (LC) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients. In this study, we use summarized symptom reports by RECOVER-Adult cohort participants linked to EHR data to characterize patients and train a computable phenotype algorithm of LC. The study included adult participants with linked FHIR-sourced EHR data. We characterized EHR diagnoses, procedures, medications, lab tests, and vital sign features associated with LC. A computable phenotyping algorithm was trained and validated against patient-reported symptoms. We assessed model discrimination and calibration in a held-out test set. We describe important model features and evaluate model discrimination and calibration. The study included 1,501 RECOVER-Adult cohort participants with linked EHR data. 376 (25%) met criteria for highly symptomatic LC based on the RECOVER Long COVID Research Index (LCRI). EHR features associated with LC included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate. The algorithm identifying patients with highly symptomatic LC had an AUROC of 0.80 (95% confidence interval (CI) 0.74-0.85), and AUPRC of 0.58 (95% CI, 0.47-0.69). These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported LC symptoms. The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.
The successful integration of Machine Learning (ML) models into clinical practice remains limited, as they often lack the standardized, quantifiable risk measures essential for clinical workflows. This study, therefore, aims to demonstrate the Unified Auto Clinical Scores (Uni-ACS) method as a means to translate ML predictions into interpretable biostatistical formats, a translation critical for clinical adoption and alignment with evidence-based practice guidelines. We employed the Uni-ACS post hoc methodology to convert ML model outputs into clinical scores and odds ratios. Validation was performed on a retrospective cohort of Chronic Obstructive Pulmonary Disease patients admitted between 2016 and 2018, predicting adverse outcomes from 2019 to 2020. The performance of Uni-ACS was benchmarked against the original ML models and logistic regression. The successful application of Uni-ACS enabled the conversion of complex ML model outputs into readily interpretable clinical scores and metrics, like odds ratios. Clinical scores derived from the ML models' SHAP values using Uni-ACS maintained a strong, interpretable predictive performance (AUROC 0.69-0.80), which was comparable to the original ML models (AUROC 0.71-0.83) and logistic regression (AUROC 0.72-0.81). Uni-ACS addresses the fundamental clinical requirement for standardized, meaningful risk stratification. It efficiently translates complex ML predictions into validated clinical scores with minimal computational burden. In conclusion, this approach facilitates the widespread adoption of ML-driven risk assessment across diverse healthcare settings while ensuring compatibility with existing clinical guidelines and regulatory frameworks.
This address was delivered by Eric Horvitz, MD, PhD, at the 2026 graduation ceremony of Columbia University School of Nursing on May 19, 2026, where he received the Second Century Award for Excellence in Health Care. The address considers the responsibilities of clinicians in shaping the future of artificial intelligence in medicine. It frames health care as an "open world," where information is incomplete, time is limited, and decisions are made under uncertainty. As AI transforms biomedicine and clinical care, the address emphasizes the importance of clinician engagement in guiding how these technologies are developed and used, and calls for systems that strengthen clinical judgment, support care teams, and advance human health, dignity, connection, and trust.
To evaluate the clinical applications and translation readiness of model-based synthetic tabular data in healthcare, and identify gaps in governance reporting that may hinder translation. We systematically searched Ovid MEDLINE and Embase (2010-August 2025; PROSPERO: CRD42025635514) for studies that generated and applied model-based synthetic tabular data in clinical contexts. Screening used a "human-in-the-loop" large language model workflow alongside independent manual review, achieving 100% sensitivity for included studies. Unlike prior reviews focused primarily on evaluation methodology, we mapped use-cases and deployment paradigms, and audited translation-readiness reporting using a predefined governance framework (validation depth, privacy, fairness, regulatory alignment). Thirty-seven studies (2019-2025) were included. GANs predominated; other approaches included VAEs, diffusion models, LLM-based synthesis, and Bayesian networks. Dataset augmentation was the primary application, often improving downstream model performance for rare outcomes. Emerging applications included synthetic control cohorts and algorithmic bias mitigation. Translation-readiness reporting was limited: 34/37 studies (92%) relied solely on internal validation, 9/37 (24%) used formal privacy models, 6/37 (16%) reported explicit fairness evaluations, and 6/37 (16%) addressed regulatory alignment. Few studies distinguished "no-release" from "delayed-release" paradigms. A systemic gap exists between methodological innovation and deployment-readiness reporting. Model-based synthetic data show clear value for augmentation and class balancing, but inconsistent reporting of validation, privacy, fairness, and regulatory considerations limits confidence in clinical deployment. We propose TRUST-SD (Transparency and Reporting for Utility, Safety, and Translation of Synthetic Data), an author-derived, preliminary, evidence-informed reporting checklist spanning 7 domains, as a starting point for community refinement and consensus-building.
Using AI algorithms can exacerbate health disparities if care or resources are allocated away from underserved populations. We evaluated an algorithm for its potential to worsen health disparities across different clinical use cases. This was a retrospective study of patients with heart failure (HF) at an academic health system using an algorithm that predicts pharmacy fill nonadherence to evidence-based HF medications. We compared prediction performance metrics (accuracy, false positive rate, false negative rate), using rate-ratios (RRs), between subgroups with and without known HF care disparities: below vs above median neighborhood-level socioeconomic status (nSES) and Black vs White race. Results were then applied to 3 hypothetical clinical use cases. Among 34 697 patients (13% Black, 10% Hispanic, 65% White), algorithm accuracy was similar across nSES and racial subgroups. The algorithm assigned more false positives for medication nonadherence among low vs high nSES (RR [95%CI] 1.50 [1.44-1.56]) and Black vs White (2.05 [1.92-2.19]) subgroups. The algorithm also assigned fewer false negatives (0.63 [0.59-0.67]) to Black vs White subgroups. When applied to 3 hypothetical use cases, worsening of existing disparities was pertinent for clinical applications where false positives could be particularly harmful (e.g, if predictions of nonadherence prompted lower treatment priority). Although accuracy was similar across demographic groups, differences in false positive and false negative rates revealed that the same prediction may worsen disparities in some use cases, but not others. Evaluation of predictions in the context of clinical use is essential to avoid unintentionally worsening inequities.
Stigma impacts outcomes across stigmatizing conditions, including substance use disorders (SUDs). Recent policy changes give patients rapid access to clinical notes in the electronic health record (EHR), which may include stigmatizing language. The objective of this study was to assess the perspectives of women with history of pregnancy and SUDs on typical language used in clinical notes. Women with a history of pregnancy and SUD were recruited through an online crowd-sourcing platform. Respondents viewed examples of clinical language and answered survey questions about perceived stigma. An inductive approach was used to analyze open-text responses, and themes were developed. Three hundred seventy survey respondents wrote a response to at least one open-text question. Thematic analysis yielded 4 major themes: (1) anticipation of future stigma facilitated by EHR documentation can affect patients' care decisions for themselves and their babies; (2) documented SUD history could have short- and long-term effects on patients' experience of stigma and discrimination, especially in labor and delivery; (3) phrases using "denies" and quotes within quotation marks could be perceived as stigmatizing and decrease trust in providers; (4) nonstigmatizing language and acknowledgement of recovery in notes can facilitate positive experiences for patients, but patients want more acknowledgement of recovery and positive language. Electronic health record documentation can modulate stigma experiences for women during and after pregnancy through stigmatizing language in clinical notes and facilitating discrimination, decreasing trust in providers and negatively impacting health outcomes. Raising providers' awareness of nonstigmatizing and positive language or implementing technology to prompt nonstigmatizing terminology could contribute to positive experiences among women with a history of pregnancy and SUD.
This study aimed to identify trajectories of multimorbidity following acute myocardial infarction (AMI), using explainable temporal machine-learning methods, and assess their clinical, prognostic, and biological significance. Dynamic Time Warping k-means clustering was applied to post-AMI diagnostic sequences from 12 701 UK-Biobank participants. Latent Dirichlet Allocation characterised cluster themes. Multiclass classifiers (CatBoost, XGBoost, random forest, logistic regression) trained on pre-AMI diagnoses and demographics predicted trajectory membership, with SHAP interpretability. SMART scores and Cox models evaluated 5-year mortality; Phenotype-Wide Association (PheWAS) and Reactome pathway enrichment were used to identify associated biological mechanisms. Three trajectories of multimorbidity were identified: acute cardiorenal-respiratory with metabolic disease (ACUTE-CARD; 63.4%), cardiometabolic disease with arrhythmic-ischemic burden (CARDIOMIX; 13.5%), and smoking-related multisystem multimorbidity (SMO-CARD; 23.1%). XGBoost achieved the highest discrimination (AUC-ROC 0.906; 95% CI, 0.895-0.916), with CatBoost showing comparable performance (AUC-ROC 0.900; 95% CI, 0.889-0.910). SMO-CARD had the highest 5-year mortality (43.9%). The established SMART cardiovascular risk score remained the dominant predictor of mortality while the identified trajectories provided complementary prognostic signals; these associations were weaker after full adjustment for potential confounders with each profile displaying distinct genetic and pathway signatures. The SMART score outperformed in capturing mortality risk, whereas the trajectories complement it by revealing nuanced clustering of risk factors across organ systems and in identifying trajectory-specific intervention priorities. Explainable temporal modeling of EHR data reveals clinically interpretable, biologically grounded multimorbidity trajectories after AMI that complement established risk scores and provide a reproducible approach to mechanistic phenotyping and precision care.