Enhancing the effectiveness of human-AI collaborative consultation in online health communities (OHCs) constitutes a core requirement for optimizing the allocation of medical resources and promoting the sustainable development of medical services. Nevertheless, the pathways to improving such effectiveness remain insufficiently understood. This study aimed to conduct an in-depth exploration of the multiple factors influencing the effectiveness of human-AI collaborative consultation in OHCs and to assess the causal relationships among these factors, thereby providing a theoretical foundation and practical guidance for advancing the clinical application of human-AI collaboration. Grounded in the information ecology theory, we constructed an analytical framework encompassing four dimensions: information human, information, environment, and technology. We collected 296 valid questionnaire responses and used fuzzy set qualitative comparative analysis to systematically investigate the configurational mechanisms through which these four types of factors jointly influence consultation effectiveness. The findings are as follows: (1) no single factor constitutes a sufficient condition for high consultation effectiveness (consistency <0.9); rather, such effectiveness emerges from the synergistic interplay of multiple configurations involving technology, information, human information, and environment; (2) five distinct configurational pathways lead to high consultation effectiveness, demonstrating clear equifinality; (3) system responsiveness, information usefulness, perceived service empathy, perceived service accuracy, and perceived service effectiveness are all core or important conditions across these pathways; and (4) substitutability exists among antecedent conditions-specifically, perceived uncertainty and social norms, as well as operational convenience and platform ethical norms, can substitute for one another in different configurations to enhance patient satisfaction jointly. This study reveals multiple pathways to achieving high effectiveness in human-AI collaborative consultation within OHCs. It not only offers a novel theoretical perspective for understanding the complex mechanisms of human-AI collaboration in medical contexts, but also provides significant practical implications for the design optimization of digital health platforms and related policy formulation.
Sleep disorders represent a significant public health burden associated with cardiovascular and neurocognitive morbidities. While AI technologies offer potential for personalized sleep medicine, clinical integration remains limited. This translational disparity is often attributed to a lack of human-centered design, specifically insufficient stakeholder engagement in the development and implementation of these technologies. Current research frequently prioritizes algorithmic performance over usability and patient trust. This scoping review systematically maps the extent and nature of human-centered AI (HCAI) research within sleep medicine across different AI modalities, evaluating how diverse stakeholders are involved in the design, validation, and implementation of AI tools, including patients, clinicians, and technologists. Following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines, we searched 8 databases (PubMed, Web of Science, Embase, Scopus, IEEE Xplore, ACM Digital Library, APA PsycINFO, and CINAHL) for literature published up to June 18, 2026. We identified primary research describing the design, development, or evaluation of AI technologies for sleep health with explicit human-centered components. Included studies (n=34) were categorized based on AI technology type and the method of stakeholder engagement. Data were extracted and synthesized using a thematic analysis approach. Based on the included studies, the analysis reveals an uneven distribution of research focus across technological domains as descriptive patterns rather than definitive trends. Research on generative AI (GenAI) is predominantly restricted to downstream expert auditing of output accuracy (comprising 7/11, 64% of GenAI studies), with a noticeable gap in upstream participatory design involving patients. Conversely, deep learning research primarily focuses on technical explainable AI methods to address algorithmic opacity for clinicians, yet lacks progression to real-world clinical implementation. Mobile health and wearable technologies (17/34, 50%) demonstrate the most balanced HCAI ecosystem, evidencing a complete translational cycle from upstream co-design to downstream clinical implementation. Furthermore, an emerging trend is observed where AI is evolving from an automated diagnostic tool into an interactive therapeutic agent, with recent studies indicating that lay users may perceive responses from large language models as more empathetic than those from physicians. Lacking formal quality appraisal, our findings reflect research activity patterns rather than confirmed clinical effectiveness. Nevertheless, this scoping review innovatively applies the HCAI framework to the sleep AI lifecycle. Unlike existing reviews prioritizing algorithmic performance metrics over usability, clinical workflow integration, and patient trust, this study systematically maps these essential sociotechnical factors. It contributes to the field by revealing distinct methodological disparities and the urgent need for upstream participatory design, particularly for GenAI. In the real world, establishing standardized protocols for human-AI interaction, ensuring algorithmic transparency, and addressing demographic biases are essential to foster the clinical trust required for effective AI adoption.
Community college (CC) students face significant mental health concerns but are unlikely to receive treatment. Barriers to mental health service uptake among CC students have been delineated, but few studies have identified strategies to improve uptake. Text messaging has been used to address engagement barriers to mental health services among adolescents and adults, but little research has explored this strategy for CC students. The goal of this study was to partner with CC students to co-design and conduct pilot usability testing of a text messaging intervention to address barriers and increase uptake of a mental health screening and treatment program, called Screening and Treatment for Anxiety and Depression (STAND), offered to CC students. We conducted 2 parallel sets of 4 co-design focus groups with CC students who had varying levels of engagement with STAND. We used rapid qualitative analysis to extract key themes, create text message prototypes and refine them, and present updated prototypes to gather feedback across workshops. We also assessed six usability factors on a 5-point Likert scale: satisfaction, helpfulness, attractiveness, readability, comprehension, and likelihood of getting started with STAND after receiving texts. Key themes emerged about perceptions of texting, barriers to STAND, a basic framework for the text message intervention, feedback about the format of messages, and feedback about the content of messages. Students expressed positive regard for text messaging and general agreement on key barriers to STAND. Students codeveloped a framework for the intervention, including (1) delivering introductory texts to engage students in the text messages, (2) providing a personalized approach for students to select barriers most salient for them, and (3) delivering tailored content designed by students to address each barrier. Across workshops, several themes emerged with regard to how messages should be formatted and delivered, including the following: use short messages; use not too many messages; use relevant language; use images, memes, and short videos; and make messages "human-like." Themes related to the content of messages included the following: reminders that you are not alone, knowledge that STAND has worked for other students, expressing understanding of student context and stressors, and providing an option to speak to a team member. Mean ratings on usability factors ranged from 3.88 (SD 0.64) to 4.25 (SD 0.46). This study describes a process for co-designing a text messaging mental health engagement intervention with CC students that is grounded in a human-centered design approach. Further research is needed to rigorously test this intervention and make iterative refinements to improve response and effectiveness.
AI-powered clinical decision support systems (CDSS) have shown promise in improving prediction, monitoring, and treatment optimization across clinical domains, including HIV care. However, translating AI outputs derived from electronic health records into clinically meaningful, trustworthy, and actionable decision support remains challenging, underscoring the need for more human-centered and socioecologically grounded CDSS design. This study aimed to explore how we can effectively translate the outputs of machine learning models based on HIV electronic health records into a real AI-powered CDSS for HIV care. Using the human-in-the-loop method, we engaged a set of stakeholders, including HIV physicians, nurse practitioners, infectious disease pharmacists, social workers, and case managers. Stakeholders interacted with an AI-powered CDSS prototype to identify barriers and challenges to adoption, as well as to inform a more holistic and context-aware AI-powered CDSS design. We conducted a field study at Prisma Health in South Carolina that included pre- and postsurveys, interactive usability testing sessions, think-alouds, and in-depth interviews with 16 clinicians providing HIV care between March and September 2025. We analyzed survey responses using descriptive statistics, and then transcribed and analyzed think-aloud and interview data using an etic and emic approach. Clinicians identified multiple challenges and design considerations for AI-powered HIV CDSS, demonstrating that clinician-AI interaction is inherently sociotechnical and embedded across multiple socioecological levels. While clinicians relied on familiar clinical indicators as cognitive anchors for interpreting AI predictions, they emphasized that social determinants of health were central to their own risk assessment and clinical decision-making. Additionally, clinicians' trust in AI is conditional and develops over time, with explainability and actionability emerging as critical factors for translating predictions into meaningful clinical interventions. Findings highlight the need to move beyond technically accurate predictions toward AI-powered CDSS designs that align with clinicians' cognitive practices and socioecological realities of HIV care. By extending a sociocognitive framework through empirical grounding in HIV clinical practice, this study offers design insights for developing AI-powered CDSS that are trustworthy, context-aware, and capable of supporting actionable decision-making in HIV care settings and beyond.
Artificial intelligence (AI) tools have the potential to enhance personalized clinical care, particularly in radiology. However, their integration into clinical workflows remains complex, especially in pediatric oncology, where early cancer detection is critical. Children with Li-Fraumeni Syndrome (LFS), a rare cancer predisposition disorder, undergo regular surveillance whole-body MRI (wbMRI), which presents an opportunity for AI-assisted tumour detection. We evaluated the feasibility of an AI-assisted overlay for highlighting tumour-like regions in pediatric surveillance wbMRI and explored how access to the overlay influenced radiologist workflow, candidate-lesion marking behaviour, follow-up recommendations, and perceived workload. We developed a patch-based AI segmentation model trained on augmented 2D slices from 675 surveillance wbMRI volumes of pediatric patients with LFS. The model was designed to highlight regions with high tumour probability. A reader study was conducted with two radiologists who independently reviewed wbMRI cases both with and without AI assistance. We measured evaluation time, number and location of tumours identified, type of follow-up recommendation, and subjective feedback using structured questionnaires. AI assistance altered interpretation workflows for both radiologists, with mixed effects. On average, the time required to evaluate each case increased when using the AI tool for both radiologists. However, one radiologist had an increase in the number of candidate lesion locations selected with the tool, and one had a decrease in the number of candidate lesion locations selected with the tool. Subjective feedback indicated that one of the radiologists found a greater difference in their perception between performing with the tool versus without the AI tool; however, both radiologists felt that the task was less difficult and less stressful with the AI tool. Inter-rater variability was evident, underscoring the need for personalized calibration of AI tools. AI-assisted wbMRI interpretation can improve tumour detection in pediatric cancer surveillance by reducing false negatives. However, its influence on workflow efficiency and inter-radiologist variability highlights the importance of careful implementation. Successful integration requires addressing challenges such as improving the predictive precision of AI models, offering intuitive end-user designs and instructions, and building trust in AI outputs. AI outputs can influence workflow and behaviour in reader-specific ways. Clinical translation will require larger, randomized, multi-reader studies and model refinement to reduce false positives and quantify lesion-level reader performance. This can help ensure better patient outcomes in addition to reduced clinician burnout.
AI tools have the potential to enhance personalized clinical care, particularly in radiology. However, their integration into clinical workflows remains complex, especially in pediatric oncology, where early cancer detection is critical. Children with Li-Fraumeni syndrome (LFS), a rare cancer predisposition disorder, undergo regular surveillance whole-body magnetic resonance imaging (wbMRI), which presents an opportunity for AI-assisted tumor detection. We evaluated the feasibility of an AI-assisted overlay for highlighting tumor-like regions in pediatric surveillance wbMRI and explored how access to the overlay influenced radiologist workflow, candidate-lesion marking behavior, follow-up recommendations, and perceived workload. We developed a patch-based AI segmentation model trained on augmented 2D slices from 675 surveillance wbMRI volumes of pediatric patients with LFS. The model was designed to highlight regions with high tumor probability. A reader study was conducted with 2 radiologists who independently reviewed wbMRI cases both with and without AI assistance. We measured evaluation time, number and location of reader-marked candidate lesions, type of follow-up recommendation, and subjective feedback using structured questionnaires. AI assistance altered interpretation workflows for both radiologists, with mixed effects. On average, the time required to evaluate each case increased when using the AI tool for both radiologists. However, one radiologist had an increase in the number of candidate lesion locations selected with the tool, and one had a decrease in the number of candidate lesion locations selected with the tool. Subjective feedback indicated that one of the radiologists reported lower mental demand with the AI tool, while both radiologists reported lower stress with the AI tool. Interrater variability was evident, underscoring the need for personalized calibration of AI tools. AI-assisted wbMRI interpretation can improve tumor detection in pediatric cancer surveillance by reducing false negatives. However, its influence on workflow efficiency and interradiologist variability highlights the importance of careful implementation. Successful integration requires addressing challenges such as improving the predictive precision of AI models, offering intuitive end-user designs and instructions, and building trust in AI outputs. AI outputs can influence workflow and behavior in reader-specific ways. Clinical translation will require larger, randomized, multireader studies and model refinement to reduce false positives and quantify lesion-level reader performance. This can help ensure better patient outcomes in addition to reduced clinician burnout.
Blended therapy (BT) combines digital applications with face-to-face treatment and has become an increasingly important component of psychiatric care. Evidence indicates that BT can achieve outcomes comparable to or even superior to those of traditional face-to-face therapy. Despite certain advantages, routine implementation of BT remains challenging, and clinical practice suggests that while some inpatients engage with BT, many either discontinue early or do not initiate its use at all. To better understand these patterns, this multicentric, retrospective observational study investigates factors associated with noninitiation and dropout among inpatients who are offered BT.  In this study, data from 278 inpatients were analyzed to examine the influence of sociodemographic variables, comorbidities, and symptom severity on the uptake and continued use of BT. The objective was to identify predictors of noninitiation and dropout.  Multivariable logistic regression models were conducted to identify significant predictors of noninitiation and dropout among inpatients using the transdiagnostic, cognitive behavioral therapy-based electronic mental health platform Minddistrict, which offers modules targeting psychoeducation, cognitive restructuring, and behavioral activation. Data were collected from 2 psychiatric hospitals between January 2020 and May 2024. The sample consisted predominantly of patients diagnosed with depression (182/278, 65.7%) and posttraumatic stress disorder (61/278, 21.9%), alongside various comorbid conditions. The findings indicate distinct patterns of association for noninitiation and dropout. Of the 278 patients, only 5 (1.8%) completed all the assigned modules, and one-third of the patients never initiated the platform at all. Specifically, increasing age was linked to a lower risk of noninitiation (odds ratio [per year age difference] 0.98, 95% CI 0.96-1.00; P=.01), while the presence of a comorbid anxiety disorder was associated with a reduced risk of dropout (odds ratio 0.23, 95% CI 0.08-0.66; P=.007). Several variables showed no association with either noninitiation or dropout across all analyses, including sex, overall symptom severity, and certain comorbidities such as personality disorders and depression.  In this preselected inpatient sample, uptake of BT was very limited. Older age was associated with lower noninitiation, and comorbid anxiety disorders were associated with a lower likelihood of dropout. These findings may help inform future prospective studies on how BT can be introduced and supported more effectively in inpatient psychiatric care. As access to BT was granted selectively by therapists, the results should be interpreted as predictors of engagement within a selected sample rather than general predictors of BT uptake among all psychiatric inpatients.
Digital multidomain interventions hold promise for dementia risk reduction; however, populations at higher dementia risk, including those experiencing socioeconomic and educational disadvantage, remain underrepresented in trials, and engagement with digital interventions often declines over time. Coproduction and blended models that combine digital tools with human support may improve reach, acceptability, usability, and sustained engagement. Designing interventions that are usable and acceptable for individuals facing structural, educational, or digital barriers (underserved groups) is therefore likely to produce solutions that are both accessible and scalable for the wider midlife and older adult population. This study aims to describe the coproduction process used to develop ENHANCE (Tailored Intervention for Brain Health and Cognitive Enrichment)-a coach-supported digital intervention targeting 10 modifiable dementia risk factors in older adults from underserved groups-and report key outputs and lessons learned for equitable digital prevention design. We coproduced ENHANCE between July 2023 and February 2025 using a multistage development process guided by the Medical Research Council framework for complex interventions and the Double Diamond design model. The person-based approach informed user-centered guiding principles (key design objectives), while behavior change content was operationalized using behavioral change theories. Coproduction followed 4 phases. The Discovery phase explored barriers to engagement with existing digital materials and identified candidate components for each dementia risk-factor module. The Define phase translated these insights into guiding principles and blueprints of each risk-factor module integrated with behavioral change components. The Design phase involved iterative co-production and usability testing of prototypes. The Delivery phase evaluated a high-fidelity prototype through a 1-week usability study with coaching support. Contributors included 162 research participants recruited from underserved community settings, 33 patient and public involvement contributors, and 4 human-computer interaction experts. Throughout development, coproduction focused on reducing literacy, digital confidence, and cultural barriers to maximize usability across diverse adult populations. Coproduction produced (1) evidence-informed module strategies for targeted dementia risk factors; (2) a set of guiding principles to ensure low-literacy, culturally relevant, and accessible content, supporting both equity of access and wider population usability; (3) a meadow-themed app integrating tailored check-ins, educational videos, cognitive training games, and in-app messaging; and (4) a structured coaching model, including onboarding, brief follow-up, and accompanying coaching manuals. Iterative testing and refinement improved navigation, simplified language, reduced text burden, and ensured the use of familiar and accessible game formats, resulting in a feasibility-ready prototype. ENHANCE is a coproduced, coach-supported digital intervention designed to be accessible for underserved midlife and older adults at increased dementia risk, with design features to support accessibility, engagement, and scalability across the wider aging population. The development process illustrates how integrating coproduction with behavioral science and usability methods can support principled intervention design for equitable digital dementia prevention.
AI adoption in health care has accelerated rapidly, with ambient documentation tools, diagnostic imaging AI, and clinical decision support systems (CDSSs) entering routine practice. However, the cognitive demands placed on clinicians supervising these systems remain understudied. Specifically, the concept of verification burden requires closer examination. Consequently, institutional decision-makers lack a structured, certainty-graded evidence base regarding the true impact of AI on clinician workload and burnout. This study aimed to systematically review evidence on cognitive workload and burnout in health care professionals that use AI-powered clinical tools, quantify pooled effects under a conservative inferential framework, and assess certainty of evidence by AI category. The study was registered in PROSPERO (CRD420261284298) and reported per PRISMA 2020 and PRISMA-S guidelines. We searched MEDLINE, Embase, Web of Science, and Cochrane CENTRAL (January 2015-2026) for studies measuring cognitive workload or burnout using validated instruments (NASA Task Load Index [NASA-TLX] and Professional Fulfillment Index [PFI]) among health care professionals using clinical AI. Risk of bias was assessed using ROB 2.0 and ROBINS-I; certainty was rated using GRADE. Meta-analyses applied Hartung-Knapp-Sidik-Jonkman adjustment with restricted maximum likelihood estimation, incorporating prediction intervals (PIs). We included 21 studies representing 2885 health care professionals across 7 countries. The synthesis demonstrated that the cognitive impact of clinical AI varies according to its specific application. Pooled analyses of ambient AI documentation showed statistically significant reductions in NASA-TLX temporal demand (SMD -1.46, 95% CI -2.81 to -0.11; k=2; I2=31.1%) and effort (SMD -1.29, 95% CI -2.16 to -0.42; k=2; I2=0%), PFI work exhaustion (MD -0.35, 95% CI -0.58 to -0.12; k=3; I2=0%; 95% PI -1.03 to 0.33), and burnout prevalence (OR 0.47, 95% CI 0.25-0.86; k=3; I2=0%; 95% PI 0.06-3.82). Two pools favored ambient AI but did not reach significance at k=2: NASA-TLX mental demand (SMD -1.29, 95% CI -3.64 to 1.07) and documentation time (SMD -0.24, 95% CI -1.10 to 0.61). Diagnostic imaging AI and CDSS showed mixed or paradoxically increased workload. GRADE certainty was moderate for cognitive workload reduction with ambient AI, low for burnout reduction with ambient AI, and very low for imaging AI and CDSS outcomes. This review combines validated workload instruments, meta-analysis, and PIs in health care AI, delivering a GRADE certainty assessment across 5 AI categories that prior accuracy- or efficiency-focused reviews have not provided. Ambient AI documentation was associated with reduced cognitive workload and burnout, but only in voluntary early-adopter cohorts and based on few studies; the conservative CIs were wide and, where estimable, PIs crossed the null. Findings inform institutional pilots with prospective workload measurement, regulatory human-factors evaluation of AI medical devices, and human-centered AI design. Net benefit on the health care workforce remains an open empirical question.
Preterm birth (PTB), or birth before 37 weeks of gestation, remains a significant public health issue in the United States, particularly in Detroit, Michigan. Growing evidence suggests that volatile organic compounds (VOCs), aromatic or chlorinated organic compounds that vaporize readily, may influence PTB risk. However, much of this prior work is limited by indirect VOC exposure estimates (eg, assignment based on maternal residential address), single-point or cumulative exposure estimates during pregnancy, or limited consideration of potential mechanistic factors. The Center for Leadership in Environmental Awareness and Research (CLEAR) birth cohort has been designed to test the hypotheses that prenatal VOC exposures, measured as VOC metabolites in maternal urine, increase the risk of PTB; that VOC exposures are associated with maternal inflammation and placental function measures; that associations between prenatal VOC exposures and PTB may be mediated by these maternal inflammation and placental function measures; and that there are neighborhood-level factors that may increase the risk of VOC exposure during pregnancy. A prospective cohort of 1075 pregnant patients receiving prenatal care at Henry Ford Health will be recruited. Pregnant patients residing in Detroit or receiving prenatal care at a Detroit-based Henry Ford Health women's health clinic are eligible. Pregnant patients are followed until delivery. Up to 3 urine and blood samples (collected during early, mid, and late pregnancy) are obtained for measurement of VOC metabolites and inflammatory biomarkers, respectively. The placenta is obtained after delivery for epigenomic and transcriptomic measurement. Surveys are administered to pregnant participants to assess a variety of lifestyle, psychosocial, medical, residential, and other factors. Address information collected from both surveys and electronic medical records across pregnancy will be used to identify potential sources of VOC exposure. The electronic medical record is used to obtain medical and delivery data, including infant sex, date of delivery, and gestational age (GA) at delivery. PTB, the primary study outcome, is defined as GA at delivery <37 weeks. A nested case-control approach (frequency matching PTB cases 1:1 with full-term controls [GA at delivery ≥37 weeks] based on infant sex and maternal race) will be applied. Statistical methods, including logistic regression, linear mixed methods, and geographically weighted regression models, as well as chemical mixture approaches, will be used. Funding began September 2022 and recruitment commenced November 2023. Through April 22, 2026, a total of 468 pregnant patients have consented to participate in the CLEAR birth cohort, and recruitment is ongoing. The CLEAR cohort will provide novel data on the role of VOCs during pregnancy in the risk of PTB. Additionally, the role of VOC exposures during pregnancy in maternal inflammation and placental function will be examined. Finally, potential sources of VOC exposures, which could be targets for environmental remediation, will be identified.
Clinician burnout has reached crisis levels in emergency medicine, with clinical documentation burden identified as a central contributing factor. Ambient artificial intelligence (AI) scribes offer a promising approach to reduce this burden, but objective evidence in the emergency department (ED) setting remains limited, and prior reports have been constrained by short observation windows and low adoption. This study aimed to evaluate the association between ambient AI scribe use and on-shift documentation time during a 13-month staged rollout in a busy ED, accounting for physician- and patient-level factors. We conducted a retrospective cohort study at a tertiary academic ED from February 2025 to March 2026. The analytic cohort comprised 10,344 encounters managed by 100 attending physicians across 4 ED care settings. We restricted analysis to encounters managed by a single attending physician and excluded those with human scribes. The comparison group comprised encounters in which the ambient AI scribe was not used; use was determined entirely at attending physician discretion on an encounter-by-encounter basis. The primary outcome was on-shift documentation time derived from electronic health record audit logs. We used mixed-effects linear models with physician random intercepts to adjust for patient and encounter characteristics. Ambient AI scribe use was associated with a 72.6-second reduction in on-shift documentation time per encounter (95% CI 63.8-81.4; P<.001). The effect was similar in magnitude for high-use physicians (use rates of ≥18.2%, which was the cohort mean; -71.6 seconds) and low or moderate users (-64.2 seconds), with no statistically significant difference (P=.51). Note character count decreased by 690 characters (95% CI 273-1107; P=.001); after-shift documentation time increased modestly by 9.1 seconds (95% CI 2.9-15.3; P=.004). Negative control outcomes were largely null, and a within-clinician placebo permutation test yielded a distribution centered at 0 (mean -0.8 seconds), inconsistent with the observed effect arising from confounding alone. In this single-center analysis, ambient AI scribe use was associated with a statistically significant reduction in on-shift documentation time (P<.001), equivalent to approximately 24 minutes per 8-hour shift if used across 20 encounters. These findings extend prior descriptive work with adjusted inferential evidence and support the clinical relevance of ambient AI scribes for ED documentation burden, although the magnitude of benefit varies by physician, patient, and workflow factors.
Primary care is becoming increasingly complex, with primary care physicians (PCPs) facing rising workloads driven by workforce shortages, growing administrative demands, and expanding clinical responsibilities. Recent advances in large language models (LLMs) offer new opportunities to support PCPs across clinical, administrative, and communication-related tasks within their workflows. Understanding how these technologies are perceived and used in primary care practice is, therefore, critical to inform their safe, effective, and human-centered implementation. This study aimed to explore Dutch and US PCPs' perceptions and experiences regarding the use of LLMs in clinical practice, with particular attention to clinical usability, communication and teamwork, and implications for everyday workflows. We conducted a qualitative study using semistructured interviews with 15 PCPs from the United States and the Netherlands. Data were collected between February and June 2025 and analyzed using reflexive inductive thematic analysis. Ten themes emerged related to the use of LLMs in primary care clinical practice, each theme consisting of a set of subthemes. We found that LLMs are being integrated into primary care as both clinical and communication support tools, assisting with diagnostic reasoning, administrative tasks, workload management, and interprofessional and patient communication. While PCPs reported perceived benefits, they also expressed concerns related to safety, efficiency, authenticity, and the preservation of the therapeutic relationship, highlighting the need for careful and context-sensitive use. Our findings suggest that LLMs are already being integrated into primary care in diverse ways, with their value shaped by both contextual factors and clinician judgment. Understanding how clinicians navigate LLM use in everyday practice is essential to ensuring that LLMs support high-quality, patient-centered primary care and inform organizational policy and LLM design.
Cardiovascular disease remains the leading global cause of mortality, driven by interrelated behavioral, biological, and psychosocial risk factors, despite the availability of effective prevention and treatment strategies. Persistent policy inertia, systemic fragmentation, and adverse social and commercial determinants have limited national responses. Addressing these gaps necessitates place-based, systems-oriented approaches that mobilize local assets, engage multisector stakeholders, and incorporate adaptive evaluation. The Springfield Healthy Hearts initiative exemplifies such an approach by positioning Greater Springfield as a "living laboratory" for coordinated cardiovascular health action through a comprehensive data framework, providing a replicable model for other communities. This protocol outlines the Springfield Healthy Hearts Data Framework, a multicomponent system for dynamically guiding, implementing, and evaluating coordinated action for heart health. The data framework was developed through a structured co-design process involving community members, expert researchers, health professionals, and representatives from local implementation partners. The framework comprises four integrated components: (1) Project evaluation, applying pragmatic frameworks to assess coordinated action projects; (2) Community evaluation, a repeated cross-sectional evaluation of Springfield residents, workers, and regular visitors to capture individual-level behavioral, biological, and psychosocial cardiovascular disease risk factors, as well as engagement with coordinated action projects; (3) City evaluation, ongoing monitoring of suburb- and city-level indicators across 4 domains (sociodemographic characteristics, built environment, food and commercial environment, and health services); and (4) Data synthesis, to utilize data across all levels to inform a continuous learning system. Project evaluations will use both quantitative and qualitative methods, including realist evaluation where appropriate. Community evaluation will be analyzed using descriptive statistics, mixed effects models, and subgroup analyses, with missing data addressed via multiple imputation. City-level data will be analyzed descriptively and dynamically to detect temporal trends and contextual changes. Initial funding for the Springfield Healthy Hearts Data Framework was secured in June 2025. Co-design workshops were conducted between November 2025 and February 2026 (n=15 participants), informing the design and prioritization of framework components. Community evaluation data collection is scheduled to commence in September 2026 and conclude in August 2027. Data cleaning and preliminary analyses are anticipated in late 2027, with first results expected to be disseminated in early 2028. The Springfield Healthy Hearts Data Framework is a replicable model for other communities aiming to implement city-wide, coordinated approaches to heart health action. Findings will be disseminated through peer-reviewed publications, community reports, interactive dashboards, and policy briefs.
Health services increasingly face decisions about how to integrate immersive technologies into routine practice. International guidance highlights the need for structured governance in digital health, yet extended reality (XR) initiatives are often launched through isolated pilots without a clear assessment of organizational readiness or implementation risk. Although factors influencing XR adoption are well documented, health care organizations and system-level decision-makers still lack practical, governance-oriented tools to translate these determinants into structured strategic decisions made before implementation. This study aims to develop multicriteria decision analysis for extended reality (MCDA-XR), a strategic governance framework that translates behavioral, organizational, and technical implementation determinants into a structured decision-support process for health care organizations. The study followed a sequential mixed methods design covering the first 2 phases of a 3-stage framework development and validation project. Phase 1 (identification) defined strategic criteria by integrating theoretical perspectives on organizational complexity, behavior change, technology acceptance, and immersive safety, together with a targeted review of XR implementation evidence. Phase 2 (construction) refined the framework through participatory sessions. A multidisciplinary group of 33 stakeholders, including professionals and managers from hospital and primary care settings, and postgraduate students, evaluated the proposed criteria for strategic relevance and operational clarity. This process resulted in a refined 10-criterion structure and the establishment of a dual-score assessment logic. Phase 3 (validation), planned as a subsequent step, will examine how the framework performs when applied prospectively in clinical settings. The development process yielded a framework comprising 10 operational criteria grouped into 3 conceptual domains (human, organizational, and technical). Stakeholder ratings indicated high strategic relevance across all criteria, with mean scores ranging from 4.03 (SD 0.95) for workflow integration to 4.61 (SD 0.56) for safety and comfort. The final instrument applies a dual-assessment approach in which each criterion is rated separately for strategic importance and organizational readiness. Mapping these dimensions enables organizations to identify priority gaps, particularly areas of high importance and low readiness, and to distinguish between manageable constraints and critical barriers requiring targeted preparatory action prior to implementation. MCDA-XR addresses a key governance gap in XR implementation by providing a structured way to align adoption decisions with institutional priorities and operational constraints. Rather than relying on descriptive feasibility assessments, the framework is intended to support explicit prioritization and action-oriented decision-making at the organizational level. MCDA-XR is positioned for Phase 3 evaluation, which will examine the practical utility, interpretability, and implementation relevance of the framework when applied prospectively in real-world clinical deployments.
Widespread antiretroviral therapy has greatly extended the life expectancy of people living with HIV (PLWH), making cardiovascular disease (CVD) one of their primary comorbidities. Nevertheless, significant cross-specialty knowledge gaps persist in routine clinical practice. Siloed disciplinary expertise results in low clinical adherence to guideline-recommended risk management interventions, highlighting an urgent demand for integrated, evidence-based tools that break down interdisciplinary barriers. Large language models (LLMs) have demonstrated robust medical knowledge retrieval and reasoning capacity in recent years, yet no systematic evaluation has determined whether these models can bridge such knowledge gaps and facilitate multidisciplinary collaborative CVD management for PLWH. This study compared the performance of four mainstream AI models (Deepseek-V3, Deepseek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for PLWH. Based on four authoritative domestic and international guidelines on HIV and CVD care, a structured 25-question assessment battery was developed via two rounds of Delphi expert consultation, with standard reference answers and an evaluation framework finalized through expert consensus. Responses of the four LLMs were generated with standardized prompts, while 12 clinicians answered identical questions in one-on-one structured interviews, with all verbal replies transcribed verbatim. Six multidisciplinary experts independently rated all responses across four dimensions: accuracy, completeness, readability and reliability, using a 4-point ordinal scale ranging from 1 (poor) to 4 (excellent). Cumulative link mixed models (CLMMs) were applied to analyze intergroup differences. All AI models achieved statistically significantly higher scores than clinicians across all evaluation dimensions (p < 0.01). The AI group had mean scores of 3.44-3.68 (median = 4, CV: 0.145-0.178). Restricted by individual factors including specialty background, knowledge reserve, clinical experience, clinicians obtained lower mean scores of 1.78-2.05 (median = 2, CV: 0.428-0.473) with markedly greater score dispersion. Among all AI models, Deepseek-R1 delivered the optimal performance and showed statistically significant advantages over ChatGPT-4o, ChatGPT-o4-mini and Deepseek-V3 (all p < 0.01). Specialty-stratified CLMM analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (OR = 0.92, 95% CI: 0.84-1.01, p = 0.094). Dimension-specific CLMMs combined with Wilcoxon rank-sum tests confirmed that cardiologists only earned significantly higher scores in the accuracy dimension (OR = 0.81, 95% CI: 0.67-0.97, p = 0.0261). Domain-specific performance divergence was observed: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists achieved higher scores on drug adverse effect evaluation (2.23 vs 1.65). This structured Q&A study on CVD management for PLWH found that LLMs outperformed human clinicians on all assessment metrics, with Deepseek-R1 attaining a distinctly superior composite score. The findings support the promising potential of Deepseek-R1 as a cross-disciplinary decision-support tool: it integrates multi-domain complex clinical knowledge, which may help address cross-specialty knowledge barriers, could improve the completeness and precision of clinical information output, and may enhance communication and decision-making efficiency for patients with complicated multimorbidity. To maximize clinical benefits, AI systems should be integrated into multidisciplinary care workflows alongside targeted clinical training to optimize the management of complex comorbidities among PLWH.
Digital cognitive interventions (DCIs) have emerged as scalable approaches for treating cognitive dysfunction across psychiatric, neurological, and aging populations. Despite growing evidence of efficacy, little is known about which intervention components drive therapeutic effects or through which neurocognitive mechanisms they operate. As a result, null findings are often difficult to interpret, making it unclear whether interventions failed to engage their intended targets, or whether the targets themselves are not causally related to meaningful outcomes. This limits intervention refinement, comparative evaluation, and precision personalization. Here, we argue that DCI research should shift from broad efficacy testing toward mechanistic trials designed to identify active ingredients-the intervention components responsible for engaging prespecified neurocognitive targets and producing clinically meaningful benefits. We propose adapting dismantling design methodology from psychotherapy research in order to integrate Research Domain Criteria constructs, mechanistic neuroscience, and high-resolution digital behavioral data to identify factors driving cognitive and functional outcomes. This approach aligns with the National Institute of Mental Health experimental therapeutics framework by explicitly linking target specification and target engagement with downstream clinical and functional outcomes. Mechanistic dismantling trials can determine whether specific DCI features, including adaptive difficulty, reward schedules, feedback contingencies, task variability, cognitive targets, and human support, are necessary, sufficient, or synergistic for engaging neural circuitry and producing durable and clinically meaningful transfer. Beyond optimizing intervention design, such studies may transform null or negative trials into mechanistically interpretable findings, while clarifying disease mechanisms and supporting the development of personalized, optimized, and usable DCIs.
Implementing digital mental health interventions (DMHI) for those with psychosis is a persistent challenge. A process evaluation, or studies conducted alongside trials, is one research method that may address this issue. However, a synthesis of process evaluation data in this area is missing. This study aimed to understand what is known about context, implementation, and mechanisms of impact by synthesizing process evaluation data from trials evaluating DMHIs used by people with psychosis. A scoping review using a 2-phase search strategy underpinned by the Medical Research Council (MRC) process evaluation framework was conducted. Database searches of Cochrane Central Register of Controlled Trials and PsycInfo in 2024 and 2025 first identified an index sample of peer-reviewed trials predominantly conducted in the United Kingdom (≥50% of samples from the United Kingdom in multicountry studies). Next, papers linked to the index sample were retrieved and included if they reported process evaluation data as operationalized in the MRC framework. Two authors independently screened references, extracted summary data, and assessed the quality of index trials. One author qualitatively synthesized process evaluation data using a deductive framework synthesis approach using the MRC framework. Findings were triangulated with senior authors and presented as a narrative synthesis. Searches identified 14 DMHIs and 45 papers reporting process evaluation data, though only 2 were labeled as such. Qualitative syntheses of process evaluation data generated five themes aligned with the MRC framework: (1) enhancing fit and supporting delivery (implementation strategies); (2) DMHI implementation varied across users, staff, and delivery settings (implementation outcomes); (3) helping users to respond in more helpful ways (mechanisms); (4) addressing perceived and actual implementation factors (context); and (5) limited impact of user characteristics on DMHI outcomes (context). There is preliminary evidence that DMHIs can be delivered to people experiencing psychosis within trial settings, although use varied between individuals. Future implementation efforts may benefit from addressing contextual factors influencing DMHI use, including users' treatment needs and preferences, everyday demands, and staff availability for blended interventions. Future research could evaluate implementation strategies, validate how and for whom DMHIs work, and embed process evaluation in trials. PROSPERO CRD42024439117; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024439117.
Digital decision-support tools for labor care remain limited, with few technologies successfully addressing the complex, time-sensitive decisions required during labor triage. Fit4Labour is a clinician-facing, data-driven research tool, currently under development, that combines computerized cardiotocography interpretation with maternal and fetal risk factors to generate individualized risk scores at labor onset. Its primary aim is to support clinicians in identifying fetuses who may require closer monitoring or expedited delivery, while simultaneously providing reassurance in low-risk cases. By promoting consistent communication and timely escalation of care, the Fit4Labour tool seeks to strengthen clinical decision-making. Understanding and addressing usability and implementation barriers will be critical to its adoption in clinical practice. This study aims to evaluate whether a digitally co-developed labor decision-support tool (Fit4Labour) maintains usability and implementation readiness across NHS hospitals with differing clinical contexts. We conducted a convergent parallel mixed methods study in 3 United Kingdom hospitals (December 2022 to May 2025). Phase 1 involved iterative co-development with midwives and doctors at Oxford University Hospitals NHS Foundation Trust; Phase 2 validated the locked version at Birmingham Women's and Children's NHS Foundation Trust and Buckinghamshire Healthcare NHS Trust. Participants completed scenario-based usability sessions evaluated with the System Usability Scale (SUS) and Single Ease Question (SEQ), and task completion time, followed by focus groups and interviews analyzed thematically. Twenty-six health care professionals participated: 12 in co-development (7 midwives, 5 doctors) and 14 in validation (8 midwives, 6 doctors) phases. During co-development at Oxford, the tool met the "excellent" usability threshold (mean SUS 82.1, SD 12.3), indicating readiness for the validation phase. The locked version (v4.0) independently met the "excellent" threshold at both validation sites (combined mean SUS 85.8, SD 10.2; Birmingham 80.7, SD 10.8; Buckinghamshire 90.8, SD 7.2). Task completion times were comparable across validation sites (Birmingham 10.3, SD 1.6 min; Buckinghamshire 9.2, SD 1.9 min), while SEQ scores were consistently high across all scenarios (mean 6.1/7, SD 0.8). Thematic analysis identified 12 themes within 3 domains: clinical integration and workflow, technology adoption and implementation, and patient safety and decision-making. Participants described the Fit4Labour tool as a supportive tool, "like a co-pilot," improving confidence in decisions with the potential to aid triage assessment. Perceived limitations included an incomplete risk factor profile and the need for minor technical adjustments or integration with existing hospital systems. Through systematic co-development, the Fit4Labour tool met the established usability benchmark at 2 independent NHS hospitals with markedly different clinical contexts. Clinicians viewed the tool as a supportive aid providing a shared language for risk communication and enhanced decision-making while preserving clinical autonomy. Whether these usability findings translate to improved clinical outcomes in real-world practice requires prospective evaluation.
Posttraumatic stress, along with comorbid mental health challenges and hazardous alcohol use, disproportionately affects people living with HIV. The drivers of these stressors are both intraindividual, rooted in early life adversity and firsthand violence exposures, and contextual, often place-based. Imparting effective coping skills and distinguishing between changeable and unchangeable stressors can improve stress management in the short term, with cascading effects on key HIV continuum of care end points, such as antiretroviral therapy adherence. However, problem- and emotion-based coping skills, delivered via traditional linear in-person group modalities, may falter in the moment. To address this, we adapted the evidence-based Living in the Face of Trauma intervention into an iOS- and Android-native app, featuring daily diary-triggered coping skills recommendations, self-guided Living in the Face of Trauma psychoeducational sessions, and a customizable geofencing function. This mixed methods study aimed to examine the acceptability, feasibility, and user experiences of NOLA (New Orleans, Louisiana) Gem, focusing on user interaction costs relative to geographic ecological momentary assessment (GEMA) alone and refining future optimization options. People living with HIV (N=32) were recruited across New Orleans and initially randomized 1:1 to treatment (NOLA Gem + GEMA) versus control (GEMA) for 21 days. Feasibility was assessed via enrollment and attrition rates. At the immediate postassessment, participants completed acceptability and usability measures and a brief structured usability interview. Analyses included descriptive statistics, bivariate logit modeling, and synergistic human-large language model deductive coding. In total, 30 participants (n=22 in the GEMA + NOLA Gem treatment arm) completed the pilot, representing 94% (n=29) of baseline enrollees. Acceptability was very high across the board: 100% (n=30) of users considered NOLA Gem "very" or "somewhat" successful in addressing their daily lives, with 91% (n=28) endorsing increased calm and emotional well-being. In addition, 50% (n=11) of NOLA Gem users were "extremely likely" (Net Promoter Score=10/10) to recommend the app to friends. Eight (27%) GEMA and GEMA + NOLA Gem users reported privacy concerns. Eleven (50%) NOLA Gem users received geofencing alerts; perceptions of this feature's helpfulness were mixed. No statistically significant sociodemographic or clinical predictors of disparate acceptability or increased privacy concerns were found. No additional frictions were evidenced by GEMA + NOLA Gem versus GEMA users. Qualitatively, NOLA Gem users praised the just-in-time mindfulness, breathing, problem-solving skills delivery, and broader stress control and self-insight benefits. A subset of users pointed out the burdensome length and sometimes inconvenient timing of the daily diaries. Recommendations for next-generation personalization included user-specific dynamic daily diary and geofencing prompt tailoring. Our small pilot study demonstrated high NOLA Gem acceptability and feasibility, as well as a rich and beneficial user experience among people living with HIV, with clear and actionable opportunities for improvement.
Multidomain dementia-prevention interventions delivered via apps have the potential to reach large populations. However, existing trials have tended to recruit more socioeconomically advantaged participants, raising concerns that the resulting interventions may be less usable for older adults from minority ethnic, lower educational, or lower socioeconomic backgrounds, who are at higher risk of dementia. The Tailored Intervention for Brain Health and Cognitive Enrichment (ENHANCE) app was designed to address this by prioritizing accessibility and engagement across diverse user groups. This study evaluated the usability and user experience of the ENHANCE prototype during a 1-week at-home supported-use test and explored factors influencing use and engagement among older adults. We purposively recruited adults aged 60-80 years without dementia for a 1-week mixed methods usability evaluation through community settings, including groups underrepresented in dementia-prevention trials. Participants had at least 1 of 10 prespecified dementia risk factors, attended a face-to-face onboarding session with a coach, used the app at home for 7 days with ongoing coach support, and completed a posttest interview and satisfaction survey. We analyzed quantitative data, including app usage metrics and survey responses descriptively, and used reflexive thematic analysis of qualitative data from onboarding sessions, posttest interviews, coaching calls, and in-app messages. Ten participants participated in the study. The mean age was 68 (SD 6) years, and 7 were females. Participants represented a wide range of deprivation (Index of Multiple Deprivation deciles: mean 4, SD 2, range 1-8), with 6 from ethnic minority backgrounds. All met prespecified minimum-use targets (watching a module video, completing a check-in, and playing assigned games at least once). Many demonstrated additional voluntary engagement: 5 rewatched the risk factor video, 7 used the in-app messaging feature, and among the 5 participants with the hypertension module, blood pressure was logged on an average of 5 out of 7 days. Survey responses indicated high satisfaction, perceived usefulness, and ease of use; 9 participants intended to continue using the app and would recommend it to peers. Qualitative analysis identified engagement facilitators, including rewarding game design, familiar interfaces, appropriately challenging gameplay, consistent virtual rewards, trusted expert information combined with peer stories, and coach support. Barriers included unclear visual cues, insufficient accommodation of motor or sensory impairments, and visual discomfort in some games. Older adults found the ENHANCE prototype usable, acceptable, and engaging over 1 week. Human coaching, inclusive design, and integration of expert and peer narratives were highlighted as key engagement drivers. These findings support further feasibility testing to examine longer-term engagement and provide design insights for more inclusive digital health interventions.