With the global proliferation of chronic diseases and sudden infectious outbreaks, the use of artificial intelligence and digital technologies to enhance the sustainable growth of public health services has become a key research focus. This study examines the impact and transmission mechanisms of artificial intelligence and digital upgrading on the long-term development of China's public health services. It visualizes and performs regression analysis on panel data from 30 provincial-level units in China from 2012 to 2024, using kernel density estimation, standard deviation ellipse, and dual machine learning models. The following conclusions are drawn: (1) The sustainable development of public health services in China shows a regional distribution pattern, with higher levels in the east and lower levels in the west. Although overall levels have improved over time, regional disparities have widened. Hotspots for sustainable growth also show a spatial development trend toward the southeast. (2) Artificial intelligence and digital upgrading significantly positively impact the sustainable growth of public health services in China. A one-unit rise in artificial intelligence and digital enhancement results in gains of 0.014% and 0.080% in the sustainability of public health services, respectively. (3) The positive effects of artificial intelligence and digital upgrading on the sustainable development of public health services exhibit heterogeneity across economic zones, resource endowments, and the North-South regional division. (4) Digital upgrading and artificial intelligence significantly enhance the development of green technological innovation, green patent technology innovation, and green utility model innovation. Through this pathway, the sustainable development performance of public health services will be further improved. Digital upgrading and artificial intelligence improve the sustainable development of public health services in China through multiple pathways, including spatial distribution dynamics, direct positive effects, heterogeneous regional impacts, and enhanced green technological innovation.
To propose the Health Research Standard Process for Artificial Intelligence framework, adapted from the Cross-Industry Standard Process for Data Mining approach, and apply it using time series and artificial intelligence techniques to analyze data from the Sistema de Informação de Agravos de Notificação (SINAN - Notifiable Diseases Information System) on violence against women in São Paulo, exploring sociodemographic, spatial, and predictive dimensions. The methodology was adapted to health research, replacing the original Business Understanding stage with Research Understanding by incorporating essential technical-scientific elements. We analyzed 80,148 reports of violence against women aged between 20 and 59 living in the city of São Paulo between 2013 and 2023. The analyses included descriptive statistics, calculation of prevalence ratios, decomposition and predictive modeling of time series capturing trends and seasonality, and geospatial analysis of notifications. We observed temporal patterns and sociodemographic characteristics of the victims, with increasing trends in reports of physical, psychological, and sexual violence after 2015 and seasonal behavior. Among the associations, there was a 32% increase in the prevalence of sexual violence when the aggressor was under the influence of alcohol and a 2.47 times greater risk of sexual violence among pregnant women. The geospatial distribution revealed concentrations of notifications in peripheral areas of the municipality, such as the south, east, and central zones. Predictive modeling indicated that the upward trend will persist over the next 24 months, with estimated rates reaching up to 12 cases per 100,000 inhabitants. The applicability of the Health Research Standard Process for Artificial Intelligence was effective as a model for data analysis and the development of predictive algorithms in public health. It was possible to propose a replicable methodological matrix for future research and evidence-based interventions.
The growing bases of Artificial Intelligence (AI) applications ranging from diagnostic support to immersive training have rapidly advanced in the dental education field. Endodontics, by its very nature of relying so highly on a proper diagnosis and careful technical execution, is an indication through which AI may be best poised to succeed the most in specialty care. The aim of this systematic review was to assess the role of artificial intelligence (AI): machine learning (ML), deep learning (DL), virtual/augmented reality (VR/AR) and large language models (LLMs) related to endodontic education based on available evidence published until September 2025. This review was performed in accordance with the PRISMA 2020 guidelines. Publication databases were reviewed included PubMed, Scopus, Web of Science and Cochrane. Inclusion Criteria: Studies that evaluated any form of AI for didactic, preclinical or clinical education in endodontics and/or patient-centered education were included. Study characteristics, AI domains, applications and outcomes were extracted. Risk of bias and methodological quality were evaluated according to study design using RoB 2, ROBINS-I, AXIS, and AMSTAR-2 tools. Fifteen studies were included. Radiographic interpretation augmented by AI improved sensitivity and specificity to reduce false positive reporting especially for junior clinicians. In preclinical training, VR/AR simulations have shown to improve psychomotor skills, confidence and knowledge acquisition. LLMs can be useful in producing exam questions and case-based Q&A, although the accuracy and discriminatory ability varied. AI mediated Patient education interventions led to anxiety reduction and comprehension. There was heterogeneity of outcome measures, dataset bias; reliability and transparency issues. AI holds promise for use in diagnostic, didactic and preclinical endodontic education. They must be safely implemented in a controlled format, under the supervision of faculty and with objective evaluation metrics in place. AI provides quantifiable benefits in endodontic education by improving accuracy of diagnosis, assisting decision-making and facilitating dental students training using VR/AR simulation. Some interventions using AI in curricula may allow the student to acquire skills faster, feel more confident, and transfer these benefits to improved patient communication. But we need to make sure our integration is backed up with faculty monitoring, transparent AI models and rigorous validation before putting it in any production environment or relying on it too heavily for exam outcomes.
Artificial intelligence (AI)-assisted endoscopy represents a promising approach for lesion detection, yet frequent false-positive detections impair clinical utility by disrupting examinations and diminishing physician confidence. Linked-color imaging (LCI), an image-enhanced endoscopy technique that amplifies mucosal and vascular contrast, may address this limitation. This investigation evaluated whether LCI reduces false-positive AI detections compared with white-light imaging (WLI). This retrospective study analyzed consecutive AI-assisted upper endoscopies performed between March 2024 and June 2025. WLI and LCI were performed sequentially within the same endoscopic session in each patient. False-positive AI detections were compared between modalities using two computer-aided detection (CAD) versions. Propensity score adjustment was used as a sensitivity analysis for baseline differences between CAD Versions I and II. Of 66 initially screened cases, 63 remained after excluding patients with prior gastric surgery. LCI reduced false-positive AI detections compared with WLI (median 2 vs. 5; p < 0.001). In CAD version-stratified sensitivity analyses, LCI reduced false-positive AI detections in both Version I (5 to 2; p = 0.01) and Version II (2 to 0; p = 0.03). This reduction remained consistent across atrophic grades. Both imaging modalities identified all gastric lesions, achieving 100% detection sensitivity. LCI assessment performed after WLI observation yielded fewer false-positive CAD-EYE detections while maintaining lesion detection sensitivity. However, because the observation sequence was fixed, these findings should be interpreted cautiously and require confirmation in prospective or counterbalanced studies. Trial Registration: N/A (retrospective study).
Advances in three-dimensional (3D) transthoracic echocardiography (TTE) with artificial intelligence (AI) enable AI-assisted semi-automated right ventricular (RV) function analysis, including two-dimensional (2D) measurements. However, data on the agreement of these semi-automated analyses in routine clinical practice remain limited. This study evaluates the agreement of AI-based 3D TTE semi-automated measurements compared to manual 2D TTE measurements. We enrolled 201 patients who underwent both 2D and 3D TTE between July and November 2023. AI-assisted semi-automated measurements of right ventricular morphology and functional parameters were performed using 3D Auto RV software. While echocardiographic images were acquired manually by the operators in a conventional manner, the AI-assisted analysis was applied specifically during the post-acquisition offline processing stage. The feasibility of RV measurement using the AI-based software was 84.8% (201/237) among patients with available 3D datasets, and the overall applicability was 67.4% (201/298) across all consecutive patients undergoing routine echocardiography. Among the 201 analyzed cases, fully automated analysis without manual correction was feasible in 10.9% (22/201), while the remaining 89.1% required manual adjustments of the endocardial borders to ensure clinical accuracy. Time for analysis was significantly shorter with the AI-assisted semi-automated method compared to the manual 2D method (30% reduction, p < 0.001). AI-based 3D TTE semi-automated measurement has the potential to enhance the efficiency and reproducibility of RV function assessment, reducing examiner workload and improving clinical workflow.
Artificial intelligence (AI) is rapidly integrating into clinical radiology. As primary diagnosticians, radiologists increasingly interpret AI-generated analyses and are expected to oversee the monitoring and governance of deployed AI systems. Although AI literacy among radiologists is improving, several technical aspects of AI remain insufficiently accessible. One such concept is uncertainty quantification (UQ), which estimates the reliability of AI predictions and can signal when outputs should be interpreted with caution. This review introduces key UQ concepts relevant to radiology, distinguishing between aleatoric uncertainty and epistemic uncertainty arising from data variability and knowledge gaps. We summarize commonly used UQ approaches in current research and practice. Furthermore, through a narrative review of selected recent AI imaging studies, we illustrate how UQ methods are applied in practice and highlight methodological trends, findings, and limitations. Although UQ has the potential to improve the safety and interpretability of AI-assisted screening, challenges remain, including calibration, threshold selection, computational cost, and the need for prospective clinical validation.
The integration of artificial intelligence (AI) into the medical and spiritual care of the sick is expected to challenge a range of religious values and norms. Various religious authorities have expressed concerns about AI's integration, which include the questionable ability of AI to express genuine empathy, the risk of humans abandoning caretaking roles, the misuse of AI for religious rites, the attributing of spiritual and supernatural qualities to machines leading to idolatry, the undermining of provider responsibility and accountability, the dereliction of human intellect, the risk of bias and inappropriate proselytization, and threats to health equity. This manuscript identifies common concerns across major religious traditions, including Judaism, Christianity, Islam, Hinduism, and Buddhism. AI integration should align with values that promote human health and flourishing. Such values are often rooted in religious teaching, and understanding the reservations that religious communities have toward AI will help ensure such alignment.
This quality improvement study evaluates the use of a video-based artificial intelligence framework that uses multi-instrument tracking and clinically informed spatiotemporal kinematic features to provide objective, granular assessment of laparoscopic cholecystectomy skill.
This study provides a comprehensive analysis of the application of artificial intelligence (AI) in diagnosing acute traumatic wrist joint injuries (WJIs), including fractures and ligament damage. AI has demonstrated significant potential in identifying fractures and ligament damage. The study highlights the use of various AI technologies and algorithms, including Convolutional Neural Networks (CNNs), Gradient Class Activation Mapping (Grad-CAM), deep learning models, object detection models, automated assessment algorithms, traditional machine learning techniques, data augmentation and preprocessing, Natural Language Processing (NLP), and integration with other imaging modalities. Compared with traditional diagnostic methods, AI offers substantial benefits, such as efficient processing of large datasets, minimizing diagnostic errors and missed cases, aiding in interpreting complex fracture patterns, optimizing workflow, enhancing diagnostic efficiency, and providing comprehensive diagnoses through multimodal integration. AI has significantly improved the precision and efficiency of fracture detection, reduced unnecessary imaging procedures, expedited the diagnostic and reporting process, optimized resource allocation, and improved patient outcomes. However, clinical application of AI faces challenges, including ethical considerations, regulatory hurdles, data privacy and security issues, algorithm transparency and interpretability problems, and unclear liability definitions. The future of AI in diagnosing WJIs is promising, with potential advancements in fracture detection, treatment planning, and rehabilitation strategies. However, challenges such as model validation and training of healthcare professionals must be addressed to fully integrate AI into orthopedic practice and advance the management of WJIs.
Generative Artificial Intelligence (GenAI) has catalyzed a transformation in medical education. Understanding learners' perceptions is essential to guide their responsible integration into curricula. A cross-sectional survey was administered to 1039 undergraduate medical students across Saudi Arabia. A purpose-developed, pilot-tested instrument (Cronbach's α = 0.71 for dichotomous items) assessed students' familiarity with computational language models, perceptions of their educational utility, and attitudes toward technology-enhanced pedagogical approaches. Descriptive statistics, Kolmogorov-Smirnov testing for normality, and multivariable binary logistic regression (two-sided, α = 0.05) were performed using IBM SPSS Statistics v28.0. Among 1039 participants (64.3% male; median age 22 years [IQR 20-24]), 57.2% (595/1039) reported familiarity with computational language models in medical education, and 70.1% (728/1039; 95% CI: 67.2-72.9) supported curricular integration. A strong majority (86.4%; 898/1039; 95% CI: 84.2-88.4) anticipated impact on the future of medical education. While 73.4% (763/1039; 95% CI: 70.6-76.0) perceived benefit for basic science education, only 41.6% (432/1039; 95% CI: 38.6-44.6) recognized utility in clinical skills training. Only 29.8% (310/1039; 95% CI: 27.0-32.7) considered these tools superior to human instruction. Key concerns included distrust in output reliability (52.6%; 547/1039; 95% CI: 49.5-55.7) and awareness of reference fabrication (64.0%; 665/1039; 95% CI: 61.0-66.9). Saudi medical students express strong interest in GenAI-particularly for basic sciences and simulation-but perceive it as complementary rather than superior to human instruction. Findings reflect learner perceptions, not measured educational effectiveness. Implementation should prioritize reliability, ethical use, and preservation of humanistic competencies.
Global maternal health outcomes remain inequitable, particularly in low- and middle-income countries, due to persistent gaps in access, care continuity, and postpartum follow-up. Digital health tools, especially mobile health and artificial intelligence, offer promising avenues to enhance education, monitoring, risk stratification, and clinical decision-support. This umbrella review synthesizes evidence from existing reviews on AI applications in maternal health, focusing on: (1) AI for predicting and stratifying risks of pregnancy-related complications and mortality, (2) AI-driven preventive and supportive interventions from pregnancy through postpartum, and (3) implementation barriers to real-world scale-up. It also proposes a practical roadmap for developing a scalable AI-enabled maternal health solutions.We searched PubMed, Scopus, and Web of Science for English-language reviews (2000-2025). Two reviewers independently screened records, assessed eligibility, and evaluated methodological quality using AMSTAR 2; low-quality reviews were excluded. Thirty-four reviews were included. Commonly employed AI methods included logistic regression, random forests, gradient boosting, support vector machines, and deep learning. Evidence indicates AI's potential for risk prediction in hypertensive disorders, gestational diabetes, preterm birth, perinatal mental health, and other complications. However, a significant translation gap persists: external validation, model interpretability, clinical workflow integration, fairness auditing, and prospective impact evaluation were inconsistently addressed. This umbrella review suggests that AI tools can predict pregnancy complications, but a major translation gap remains-lack of external validation, interpretability, workflow integration, fairness audits, and prospective impact assessment. Addressing these gaps is essential for equitable scale-up, especially in low- and middle-income countries. We propose a practical roadmap to guide future development from pregnancy through postpartum.
Purpose To characterize Predetermined Change Control Plan (PCCP) adoption and documentation transparency among U.S. Food and Drug Administration (FDA)-cleared radiology artificial intelligence/machine learning (AI/ML)-devices (2015-2025). Materials and Methods A cross-sectional systematic scoping review was conducted with linked data from FDA AI/ML-enabled device databases through April 2026. PCCP documentation completeness was scored by two independent observers (intraclass correlation coefficient: 0.93) using an 8-point rubric (possible scores, 0-8) derived from FDA's final PCCP guidance (December 2024). Identified PCCP devices underwent manual verification against FDA regulatory summaries. Results Among FDA-listed AI/ML-device submissions, 1080/1394 (77.5%) were radiology submissions, and nearly all cleared via 510(k) review. Across all FDA panels, 170 devices were cleared with a PCCP. Radiology led AI/ML-specific PCCP adoption (34/37 radiology PCCP devices, 91.9%). Among the PCCP-cleared radiology AI devices, 22/34 (65%) were cleared in 2025 alone following the final FDA guidance. Discrepancies between FDA's public database and individual summaries required manual adjudication for 9/34 (27%) of devices. PCCP documentation scores ranged from 0 to 8 (mean, 5), with most modifications focused on data retraining, compatibility expansion, and algorithm optimization. Continuous monitoring of device performance and predefined drift triggers for retraining were absent from public summaries. Conclusion PCCP adoption increased in radiology after issuance of the final FDA guidance, yet public lifecycle controls, particularly monitoring performance metrics and trigger thresholds, were limited. Standardized PCCP reporting of lifecycle controls are suggested as a condition of PCCP authorization to enable systematic postmarket monitoring as this pathway scales. ©RSNA, 2026.
Artificial intelligence (AI) is conquering medicine in many fields. With geriatric patients, it is important not only to understand the decision tree of, for example, a tumor disease, but also to assess their functionality and functional deficits. For potential (urological) care of geriatric patients, AI-supported tumor boards, AI-simulated human interaction, and AI-based polypharmacy tools were identified. The advantages and disadvantages, risks and benefits, and practical aspects are weighed against each other. The specific characteristics of geriatric patients arise from their multidimensionality. These often involve functional impairments that can only be systematically assessed. The challenge will be to integrate these functional deficits, along with appropriate cut-offs, life expectancy, patient goals, anamnesis, and the potential risks of any therapy, into treatment decisions, in addition to the characteristics of a (tumor) disease. Ensuring the psychosocial care of nursing home residents during times of staff shortages also presents a challenge. "Learning" and "responsive" cuddly toys and robots could provide valuable support in this area. A third application of AI is certainly a comprehensive medication tool that compares multiple databases with the patient's specific needs. AI applications in medicine and urology are inevitable. Their integration into the care of the particularly complex and vulnerable geriatric patient requires careful consideration: What information is needed regarding the patient, their illness, therapy, and therapy risks, and how can this information be applied? The question of whether AI is a curse or a blessing can perhaps currently be rephrased as "opportunities, but also risks." FRAGESTELLUNG: Künstliche Intelligenz (KI) erobert die Medizin in vielen Feldern. Bei dem geriatrischen Patienten gilt es, nicht nur den Entscheidungsbaum etwa einer Tumorerkrankung nachzuvollziehen, sondern auch seine Funktionalität bzw. funktionellen Defizite zu erfassen. Für eine mögliche (urologische) Betreuung geriatrischer Patienten konnten das KI-gestützte Tumorboard, die KI-simulierte menschliche Zuwendung und das KI-basierte Multimedikationstool identifiziert werden. Für und Wider, Gefahren und Nutzen und praktische Aspekte werden gegeneinander abgewogen. Die Besonderheiten des geriatrischen Patienten generieren sich aus seiner Multidimensionalität. Häufig handelt es sich um Funktionsausfälle, die nur mit Assessments strukturiert erfasst werden können. Die Herausforderung wird sein, diese Funktionsdefizite mit entsprechenden Cut-offs, Lebenserwartung, Zielen des Patienten, anamnestischen Hinweisen, den möglichen Gefährdungspotentialen einer eventuellen Therapie zusätzlich zu den Charakteristika einer (Tumor)Erkrankung in eine Therapieentscheidung einfließen zu lassen. Die psychosoziale Betreuung eines Heimbewohners in Zeiten von Personalknappheit zu gewährleisten, stellt ebenso eine Herausforderung dar. „Lernende“ und „reagierende“ Kuscheltiere und Roboter könnten hier unterstützend wirksam sein. Eine dritte Anwendung von KI stellt sicherlich ein umfassendes Medikationstool dar, das multiple Datenbanken mit den Besonderheiten des Patienten abgleicht. Die KI-Anwendungen in der Medizin und der Urologie werden nicht aufzuhalten sein. Ihr Einzug in die Betreuung des besonders komplexen und vulnerablen geriatrischen Patienten bedarf einer gründlichen Vorüberlegung: Welche Informationen werden im Hinblick auf Patient, Erkrankung, Therapie und Therapierisiken benötigt und wie lassen sich diese anwenden? Die Frage nach Fluch oder Segen lässt sich derzeit vielleicht in „Chancen, aber auch Risiken“ umformulieren.
Falls among older adults are a leading cause of morbidity and loss of independence. Wearable sensors combined with machine learning (ML) offer opportunities for objective fall risk evaluation, but low model transparency limits clinical adoption. Interpretable and explainable artificial intelligence (XAI) methods can address this constraint, yet their application in wearable sensor-based fall risk assessment has not been systematically examined. A PRISMA 2020-compliant systematic review was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore. Studies were eligible if they included older adults, employed wearable sensors, hybrid sensor systems, or structured clinical assessment instruments, applied AI/ML for fall risk assessment (not detection), and incorporated intrinsic interpretability or post-hoc XAI. Data were extracted on population characteristics, sensor modalities, outcome definitions, ML algorithms, explainability strategies, validation methods, and predictive performance. Eleven studies (2019-2025, total n = 5,484) met inclusion criteria. Inertial Measurement Units predominated. Fall risk definitions were heterogeneous, spanning retrospective fall history, clinical balance scales, and prospective diaries. Explainability was implemented almost exclusively at the global level, with only two studies providing both global and local explanations. On comparable tasks, interpretable models achieved accuracy of 0.65-0.92 and AUC of 0.70-0.92, suggesting interpretability carried no consistent performance penalty over more complex designs. All studies relied on internal validation only; none performed external validation or real-time deployment. None recruited P&O users or incorporated device-specific predictors. Current models favour interpretable architectures and achieve moderate-to-high performance, but are constrained by heterogeneous outcome definitions, absence of external validation, and global-only explainability that limits individual-level clinical utility. The evidence base does not yet support clinical deployment. Extending these findings to prosthetics and orthotics users, a clinically important downstream application, will require device-specific datasets, asymmetry-adjusted thresholds, and instance-level explanations. These represent the priority directions for the next stage of this research agenda.
Subclinical retinal microvascular remodeling may occur before clinically detectable diabetic retinopathy (DR). This study investigated artificial intelligence (AI)-derived ultra-widefield retinal vascular metrics in patients with type 2 diabetes mellitus (T2DM) with and without non-proliferative DR (NPDR), and evaluated their potential for early vascular phenotyping and diagnostic discrimination. In this observational cross-sectional single-center study, 237 participants were included: 63 healthy controls (103 eyes), 101 patients with T2DM without DR (No-DR; 201 eyes), and 73 patients with NPDR (132 eyes). Non-mydriatic 200-degree ultra-widefield fundus images were analyzed using an AI-based vascular segmentation and quantification system. AI-exported values coded as -1 were treated as missing, and sparse parameters were excluded from primary inference. Intergroup comparisons of retained vascular parameters were performed using age- and sex-adjusted mixed-effects models with participant as a random intercept, followed by Benjamini-Hochberg false discovery rate (FDR) correction. Multiparameter logistic models were evaluated using 5-fold subject-level cross-validation. After quality control, covariate adjustment, and FDR correction, 38 vascular-parameter rows remained significant. Whole-field vessel density was highest in the No-DR group, intermediate in NPDR, and lowest in controls [control, 0.017 (0.008-0.023); No-DR, 0.026 (0.020-0.032); NPDR, 0.021 (0.015-0.026); FDR P < 0.001]. Fractal-dimension metrics showed similar early alterations, with arterial fractal dimension increased in No-DR and intermediate in NPDR [control, 1.287 (1.164-1.359); No-DR, 1.380 (1.328-1.420); NPDR, 1.329 (1.271-1.374); FDR P < 0.001]. Whole-field mean vessel diameter was lower in No-DR than in controls, whereas total vessel length decreased across controls, No-DR, and NPDR (FDR P < 0.001 for both). Regional heatmaps showed that significant density and fractal-dimension signals clustered mainly in the superotemporal, superonasal, and inferotemporal regions. Cross-validated models showed good discrimination for No-DR versus control (AUC, 0.853; 95% CI, 0.807-0.893), NPDR versus control (AUC, 0.785; 95% CI, 0.720-0.841), and No-DR/NPDR versus control (AUC, 0.830; 95% CI, 0.788-0.873), but more modest discrimination between NPDR and No-DR (AUC, 0.658; 95% CI, 0.596-0.720). AI-derived ultra-widefield retinal vascular metrics demonstrate early, spatially heterogeneous microvascular remodeling in T2DM before clinically apparent DR. Vessel density, fractal dimension, vessel diameter, and vessel length provide complementary information, and multiparameter vascular modeling may support early detection and risk stratification. External validation and longitudinal studies are required before clinical implementation.
To develop an agentic artificial intelligence (AI) framework that streamlines and standardizes eye bank operations by automating donor screening, image analysis, and tissue suitability assessment under expert supervision. A modular, on-premises, multi-agent system was designed consisting of three AI agents. The Donor Screening Agent analyzes medical histories to identify contraindications for ocular donation. The Corneal Image Analysis Agent interprets multimodal images, including specular microscopy, slit-lamp, and optical coherence tomography, to quantify endothelial cell density, morphology, and clarity. The Suitability Assessment Agent integrates outputs from both preceding agents to generate structured recommendations for transplantation, research use, or rejection. Each agent operates locally to ensure privacy, real-time processing, and compliance with health information regulations. The proposed architecture replaces fragmented manual workflows with a unified, intelligent system that reduces variability, accelerates evaluations, and enhances transparency. By enabling consistent and rapid analysis of donor data and corneal images, the system can improve tissue utilization and decrease discard rates. A modular, low-latency design supports deployment in diverse settings, including low- and middle-income countries, where connectivity and infrastructure may be limited. Agentic AI offers a scalable pathway toward a standardized eye banking process by combining automation with human oversight. This approach has the potential to improve decision quality, operational efficiency, and equity in corneal transplantation. It carries implications for both limited resource settings and the broader context of transplantation medicine.
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
Systematic literature reviews offer high potential for efficiency gains from artificial intelligence (AI), now integrated into several systematic literature review software platforms. Validation studies show acceptable sensitivity, specificity, and accuracy for AI-assisted systematic literature reviews of clinical trial publications. Unlike trials, economic model publications lack consistency in content, terminology, and structure. We aimed to test the efficiency and accuracy of AI-assisted search, screening, and data extraction when applied to a systematic literature review of economic evaluations. A previously conducted manual systematic literature review of economic evaluations for chronic rhinosinusitis with nasal polyps was replicated using a machine learning-based inclusion prediction model (Robot Screener) and a large language model-based criteria screener (Smart Screener) within Nested Knowledge software, with performance benchmarked against the original human-conducted systematic literature review. The AI-generated search retrieved 22/43 (51%) PubMed articles from the original systematic literature review. Accuracy exceeded 95% for title/abstract screening but fell below 80% for full-text screening. Extraction was reliable for high-level model descriptors and general study characteristics, but less so for model structures, health states, outcomes, and distinguishing sensitivity from scenario analyses and complex modeling assumptions for duration of response, discontinuation, surgery, and mortality. Estimated time savings ranged from ~20% (data extraction) to 60% (title/abstract screening and searches), varying by task human validation requirement. Artificial intelligence-driven tools performed well for title/abstract screening and general data extraction but were less accurate for full-text screening and interpretation of modeling choices. They can increase systematic literature review efficiency for economic evaluations but fall below the reliability seen for systematic literature reviews of clinical trials. Artificial intelligence has the potential to speed up systematic literature reviews by assisting researchers with identifying articles and organizing information from the articles. In this study, artificial intelligence-assisted tools embedded within a commercial systematic literature review platform were applied to replicate a systematic review of economic evaluations, providing practical insights into their readiness for use in health economics and outcomes research and health technology assessment workflows. Artificial intelligence performed well in early-stage screening, achieving more than 95% accuracy when reviewing titles and abstracts but its accuracy dropped below 80% when screening full texts. Artificial intelligence also reliably extracted general data such as study perspective, comparators, geography, and time horizon, but it struggled with more complex elements and those that required some interpretation and experience with these types of studies, including model structure, health states, sensitivity analyses, and key modeling assumptions. Estimated time savings ranged from approximately 20% for data extraction to 60% for title and abstract screening and database searches, though savings were highly dependent on the approach taken and the extent of human validation required. Overall, the study shows that artificial intelligence can improve efficiency in aspects of the process of conducting a systematic literature review of economic evaluations by streamlining general tasks but still requires substantial human oversight.