Adding Artificial Intelligence to the PERT Improves Time to Diagnosis and Treatment of Pulmonary Embolism.
PubMed2026-08-08
There is limited data demonstrating the benefit of artificial intelligence technology in the diagnosis and triage of pulmonary embolism. Our study aims to demonstrate improved time to diagnosis of pulmonary embolism and subsequently time to anticoagulation and intervention with the goal of reducing in-hospital mortality. We hypothesized that implementation of artificial intelligence-assisted computed tomography pulmonary angiogram detection would reduce time to diagnosis, time to anticoagulation, and time to intervention compared to standard radiology-based workflow.
A single institution retrospective review from July 2018 to March 2025 was performed to identify patients diagnosed with pulmonary embolism who underwent pulmonary angiogram with mechanical thrombectomy and/or thrombolytics. Patients were divided into a pre-AI cohort (July 2018-2022) and a post-AI cohort (2022-March 2025) corresponding to the institutional implementation of Viz.ai PE (Viz.ai, San Francisco, CA), an FDA-cleared, HIPAA-compliant artificial intelligence platform for automated pulmonary embolism detection on computed tomography pulmonary angiogram. Time to diagnosis was defined as the interval from computed tomography pulmonary angiogram scan completion to artificial intelligence-generated alert (post-AI cohort) or to final radiology report issuance (pre-AI cohort). Time to anticoagulation and time to intervention were measured from the time of confirmed pulmonary embolism diagnosis. In-hospital mortality was also evaluated.
From July 2018 to March 2025, 148 patients were diagnosed with pulmonary embolism and underwent endovascular intervention. Twenty-four patients were excluded. Forty-two patients were diagnosed in the pre-AI era and 82 in the post-AI era. The median age was 65 (IQR 53-73) years and 65.5 (IQR 52-73.5) years respectively. Time to diagnosis improved significantly from 72.5 (IQR 44.5-93.3) minutes to 40 (IQR 28-68.2) minutes (p = 0.00005). Time to anticoagulation was 76 (IQR 48.3-99.8) minutes vs 61.5 (IQR 46.8-110) minutes (p = 0.824, not significant). Time to intervention improved from 1360 (IQR 1075.5-1790.2) minutes to 1224 (IQR 601.6-1563) minutes (p = 0.036). Two in-hospital deaths occurred, both in the pre-AI cohort.
Implementation of artificial intelligence-assisted computed tomography pulmonary angiogram detection using Viz.ai PE significantly improved time to diagnosis and time to intervention in patients with acute pulmonary embolism requiring catheter-directed therapy. Time to anticoagulation was not significantly different between groups; this was study was insufficiently powered to detect differences in clinical outcomes including mortality.
Journal of vascular surgery. Venous and lymphatic disorders
查看原文 ↗Artificial Intelligence in the Detection, Characterization, and Management of Renal Masses: A Narrative Review.
PubMed2026-07-01
Kidney tumors are being found more often today because CT and MRI scans are widely used and kidney masses are frequently discovered by chance. Many of these masses are benign. However, most patients still undergo surgery because doctors cannot confirm the diagnosis from imaging alone. Artificial intelligence (AI) refers to computer systems that perform tasks normally requiring human intelligence, including machine learning, in which algorithms learn patterns directly from data, and radiomics, in which quantitative features are extracted from medical images to support diagnosis. These tools offer new ways to improve how kidney masses are detected, characterized, and treated. This narrative review summarizes current evidence on AI applications across the renal mass pathway, covering automated detection, imaging-based characterization, pathological staging, prognosis prediction, and AI-assisted surgical planning, while also outlining current limitations and future research directions. A review of the English-language literature was performed using PubMed and MEDLINE, with studies published between 2018 and 2026 prioritized. AI shows strong performance across all stages of the renal mass pathway. Deep learning models accurately detect and segment renal masses on CT and MRI. Radiomics-based classifiers distinguish benign from malignant lesions and predict tumor subtype without biopsy. Multimodal AI models predict survival with high accuracy and outperform established clinical scoring systems. AI-assisted surgical planning tools support nephron-sparing surgery and predict postoperative kidney function. Wider clinical use requires better prospective validation, more diverse datasets, and improved model transparency.
Missed and true interval cancers on digital breast tomosynthesis screening mammograms: radiologists' assessment and artificial intelligence markings.
PubMed2026-08-01
Interval cancer, breast cancer detected after a negative screening examination but before the next scheduled appointment, represents a challenge in mammography screening programs due to less favorable histopathological characteristics compared to screen-detected cancer.
To determine which interval cancers from digital breast tomosynthesis (DBT) were classified as missed and true by radiologists in a review, and to stratify the findings by risk scores and markings provided by an artificial intelligence (AI) model.
In this retrospective informed consensus-based review, radiologists assessed mammograms from 46 interval cancers and classified those as false negative, minimal-sign significant or non-specific, or true negative. An AI risk score (1-7, low; 8-9, intermediate; or 10, high risk of malignancy) was available for each examination. For cases with AI risk scores of 8-10, the location of AI-detected markings was compared with the true cancer site.
A total of 17% (8/46) of interval cancers were classified as false negative, 22% (10/46) as minimal-sign significant, 20% (9/46) as minimal-sign non-specific, and 41% (19/46) as true negative. The AI model correctly identified 35% (16/46) and incorrectly located 24% (11/46). Considering false negative and minimal-sign significant as cases with the highest probability of being diagnosed at screening due to mammographic visibility, the proportion of cases correctly identified by the AI model was reduced from 35% to 22% (10/46).
About 20% of interval cancers have potential to be diagnosed earlier using AI in DBT screen-reading.
The AI-Augmented Scientific Congress Ecosystem (AISCE): Reimagining Scientific Congresses in the Age of Artificial Intelligence.
PubMed2026-07-01
Scientific congresses remain central to medical education, innovation, networking, and clinical consensus, but they are increasingly challenged by rising abstract volumes, hybrid formats, reviewer fatigue, fragmented programming, repeated speaker networks, and limited post-congress knowledge transfer. This editorial proposes the AI-Augmented Scientific Congress Ecosystem, a human-supervised model in which artificial intelligence supports the congress lifecycle rather than replacing scientific committees, reviewers, moderators, or societies. Near-term applications include abstract triage, reviewer matching, duplicate and similarity screening, program scheduling, hall allocation, attendee guidance, multilingual access, live transcription, question clustering, retrieval-grounded literature support, and post-congress synthesis. More advanced future applications include live session AI-generated debate prompts, speaker discovery, future topic prediction, congress knowledge graphs, consensus-draft generation, and consent-based digital legacy archives for medical educators. A key paradox is that congresses often contain unpublished clinical experience, early data, surgical insights, and expert debate that may precede the indexed literature on which many AI systems depend. Therefore, AI should help structure emerging knowledge, not replace expert interpretation. Ophthalmology is a suitable model field because it is visual, quantitative, technology-driven, and already engaged with AI in imaging, diagnostics, surgical planning, simulation, and education. Responsible implementation requires source grounding, auditability, privacy protection, bias monitoring, consent, conflict-of-interest governance, and final human decision-making.
Digital Pathology-Enabled Artificial Intelligence for Fibrosis Assessment in Metabolic Dysfunction-Associated Steatohepatitis: Needs, Current Progress, and Barriers.
PubMed2026-10-01
Metabolic dysfunction-associated steatotic liver disease (MASLD) is emerging as the most common chronic liver disorder worldwide. However, most of the candidate drugs, except resmetirom, have failed in the metabolic dysfunction-associated steatohepatitis (MASH) trials. The diagnosis of MASH relies on the assessment of key histological features: steatosis, ballooning, lobular inflammation, and fibrosis. Additionally, the current regulatory guidelines consider histological scores as the essential surrogate for the entry and efficacy endpoint decisions in phase IIb and phase III clinical drug trials. The reproducibility of histological scoring systems is increasingly being questioned in clinical trials. Inconsistencies in histopathological reporting, reflected by low inter-observer agreement, are largely attributable to the lack of standardized definitions, ordinal scoring structure with inadequate number of categories, use of variable criteria, ill-defined terminologies, and subjectivity in interpretation, which are further influenced by pre-analytical lab factors. These limitations result in high screen failure rates, poor risk stratification, increased placebo response, and reduced study power in drug trials, making this a serious concern. Liver pathologists are well aware of the urgent need to harmonize definitions and terminology and to transition from semi-quantitative to quantitative measurements using digital pathology (DP) integrated with computer vision and artificial intelligence (AI). Furthermore, standardized, objective, reproducible, and reliable histology-based ground truth is essential for the development and validation of non-invasive diagnostic tests for MASH. Several deep learning (DL) models and architectures, trained on stained or unstained digitized tissue slides using supervised and unsupervised approaches, have been developed, tested, and validated in MASH clinical studies. Liver fibrosis, the strongest primary prognostic determinant of clinical outcomes, has been the most extensively studied histologic feature in MASH. The AI models have demonstrated superior performance with respect to reproducibility and granularity, quantification of fibrosis progression and regression, improved confidence and concordance in Pathologists' scoring, standardization of consensus-based scoring systems, validation of AI-based tools to assist pathologists, and prediction of clinical outcomes. However, several challenges remain in the implementation of DP-enabled AI models in MASLD clinical practice and drug trials. These include ethical and legal considerations, model explainability, patient data privacy and confidentiality, financial investment, and regulatory approvals required for clinical adoption. Nevertheless, digital pathology-coupled DL models have the potential to augment current histopathological assessment in MASH.
Advancing antimicrobial resistance mitigation through artificial intelligence: A critical review and new perspectives on integrated monitoring and smart quaternary treatments of urban wastewater.
PubMed2026-08-03
Urban wastewater treatment plants (UWWTPs) play a key role in protecting environmental quality and public health by reducing pollutant discharges into receiving ecosystems. In this context, antimicrobial resistance (AMR) mitigation is emerging as a priority to limit the environmental dissemination of antibiotic resistant bacteria (ARB), antibiotic resistance genes (ARGs), and mobile genetic elements (MGEs). However, conventional UWWTP configurations are not designed to specifically control these biological determinants. Recent regulatory frameworks, within a One Health perspective, require monitoring of micropollutants that threaten water quality and public health, along with the progressive implementation of advanced treatment steps ("quaternary treatments") to remove them. Compliance with more stringent requirements, combined with the need for sustainable treatment processes, poses new challenges while creating opportunities for technological innovation. To support progress in the quaternary treatment of wastewater, the integration of advanced technologies with robust monitoring, modeling, and control systems is essential. This review critically examines recent advances in the design, management, and optimization of quaternary treatment processes through artificial intelligence (AI) algorithms, with specific emphasis on AMR mitigation. Quaternary treatments, including advanced oxidation processes (AOPs), membrane filtration, adsorption, and hybrid treatments, are discussed as promising barriers against AMR dissemination. AI models are identified as powerful tools for developing smart, adaptive, and efficient design, monitoring, and control systems for these technologies. Nevertheless, further research is needed to optimize model performance and practical feasibility in UWWTPs specifically oriented toward AMR mitigation. Key strengths, limitations, and research priorities are highlighted to support the development of next-generation smart wastewater treatment systems.
Clinician-facing artificial intelligence and guideline adherence: a structured narrative review of comparative and interventional evidence.
PubMed2026-12-01
To synthesize empirical evidence on whether clinician-facing artificial intelligence (AI) tools improve clinical practice guideline adherence and whether AI-generated recommendations are concordant with physician decision-making.
We conducted a structured narrative review aligned with SANRA and informed by narrative synthesis guidance. PubMed/MEDLINE, Embase, Scopus, Web of Science, and Google Scholar were searched for studies published from January 2020 to May 2026; earlier directly relevant studies were identified through citation tracking. Eligible studies evaluated clinician-facing AI tools, guideline adherence or concordance outcomes, or physician-versus-AI recommendations. Preprints were considered separately as emerging evidence and were not included in the peer-reviewed core synthesis.
Six peer-reviewed empirical studies met the core eligibility criteria. They included one real-world EHR-integrated pathway study, one NLP-based adherence measurement study, one randomized simulation trial of a decision-tree CDSS, and three AI-versus-physician comparative studies. Findings were most favorable for bounded tasks embedded in workflow or structured scenarios. Evidence was weaker for complex real-world decisions, and several studies did not evaluate patient outcomes. Three 2026 preprints were summarized separately as preliminary evidence.
Early evidence suggests that clinician-facing AI may support guideline-concordant care when applied to clearly defined decision tasks and integrated into clinical workflow. However, the evidence base remains small, heterogeneous, and largely indirect for diabetes care. Prospective cardiometabolic studies are needed to evaluate effectiveness, safety, usability, and patient outcomes.
The online version contains supplementary material available at 10.1007/s40200-026-02030-2.
Artificial Intelligence for the Diagnosis and Management of Neurodegenerative Diseases: A Comprehensive Review With an Emphasis on Parkinson's and Alzheimer's Diseases.
PubMed2026-07-01
Artificial intelligence (AI) is rapidly transforming research in neurodegenerative diseases, yet its clinical translation remains limited. We conducted a structured literature search across Google Scholar, PubMed, Scopus, and Web of Science, screening studies published between 2015 and April 2026 that applied machine learning (ML), deep learning (DL), and multimodal data integration to neuroimaging, biomarkers, and digital phenotyping. Our analysis revealed that AI models demonstrate strong potential for differentiating disease subtypes, predicting progression, and enhancing diagnostic accuracy, with notable advances in neuroimaging interpretation, fluid biomarker analysis, and wearable sensor data. In Parkinson's disease (PD), digital phenotyping through gait, speech, and handwriting analysis has enabled sensitive monitoring, while in Alzheimer's disease (AD), AI applied to imaging and plasma biomarkers has improved risk stratification. Despite these advances, barriers such as dataset heterogeneity, label noise, lack of external validation, and ethical concerns regarding bias, transparency, and patient trust persist. We conclude that while AI holds promise to revolutionize the care of PD and AD, real-world adoption requires multicenter validation, standardized reporting frameworks, regulatory guidance, and interdisciplinary collaboration, alongside prospective trials that embed AI tools into clinical workflows to ensure safety, equity, and effectiveness.
Transforming Cancer Treatment: Integrative Strategies Targeting the Tumor Microenvironment through Biological Innovation and Artificial Intelligence.
PubMed2026-08-08
Cancer is no longer viewed solely as a consequence of tumor-intrinsic genetic alterations but rather as a disease sustained by a complex and evolving tumor microenvironment (TME). The TME functions as an organized and dynamic ecosystem in which malignant cells interact continuously with immune populations, stromal elements, vascular networks, extracellular matrix components, and soluble mediators. These interactions critically regulate tumor initiation, progression, immune evasion, and metastatic dissemination. This review comprehensively examines the cellular and acellular architecture of the TME, emphasizing its spatial organization, metabolic reprogramming, mechanical properties, and immunological regulation across diverse tumor types. Key features of the TME include hypoxia-driven stabilization of hypoxia-inducible factors, oxidative stress-mediated immune dysfunction, metabolic competition for nutrients, extracellular matrix remodeling, and vascular abnormalities. Together, these interconnected processes establish immunosuppressive and therapy-resistant niches that promote angiogenesis, invasion, and metastatic spread. We further discuss how contemporary therapeutic strategies increasingly aim to exploit TME vulnerabilities, including immune checkpoint inhibition, adoptive cell therapies such as chimeric antigen receptor (CAR) T cells, antibody-based approaches, and rational combinatorial regimens. Emerging computational and artificial intelligence-driven frameworks are enhancing the integration of genomic, spatial, and clinical data to refine patient stratification and identify actionable microenvironmental targets. Despite substantial advances, significant challenges remain, including tumor heterogeneity, adaptive resistance mechanisms, and limited translation of preclinical findings into durable clinical benefit. Future progress will depend on integrating spatial systems biology, metabolomic and mechanobiological insights, and advanced human-relevant modeling platforms to enable precise and context-dependent TME modulation. A deeper understanding of tumor-microenvironment co-evolution is essential for the development of next-generation therapeutic strategies capable of achieving sustained clinical responses.
Perception of Medical Laboratory Professionals on the Role of Artificial Intelligence in Advancing Hematology Diagnostics.
PubMed2026-08-01
Artificial Intelligence (AI) has grown quickly in healthcare and has had a big effect on medical laboratory diagnostics, especially when it comes to diagnosing blood disorders. To figure out the pros and cons of using AI in hematology diagnostics, it is important to understand the perspective of medical laboratory workers instead of specialists toward AI use in hematology laboratory diagnostics. The goal of this study was to investigate what medical lab workers now know about AI and how they think it affects the accuracy of diagnostic hematology and patient outcomes.
Methods involved 113 participants, mostly laboratory technologists who filled out a standardized questionnaire as part of a quantitative exploratory research design. The data collection phase lasted 6 months, during which time-informed consent was sought and reminders were provided to boost response rates. Statistical tests were used including t-tests and regression analysis to determine the links between demographic factors and workers opinions on AI in hematological diagnoses.
The primary findings of this study dwell within the role of job skill and gender influence on "attitude", specifically, the study found that most of the participants believe that their professional attitude to adoption of AI in diagnoses could make their job more accurate and faster. However, there were several concerns about the quality of data, how easy it will be to understand the model, and the moral implications. There were positive relationships between knowledge, attitudes, and practices linked to AI. The secondary outcomes showed that most of the participants were young, with 45.1% being between the ages of 30 and 39. In addition, 65.5% of them had bachelor's degree, which suggests that they were comfortable with technology. Some of the participants had less than ten years of work experience, while others had more than ten years. Statistical tests demonstrated that demographic characteristics are quite important in how medical laboratory workers think about AI.
Medical laboratory workers believe that the use of AI in hematology diagnostics has its benefits, but they need more training and assistance to deal with their fears and create a space where human expertise and AI technologies can work together. By taking these views into account, healthcare organizations may better educate their staff for the changing role of AI in diagnostics, which will lead to improved patient outcomes and satisfaction.
Real or not real? Can radiologists distinguish artificial intelligence generated radiological images from real ones?
PubMed2026-07-16
Artificial intelligence (AI) models can create radiological images. We aimed to determine whether radiologists could distinguish AI-generated from real images, and factors associated with correct classification.
AI-generated images were made using an implementation of the Dreambooth fine-tuning approach applied to Stable Diffusion v2.1. Radiologists were asked to classify images as real (n = 10) or AI-generated (n = 20) and their confidence in this decision (1 least, 5 most) in an online form.
182 radiologists completed the survey. The median proportion of correctly identified images per respondent was 77.8% (interquartile range, IQR 70.0, 86.7%), with no difference between AI-generated (75.0%, IQR 70.5, 87.1%) and real images (83.4%, IQR 74.2, 92.6%, p = 0.19). Ultrasound and X-ray were more likely to be correctly identified than cross-sectional images like CT or MRI (88%, 91%, 70% and 77% respectively, p = 0.015). Mean confidence was similar for AI-generated and real images (3.50 ± 0.23 versus 3.56 ± 0.23, p = 0.49). There was no difference in classification based on number of years of experience (p = 0.57) or familiarity with AI (p = 0.37). However, radiologists with relevant specialist interests were more likely to correctly classify images (80.7 ± 1.3% versus 76.9 ± 0.8%, p = 0.012).
Radiologists were only able to correctly identify three-quarters of AI-generated images. This was impacted by sub-specialist expertise but not the number of years of experience or familiarity with AI.
Printed artificial intelligence-generated images for laser training.
PubMed2026-08-08
暂无摘要(点击查看详情)
Journal of the American Academy of Dermatology
查看原文 ↗Artificial Intelligence in Febrile Neutropenia: From Risk Scores to Real-Time Clinical Decision Support.
PubMed2026-07-01
Febrile neutropenia remains one of the most urgent and clinically challenging complications in oncology care, where timely recognition and decision-making can significantly influence outcomes. While traditional risk assessment tools have long supported clinicians in stratifying patients, they are limited in their ability to adapt to the complexity and rapidly changing nature of clinical situations. In this context, artificial intelligence (AI) is increasingly being explored as a way to complement existing approaches and support more responsive and individualized care. This editorial discusses how the role of AI in febrile neutropenia is evolving, moving from static risk scoring systems toward the possibility of real-time clinical decision support. It highlights both the potential of these tools to enhance clinical workflows and the practical challenges that still need to be addressed before widespread adoption. As these technologies continue to develop, careful integration into clinical practice may help improve how clinicians assess risk and manage patients with febrile neutropenia.
Conversational Artificial Intelligence and Neuropsychiatric Risk: A Narrative Review and Case-Based Synthesis Proposing a Delusional Feedback Loop.
PubMed2026-07-01
Conversational AI, powered by artificial intelligence, is becoming a common tool for accessing health information, educating patients, and obtaining general medical advice. These advanced systems, known as large language models, can produce responses that sound remarkably human. Nevertheless, these systems are prone to "AI confabulations," whereby they confidently generate incorrect information that could harm patients. This highlights the need to inform healthcare workers and individuals who may be prone to trusting these devices. New evidence suggests that AI may also exacerbate mental health conditions, particularly psychosis, paranoia, and related vulnerable states, especially among susceptible individuals. We conducted a targeted literature review and case-based analysis of 35 reported instances in which interactions with generative AI systems were temporally associated with the onset or worsening of psychotic symptoms. Across cases, recurrent patterns included reinforcement of delusional beliefs, amplification of pre-existing psychiatric vulnerabilities, promotion of harmful behaviors, and dissemination of unsafe medical guidance. Common contributing factors included prior psychiatric history, substance use, sleep disturbance, and prolonged AI engagement. We propose a conceptual hypothesis termed the delusional feedback loop, in which AI-generated responses iteratively validate distorted beliefs, contributing to their persistence and escalation. This process can be conceptualized as involving four components: underlying vulnerability, exposure to conversational AI, validation of distorted beliefs, and reinforcement through repeated interactions. Despite the rapid integration of conversational AI into health information seeking, there is currently no framework in the neuropsychiatric literature describing how AI interactions may relate to psychosis vulnerability. Existing reports are limited to isolated case descriptions without a common mechanism. This review addresses this gap.
Artificial intelligence for climate-health early warning systems in the Horn of Africa: opportunities, challenges, and a roadmap for action.
PubMed2026-08-08
Climate extremes, conflict, and population displacement converge in the Horn of Africa to accelerate outbreaks of climate-sensitive infectious diseases, whereas existing health surveillance systems remain fragmented and largely reactive. This Perspective examines the potential of artificial intelligence (AI) to strengthen climate-health early warning by integrating satellite earth observations, routine disease surveillance, and mobility-based vulnerability indicators into anticipatory decision support systems. Drawing on global experience and region-specific constraints, we identified critical barriers to implementation, including data fragmentation, infrastructure gaps, workforce shortages, governance silos, and unresolved ethical risks. We propose a five-layer conceptual framework for an AI-enabled Climate-Health Early Warning System (CHEWS) tailored to fragile and conflict-affected settings, alongside a phased regional policy roadmap anchored within the Intergovernmental Authority on Development (IGAD). Emphasizing data sovereignty, participatory governance, and privacy-by-design, this study positions AI-CHEWS as a feasible pathway for shifting the region from reactive outbreak responses to anticipatory public health actions that enhance climate resilience and equity.Clinical Trial Number: The authors declare that they have no competing interests.
The impact of artificial intelligence on critical thinking and clinical reasoning in health professions education: A systematic review and meta-analysis.
PubMed2026-08-04
Critical thinking and clinical reasoning underpin healthcare professionals' ability to navigate uncertainties and deliver safe and effective care. With artificial intelligence (AI) advancement and growing adoption, AI-based educational tools are increasingly used to support these cognitive competencies' development.
To synthesize randomised and controlled clinical trials on AI-based educational tools in health professions education and examine their effects on critical thinking and clinical reasoning among health professions students.
Six electronic databases were searched from January 1, 2014 to July 28, 2025 was reviewed: PubMed, Cochrane Central Register of Controlled Trials, CINAHL, Scopus, Embase and Web of Science. Two independent reviewers performed data extraction and quality assessment using standardized JBI checklists. The GRADE approach was used to assess the certainty of evidence. Studies were pooled via random-effects meta-analyses or narrative syntheses.
Fourteen randomised controlled trials and seven controlled clinical trials were included (n = 21). Meta-analyses revealed small to medium effect sizes for the surrogate clinical reasoning outcomes of performance-based assessment scores (SMD 0.68; 95% CI [0.38, 0.98], p-value = 0.00; I2 = 38%) and knowledge test scores (SMD 0.39; 95% CI [0.09, 0.69], p-value = 0.01; I2 = 79%). Critical thinking and clinical reasoning skills and dispositions were narratively synthesized, with majority of included studies favouring AI-based interventions but the evidence had low to very low certainty.
AI-based educational interventions may improve critical thinking and clinical reasoning among health profession students, but the evidence is very uncertain. This review offers preliminary insights but does not allow identification of optimal interventions or discipline-specific recommendations due to small sample sizes and substantial intervention heterogeneity. Further research is required to draw definitive conclusions.
CRD42025634074.
Comment on "Artificial Intelligence in Clinical Nutrition: A Narrative Review".
PubMed2026-08-08
暂无摘要(点击查看详情)
Clinical nutrition ESPEN
查看原文 ↗Artificial intelligence in gynecologic oncology: bridging the gap between digital innovation and clinical utility.
PubMed2026-09-01
暂无摘要(点击查看详情)
Current opinion in oncology
查看原文 ↗Evaluating the efficacy of artificial intelligence in audiology: a head-to-head comparison of ChatGPT and Gemini on hearing aid management.
PubMed2026-08-02
Hearing aid users frequently require accessible and immediate assistance for daily device management. This study aims to evaluate and compare the performance of two prominent Large Language Models (LLMs)-ChatGPT and Gemini, selected for their widespread public accessibility and market dominance-in providing accurate, comprehensible, and repeatable answers to frequently asked questions regarding hearing aids.
A comprehensive set of 44 user queries was divided into seven core categories. Responses generated by ChatGPT and Gemini were evaluated for comprehensibility and medical accuracy by three expert audiologists, utilizing official manufacturer manuals as a definitive gold standard. Inter-rater reliability was measured using the Intraclass Correlation Coefficient (ICC). Following the establishment of high consensus, repeatability was mathematically measured to objectively assess output similarity and eliminate human bias using a Natural Language Processing (NLP) algorithm (Cosine Similarity). Data were reported as medians and interquartile ranges (IQR), and pairwise statistical comparisons were conducted using the non-parametric Wilcoxon matched-pairs signed-rank test.
Inter-rater reliability was strong for both subjective domains (Accuracy ICC = 0.74; Clarity ICC = 0.74, p < 0.001). Both models demonstrated excellent performance in clarity (Median: 5.00, IQR: 0.00), with no statistically significant differences observed across any categories (p > 0.05). Regarding medical and technical accuracy, rank-based analyses revealed that Gemini showed significantly superior overall performance (p < 0.001), specifically excelling in the "Pairing" (p = 0.03) and "Before Using a Hearing Aid" (p = 0.03) categories. In the objective assessment of repeatability via NLP algorithms, both models exhibited high semantic consistency (ChatGPT Median: 0.91; Gemini Median: 0.93), with no statistically significant differences observed between them in any category (overall p = 0.30).
Both ChatGPT and Gemini generate highly comprehensible information for hearing aid users (Clarity Median: 5.00). However, their performance fluctuates depending on task complexity; while Gemini offers superior technical accuracy (p < 0.001), both maintain high but imperfect repeatability (Cosine Similarity > 0.90). Because neither model demonstrates complete diagnostic stability or perfect reproducibility across all clinical domains, they cannot currently be recommended as independent digital assistants. While LLMs show great promise as supplementary tools for patient education, professional audiological supervision remains essential to verify clinical accuracy and ensure patient safety.
Can artificial intelligence accurately assess systematic review quality? Benchmarking large language models for AMSTAR 2 appraisal in dental evidence synthesis.
PubMed2026-08-08
To evaluate the accuracy and reliability of three AI platforms ChatGPT, Perplexity, and Google Gemini in assessing the methodological quality of systematic reviews using the AMSTAR 2 checklist, compared with expert manual evaluation in dental research.
A cross-sectional comparative study was conducted to assess the performance of three AI platforms ChatGPT, Perplexity, and Google Gemini in evaluating the methodological quality of 35 systematic reviews using the AMSTAR 2 checklist. Manual assessments by a domain expert served as the reference standard. Each AI system was prompted with a standardized AMSTAR 2 query, and item-level outputs were collected for direct comparison. Key metrics included percentage agreement, error proportions, and inter-rater reliability measured by Cohen's kappa. Error proportions represent the proportion of discordant assessments out of total valid pairwise comparisons across 16 AMSTAR-2 items. Differences between LLM-generated and reference AMSTAR-2 ratings were summarized using effect estimates with corresponding 95% confidence intervals. Comparative performance across platforms was assessed based on confidence-interval overlap rather than hypothesis testing. Results are presented as effect estimates with corresponding 95% confidence intervals, without hypothesis testing or statistical dichotomization. This approach provided a robust and reproducible framework to benchmark AI-assisted quality appraisal in dental evidence synthesis.
Among 35 systematic reviews assessed, Perplexity demonstrated the highest agreement with expert AMSTAR-2 ratings (error proportion: 19.0%; weighted κ_w: 0.78, 95% CI 0.71-0.85), followed by ChatGPT (error proportion: 22.9%; weighted κ_w: 0.62, 95% CI 0.54-0.70) and Google Gemini (error proportion: 43.9%; weighted κ_w: 0.41, 95% CI 0.33-0.49). Perplexity also achieved the best sensitivity (81.3%, 95% CI 76.5-85.4%) and specificity (82.7%, 95% CI 78.1-86.5%) for correctly identifying high-quality reviews. Non-overlapping 95% confidence intervals suggest meaningful differences in performances among platforms, with Perplexity showing superior agreement across all metrics. Across all platforms, agreement was generally higher for non-critical AMSTAR-2 domains involving clear and structured reporting, whereas performance was weaker for critical domains requiring interpretation of complex methodological details, risk-of-bias considerations, and evidence synthesis procedures.
Perplexity demonstrated the highest accuracy and agreement with expert assessments of the methodological quality of systematic reviews, suggesting its potential as a supportive AI tool for AMSTAR-2-based appraisal in dental evidence synthesis. In contrast, systematic biases observed in ChatGPT and Google Gemini underscore the continued need for human oversight to ensure the validity of methodological assessments. Differences in agreement and error proportions were observed across all models when compared with expert AMSTAR-2 evaluations, indicating meaningful variability in methodological appraisal performance, reinforcing that AI-assisted appraisal of systematic review methodology should complement rather than replace expert human judgment in dental research.