Natural language processing (NLP) techniques offer promising solutions for semi-automating the time-consuming process of abstract screening in systematic reviews. The exponential growth of published literature has created significant bottlenecks, with review teams manually assessing thousands of abstracts over weeks to months. Single reviewers can miss 5-13% of relevant studies, necessitating dual screening that further increases workload. Advances in artificial intelligence, including deep learning models such as BERT and its successors, show potential for automating this critical step, but comprehensive evidence on optimal approaches, performance, and practical feasibility remains limited. This systematic review aimed to assess techniques, performance, and feasibility of NLP approaches for title and abstract screening by characterizing the range of NLP methods used, summarizing performance on key metrics like workload reduction and recall, evaluating real-world implementation feasibility, and identifying research gaps and future directions. We searched PubMed, Web of Science, Embase, CINAHL, The Cochrane Library, Scopus, and gray literature sources from inception to December 2024. The search strategy, developed with an information specialist and peer-reviewed using PRESS guidelines, targeted keywords related to natural language processing, machine learning, abstract screening, and systematic reviews. Additional sources included conference proceedings, preprint servers, reference lists, forward citation tracking, and expert consultation. We included primary studies of any design describing development or evaluation of NLP techniques for automating title and abstract screening in evidence syntheses. Eligible studies reported on NLP methods, screening performance (workload reduction, recall, precision), or implementation feasibility. Studies using only rule-based approaches without machine learning, systematic reviews of NLP methods, commentaries, and conference abstracts were excluded. No language or date restrictions were applied. Two reviewers independently screened titles, abstracts, and full texts using Covidence software, with disagreements resolved through discussion. Data extraction covered study characteristics, NLP techniques, training approaches, performance metrics, and feasibility considerations. Risk of bias was assessed using a modified ROBIS tool. Given diverse techniques and outcomes, we conducted narrative synthesis following SWiM guidelines, grouping studies by NLP approach. From 4,105 records, 19 studies met inclusion criteria, with 68.4% published since 2023, reflecting rapid field advancement. Studies employed diverse approaches from traditional machine learning (Support Vector Machines, Random Forests) to advanced deep learning models, particularly BERT variants. Most achieved >90% recall with workload reductions of 13-96%, representing substantial time savings. Deep learning models with transfer learning consistently outperformed traditional approaches. However, implementation faced significant barriers including requirements for high-quality training data, specialized computational resources, technical expertise, and user-friendly interfaces. Performance was generally better for targeted reviews with lower inclusion prevalence. NLP techniques, especially deep learning with transfer learning, show substantial promise for semi-automating abstract screening with potential for large workload savings while maintaining high recall. However, challenges remain regarding training data quality, computational requirements, technical expertise needs, and user-centered design. Realizing full potential requires interdisciplinary collaboration to develop reliable, generalizable tools integrating seamlessly with human expertise and existing workflows. Future priorities include creating standardized datasets, conducting prospective evaluations, developing user-friendly interfaces, and establishing implementation best practices to revolutionize evidence synthesis efficiency. Declarative title: Natural language processing substantially reduces abstract screening workload but requires significant expertise and specialized computing resources. The review in brief Natural language processing techniques can reduce manual abstract screening workload by 30-90% while maintaining over 90% recall of relevant studies, with deep learning models showing the greatest promise for systematic review automation. What is this review about? Problem statement: Systematic reviews are the highest level of evidence for policy and practice, but exponential literature growth has made abstract screening a major bottleneck, taking weeks to months. Single reviewers miss 5-13% of relevant studies, making dual independent screening the gold standard, at significant cost in time and resources. What is the aim of this review? This systematic review examines the techniques, performance, and feasibility of natural language processing methods for automating title and abstract screening in systematic reviews. What are the main findings of this review? What studies are included? This review includes 19 studies that evaluated NLP techniques for automating abstract screening in systematic reviews, rapid reviews, scoping reviews, and other evidence syntheses. Studies came from diverse global regions, with 68.4% published in 2023 or later. Studies varied in design and corpus size (hundreds to tens of thousands of articles), mostly from medical literature, and demonstrated generally good methodological quality. Do NLP techniques reduce workload while maintaining accuracy? NLP techniques consistently demonstrate substantial workload reductions while maintaining high recall of relevant studies. Workload reductions range from 13-96%, with most studies achieving reductions of 30-60%. Most studies achieve recall rates exceeding 90%, meaning they successfully identify over 90% of relevant articles. Precision varied widely (10-99%), primarily affecting the volume of articles requiring manual review rather than the risk of missing relevant studies. Which NLP approaches perform best? Deep learning models, particularly those leveraging transfer learning with large pretrained language models like BERT and its variants (BioBERT, PubMedBERT), consistently outperform traditional machine learning approaches. Traditional approaches such as Support Vector Machines perform well but are generally outperformed by these modern architectures. What factors affect implementation feasibility? Several key factors influence the practical implementation of NLP systems. High-quality training data and specialist technical expertise are prerequisites. Computational resources vary from standard computing for simpler models to specialized GPU infrastructure for advanced deep learning approaches. User-friendly interfaces and domain generalizability remain key challenges for broader adoption. What do the findings of this review mean? NLP offers substantial promise for reducing the abstract screening burden. Workload reductions of 30-90% could accelerate evidence synthesis and policy translation. However, successful implementation requires careful planning, technical expertise, and adequate computational resources. Standardized datasets, user-friendly tools, and clearer best-practice guidance are needed to realise this potential. How up-to-date is this review? The review authors searched for studies up to December 2024.
使用 AI 将内容摘要翻译为中文,便于快速阅读
使用 AI 分析这篇文章的核心发现、关键要点和深度见解
由 DeepSeek AI 提供分析 · 首次使用需配置 API Key
PubMed · 2026-09-01
PubMed · 2026-01-01
PubMed · 2026-06-01
PubMed · 2026-07-21