Skip to main navigation Skip to search Skip to main content

Large language models for systematic reviews were reported to perform well but rarely with verifiable safeguards: a cross-sectional study

  • Honghao Lai
  • , Bernardo Sousa-Pinto
  • , Christian Cao
  • , David Moher
  • , Janne Estill
  • , Jiayi Liu
  • , Weilong Zhao
  • , Yutong Wang
  • , Ziying Ye
  • , Bo Tong
  • , Zhenhua Yang
  • , Xufei Luo
  • , Bingyi Wang
  • , Yimeng Li
  • , Bei Pan
  • , Lu Zhang
  • , Jinhui Tian
  • , Yaolong Chen
  • , Nannan Shi
  • , Long Ge*
  • *Corresponding author for this work

Research output: Contribution to journalJournal articlepeer-review

Abstract

OBJECTIVES: The application of large language models (LLMs) to systematic review tasks is rapidly expanding, yet the transparency and methodological rigor of these evaluations remain unclear. We aimed to assess reporting transparency, methodological quality, and how authors frame claims and caveats in studies applying LLMs to systematic review tasks.

STUDY DESIGN AND SETTING: We conducted a cross-sectional meta-epidemiological study by searching PubMed, Embase, Web of Science Core Collection, IEEE Xplore, and five other databases from inception to December 1, 2025, for peer-reviewed articles and preprints. We included empirical studies evaluating generative transformer-based LLMs (eg, Gemini) for core systematic review tasks (eg, screening, data extraction) against a reference standard. We assessed reporting transparency using an adapted Chatbot Assessment Reporting Tool and methodological quality using an adapted Quality Assessment of Diagnostic Accuracy Studies 2 tool. We also analyzed the frequency and strength of claims and caveats mentioned by the authors. The study is registered with the Open Science Framework (https://osf.io/8edhb).

RESULTS: We identified and included 229 studies comprising 440 empirical tasks. Reporting transparency was moderate, with a mean item score of 0.52 (standard deviation 0.30) on a 0-1 scale, where higher values indicate more complete reporting. We observed substantial gaps in reproducibility-essential domains, including protocol information (mean score 0.12) and model details (0.30). Although 60.6% of assessments were rated as having a low risk, key safeguards against overfitting and data leakage were rarely reported; for example, locking the test set before prompt optimization, a basic protection against information leakage, was not reported in 99.8% of tasks. We identified 837 claims and 693 caveats. Authors framed claims weakly more often than strongly (66.8% vs 33.2%). Performance superiority over a comparator was the most common claim (64.8% of tasks). Readiness for practical use was claimed in 47.5% of tasks, almost always in qualified terms (93.3%).

CONCLUSION: Studies applying LLMs to systematic review tasks are reported with moderate transparency but often omit reproducibility-critical details necessary to assess leakage and overfitting. Although authors frequently make claims about performance and practice readiness, these are typically expressed cautiously. Improved reporting standards and clearer safeguards are urgently needed before routine use of LLMs in evidence synthesis can be recommended.

PLAIN LANGUAGE SUMMARY: LLMs, such as ChatGPT, are increasingly used to help carry out parts of systematic reviews, which summarize evidence to inform healthcare decisions. We examined 229 studies that tested LLMs on tasks such as screening articles, extracting data, and assessing study quality, covering 440 evaluations in total. On average, these studies reported their methods with moderate clarity, but often omitted information needed to repeat the work or judge whether the results were trustworthy. Notably, almost none described safeguards to ensure that the test data had not already influenced how the model was set up, a key step for avoiding overly optimistic results. Authors frequently described LLMs as performing well and nearly ready for practical use, though usually in cautious terms. Clearer reporting standards and stronger safeguards are needed before LLMs can be routinely relied upon in evidence synthesis.

Original languageEnglish
Article number112383
Number of pages13
JournalJournal of Clinical Epidemiology
Volume197
Early online date17 Jun 2026
DOIs
Publication statusE-pub ahead of print - 17 Jun 2026

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure

User-Defined Keywords

  • Evidence synthesis
  • Large language models
  • Meta-epidemiological study
  • Methodological quality
  • Reporting transparency
  • Systematic review

Fingerprint

Dive into the research topics of 'Large language models for systematic reviews were reported to perform well but rarely with verifiable safeguards: a cross-sectional study'. Together they form a unique fingerprint.

Cite this