Abstract
OBJECTIVES: The application of large language models (LLMs) to systematic review tasks is rapidly expanding, yet the transparency and methodological rigor of these evaluations remain unclear. We aimed to assess reporting transparency, methodological quality, and how authors frame claims and caveats in studies applying LLMs to systematic review tasks.
STUDY DESIGN AND SETTING: We conducted a cross-sectional meta-epidemiological study by searching PubMed, Embase, Web of Science Core Collection, IEEE Xplore, and five other databases from inception to December 1, 2025, for peer-reviewed articles and preprints. We included empirical studies evaluating generative transformer-based LLMs (eg, Gemini) for core systematic review tasks (eg, screening, data extraction) against a reference standard. We assessed reporting transparency using an adapted Chatbot Assessment Reporting Tool and methodological quality using an adapted Quality Assessment of Diagnostic Accuracy Studies 2 tool. We also analyzed the frequency and strength of claims and caveats mentioned by the authors. The study is registered with the Open Science Framework (https://osf.io/8edhb).
RESULTS: We identified and included 229 studies comprising 440 empirical tasks. Reporting transparency was moderate, with a mean item score of 0.52 (standard deviation 0.30) on a 0-1 scale, where higher values indicate more complete reporting. We observed substantial gaps in reproducibility-essential domains, including protocol information (mean score 0.12) and model details (0.30). Although 60.6% of assessments were rated as having a low risk, key safeguards against overfitting and data leakage were rarely reported; for example, locking the test set before prompt optimization, a basic protection against information leakage, was not reported in 99.8% of tasks. We identified 837 claims and 693 caveats. Authors framed claims weakly more often than strongly (66.8% vs 33.2%). Performance superiority over a comparator was the most common claim (64.8% of tasks). Readiness for practical use was claimed in 47.5% of tasks, almost always in qualified terms (93.3%).
CONCLUSION: Studies applying LLMs to systematic review tasks are reported with moderate transparency but often omit reproducibility-critical details necessary to assess leakage and overfitting. Although authors frequently make claims about performance and practice readiness, these are typically expressed cautiously. Improved reporting standards and clearer safeguards are urgently needed before routine use of LLMs in evidence synthesis can be recommended.
PLAIN LANGUAGE SUMMARY: LLMs, such as ChatGPT, are increasingly used to help carry out parts of systematic reviews, which summarize evidence to inform healthcare decisions. We examined 229 studies that tested LLMs on tasks such as screening articles, extracting data, and assessing study quality, covering 440 evaluations in total. On average, these studies reported their methods with moderate clarity, but often omitted information needed to repeat the work or judge whether the results were trustworthy. Notably, almost none described safeguards to ensure that the test data had not already influenced how the model was set up, a key step for avoiding overly optimistic results. Authors frequently described LLMs as performing well and nearly ready for practical use, though usually in cautious terms. Clearer reporting standards and stronger safeguards are needed before LLMs can be routinely relied upon in evidence synthesis.
| Original language | English |
|---|---|
| Article number | 112383 |
| Number of pages | 13 |
| Journal | Journal of Clinical Epidemiology |
| Volume | 197 |
| Early online date | 17 Jun 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 17 Jun 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- Evidence synthesis
- Large language models
- Meta-epidemiological study
- Methodological quality
- Reporting transparency
- Systematic review
Fingerprint
Dive into the research topics of 'Large language models for systematic reviews were reported to perform well but rarely with verifiable safeguards: a cross-sectional study'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver