← All verdicts

Do systematic reviews still matter in the age of agentic AI?

An answer on demand is a compelling promise. A trustworthy synthesis asks more of us.

Ask a clinical question. Let a group of AI agents search the literature, extract the findings, run an analysis and return an answer. If that becomes routine, why spend months producing a systematic review and meta-analysis?

It is a fair question. I think AI could shorten much of that work. But I still see a strong place for carefully conducted, published reviews. My most important reason is the scrutiny that happens beyond the team producing the answer.

First, the value of independent peer review

An answer generated for one person can arrive without anyone outside that workflow challenging it. A review submitted to a journal faces another layer of judgment: editors and external reviewers can question the search, exclusions, analysis and interpretation, then ask for revisions.

This generally happens before publication. After publication, the wider community can examine, challenge and build on the work. Cochrane’s process, for example, brings in methodological, search and clinical expertise, with lived-experience perspectives where appropriate. Other journals vary. See Cochrane’s editorial and peer-review process.

Peer review is imperfect; publication does not certify that a conclusion is correct. Still, I place considerable value on having the methods and claims questioned by people who did not generate them. An AI-assisted review can benefit from exactly this process. An on-demand answer should not be assumed to have undergone it.

The case for on-demand synthesis is strong

By agentic AI, I mean systems that coordinate several steps and tools: planning searches, retrieving records, helping screen studies, extracting data and running code. Used well, these systems could reduce repetitive work, support faster updates and make evidence easier to explore.

I would welcome that. A carefully validated workflow might also make it easier to test alternative assumptions rather than stop at a single analysis. The important question is whether each step works reliably for the review in front of us.

Cochrane’s June 2026 guidance allows AI with responsibility, transparency and human oversight, while calling for verification or validation of generative tools. That seems a sensible direction to me: evaluate the tool and its use, rather than accept or reject it simply because it is AI. Cochrane’s guidance on selecting AI tools.

The difficult part is deciding what belongs together

A systematic review is a structured process for finding, assessing and synthesizing evidence. A meta-analysis is the statistical combination of results, when combining them makes sense. A good review may conclude that pooling would be misleading.

Imagine two studies of the same treatment: one in newly diagnosed patients, another after several previous therapies. Their outcome labels may look similar, but their populations, follow-up and clinical questions may differ. We also need to recognize overlapping cohorts, missing outcomes and risks of bias. A random-effects model does not make those differences disappear.

AI can help reason through these decisions. I am not yet comfortable treating its reasoning as sufficient without expert checking. The challenge is to justify the analysis, including the decision not to pool. The Cochrane Handbook’s chapter on meta-analysis explains why these choices matter.

Finding a paper is not the same as having access

Search results may expose a title or abstract without the full text, supplementary tables or details needed for extraction. Institutional subscriptions and licensed services can help; so can open repositories. AI services can obtain lawful access too. But access and permission for automated reuse must be established, not assumed. Even within PubMed Central, availability for text mining depends on the article and its terms. PMC’s guidance on article datasets and reuse makes that distinction clear.

Databases also have operational limits. NCBI’s E-utilities, for example, ordinarily allow three requests per second without an API key and ten with one; batching and approved higher limits can help. These are manageable engineering constraints, but they complicate the promise of unlimited, instant retrieval. They affect researchers and automated systems alike. NCBI’s API guidance describes the available approaches.

Several agents do not automatically make an independent team

The value of a review team is not just its size. A clinician, statistician, information specialist and another independent reviewer can notice different problems and challenge one another’s assumptions.

I would apply the same test to an agent team: are the checks genuinely independent? Agents using the same model, inputs or earlier conclusions may repeat the same mistake. Their agreement alone does not establish correctness.

Cochrane recommends independent extraction of outcome data by at least two people, with a planned way to resolve disagreements. It also notes that extraction errors are rarely caught by journal reviewers or editors. That is a useful reminder: external peer review adds value, but cannot replace sound work inside the team. Cochrane Handbook: collecting data.

Speed still needs an audit trail

Searching, parsing documents, running models and verifying outputs all consume resources. Repeating the process for every question may or may not be economical; that depends on the system, its reuse of earlier work and the depth of checking required.

I would want a synthesis to preserve its question, protocol, search dates, selection decisions, extracted data and analysis code. If AI is involved, its tools, versions and role should be recorded too. Someone else should be able to understand how the conclusion was reached and where it might fail.

PRISMA 2020 provides a framework for transparent reporting, including automation and access to supporting materials. It is not a certificate of methodological quality. Making the work inspectable is necessary; the decisions still have to be defensible.

My verdict

Systematic reviews and meta-analyses still have an important role in the age of agentic AI. I expect their workflows to change, and I would welcome less time spent on repetitive tasks. But an answer arriving quickly does not, by itself, establish that the evidence has been handled well.

For consequential decisions, I still want expert supervision, justified methods, traceable evidence and independent scrutiny. Publishing a carefully reviewed synthesis gives the wider community something stable to examine, question and improve.

Future AI systems may meet more of these requirements. I would judge them by demonstrated performance, rather than a prediction about AGI. For now, my preference is to use agents to strengthen rigorous evidence synthesis while preserving the people and processes that hold it accountable.

← Back to all verdicts