- AI reduces the mechanical work of a meta-analysis: searching, screening, extraction, and formatting. It does not replace methodological judgment. Pooled estimates should always be computed in statistical software rather than inside a chat window.
- Hallucination is the central risk. Fabricated citations, altered numbers, and false consensus are all documented issues. Prevent these through source anchoring, dual extraction, and verification of every DOI.
- Reporting standards are under development. PRISMA-trAIce covers what to disclose when AI is used as a tool in a meta-analysis. The RAISE recommendations, supported by Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence, set expectations for responsible AI use.
Introduction
A traditional meta-analysis can take 12 months from protocol to publication, which means many are already out of date when they appear, and very few are ever updated. The gap between how fast evidence accumulates and how slowly it is synthesized is the problem that researchers are trying to solve using AI. For example, using an AI-based approach, researchers generated usable clinical evidence about ocular toxicity from hydroxychloroquine in <30 minutes.
This guide covers the full workflow of a meta-analysis where AI is used, for tasks like framing the question, searching, screening, extraction, appraisal, narrative synthesis, statistical pooling, and manuscript writing. It includes sample prompts, a hallucination prevention protocol, a comparison of free and paid tools, and the reporting standards that now govern all of it.
Step 1: Framing the Research Question and Inclusion Criteria for a Meta-Analysis
A poorly scoped research question produces a search that is either unmanageably broad or misses the evidence it needed to find.
- Translate the topic into a formal framework: PICO for interventions, PECO for exposures, SPIDER for qualitative questions.
- Check whether the question has already been answered: duplicating a recent Cochrane review is a common and avoidable error.
- Pressure-test poolability before committing: if outcome measures, populations, and designs vary too widely, plan a narrative synthesis from the start.
- Draft inclusion and exclusion criteria that 2 reviewers can apply the same way, and pilot them on 20 records before scaling.
- Decide and justify any AI use now, at protocol stage, rather than retrofitting a justification later.
- Register the protocol on PROSPERO, including the planned AI use.
Prompt for question framing:
“Convert this research topic into a PICO framework, then list 3 narrower and 3 broader variants of the question. Topic: [topic].”
Prompt for viability testing:
“I am planning a meta-analysis on [question]. List the most likely reasons this question would NOT be poolable: expected heterogeneity in outcome measures, populations, and study designs. Be specific and skeptical.”
Prompt for criteria drafting:
“Draft inclusion and exclusion criteria for a meta-analysis on [question], covering population, intervention, comparator, outcomes, study design, language, and publication window. Flag any criterion that would be hard to apply consistently across 2 reviewers.”
Step 2: AI-Assisted Literature Search and Screening
AI-assisted literature search and conventional Boolean search are quite different. AI-based search is “semantic” (i.e., based on meaning of free text) and is better at finding conceptually related work that you didn’t know the exact terms or keywords to search for. But Boolean search is what you can document, replicate, and defend in a PRISMA flow diagram. Your core search strategy in a meta-analysis should be the Boolean one, and you can use AI to support this search.
How to use AI to support a Boolean search:
- Ask AI to generate a search string, then manually check every Boolean operator, truncation, wildcard, and controlled vocabulary term/keyword.
- Feed a few seed papers into a tool like Elicit or R Discovery to surface synonyms, MeSH terms, and concepts your query would have missed and then fold those into your Boolean search strings.
- Run the search across at least 3 databases: typical combinations include PubMed, Embase, Scopus, Web of Science, and PsycINFO.
- Search grey literature and trial registries separately: AI retrieval does not fix a narrow database list.
- Record the exact strings, databases, filters, and dates of every search, because these go into the manuscript verbatim.
- Use browser extensions and reference-manager integrations so records are captured as you search rather than reconstructed later.
How to use an AI literature search tool for meta-analyses
Choose an AI tool for your meta-analysis that meets the following criteria:
- Strictly searches scientific databases (PubMed, Web of Science, etc.) and will turn up a blank result for a paper that doesn’t exist. This reduces the risk of hallucinated papers.
- Search results can be exported into a shareable file or link.
- The tool is integrated with conventional reference managers like Zotero and Mendeley.
At this stage, never use a general purpose LLM like ChatGPT, because these have the highest risk of hallucinations (they produce citations and references that look real but are not).
Run your semantic (i.e., free-text) query in your chosen AI tool. This is usually your research question, written as a full sentence. Save the output each time. Treat every paper AI has returned as candidates only. Compare them against the results of your Boolean search and include only papers that the Boolean search missed. Screen them (title first, then abstract, then full text) and record them separately for your PRISMA flow diagram.
How to use AI for literature screening
Prompt for abstract screening:
“Here are my inclusion criteria: [paste]. For each of the following 20 abstracts, output a table with columns: Study ID, Include or Exclude or Unclear, Criterion triggered, One-line justification. Do not guess when the abstract is silent; mark Unclear.”
- Require a stated reason tied to a specific criterion for every exclusion.
- Keep a documented second-reviewer or spot-check protocol: AI screening does not remove the need for human verification.
- Track AI-screened and human-screened counts separately from the start, because your flow diagram will need them.
AI-assisted data extraction
Prompt for source-anchored extraction:
“Extract the following fields from this paper into a single table row, and write NOT REPORTED where the paper does not state a value. Never estimate. Fields: [list]. After the table, quote the exact sentence or table label you took each number from.”
Extraction is where AI can do the most damage, because hallucinated numbers reach the pooled estimate without ever looking wrong. Here are some conversions by AI that you must check if you’re using AI for data extraction:
- Standard error to standard deviation and back, using the reported sample size.
- Odds ratio to risk ratio, given baseline risk.
- Median and interquartile range to approximate mean and standard deviation, flagged as an approximation in the manuscript.
- Log-scale to natural-scale reporting for ratio measures.
Step 3: Quality Appraisal and Risk of Bias
Here’s what AI can do at the quality appraisal stage:
| Study design | Appraisal tool | What AI can help with |
| Randomized trials | RoB 2 | Locate and quote sentences covering randomization, blinding, missing data, and selective reporting |
| Non-randomized interventions | ROBINS-I | Identify stated confounders and adjustment strategy |
| Observational cohort and case-control | Newcastle-Ottawa Scale | Extract selection, comparability, and outcome ascertainment details |
| Certainty of the evidence | GRADE | Draft domain-level justifications for human rating |
- What AI does well here: finding the relevant sentences fast, drafting domain justifications, and flagging where a paper lacks information.
- What AI does badly: assigning the final rating, and judging whether a protocol deviation matters clinically.
- Use AI as a pre-reader that prepares evidence for the 2 human reviewers, not as one of the reviewers.
- Appraisal is explicitly a judgment-making stage, so AI use here must be declared under both RAISE and most publisher policies.
Prompt for risk-of-bias preparation:
“Using RoB 2 domains, locate and quote the sentences in this trial report that address randomization, deviations from intended interventions, missing outcome data, outcome measurement, and selective reporting. For each domain, only quote exact sentences you find and say “not found” if you can’t find those sentences. Do not assign a risk judgment.”
Step 4: Narrative Synthesis When You Cannot Pool
Choosing not to pool is normal and often advisable. Forcing incomparable studies into one estimate produces a number that looks authoritative and means nothing.
Signals that narrative synthesis is the correct choice:
- Fewer than 3 studies report the same outcome in a compatible format.
- Outcome measures differ so much that a standardized mean difference would not be interpretable.
- Heterogeneity is extreme, and subgroup analysis cannot explain it.
- Study designs differ fundamentally, for example mixing randomized trials with uncontrolled before-and-after studies.
Where AI genuinely helps in narrative synthesis:
- Clustering findings into candidate themes across dozens of results sections.
- Building effect-direction tables and vote-counting matrices.
- Drafting SWiM-compliant text for a literature synthesis that you then verify and rewrite.
- In meta-synthesis, coding qualitative themes and surfacing disconfirming cases you may have skimmed past.
The specific risk here is that language models are trained to produce coherent prose. But this can result in conflicts in the literature being glossed over or “smoothened”. Always instruct AI to preserve contradiction explicitly.
Step 5: Statistical Pooling and Interpretation
This is the stage with the hardest boundary in the entire workflow. AI can write the code that runs the analysis. AI must never produce the numbers by reasoning over them in a chat window. All calculations must be done on conventional software like R (metafor and meta packages) or Rev Man.
Decisions the researcher must make before touching any tool:
- Fixed-effect or random-effects, decided in the protocol and justified, not chosen after seeing the results.
- Effect measure by data type: standardized or raw mean difference for continuous outcomes; odds ratio, risk ratio, or hazard ratio for binary and time-to-event outcomes.
- Which subgroup analyses and meta-regressions are pre-specified, and which will be labeled exploratory.
- How publication bias will be assessed, and at what minimum number of studies those tests become meaningful.
What AI can do at this stage:
- Writing and commenting R code using metafor, meta, or dmetar, and the Python or Stata equivalents.
- Debugging code that runs but produces implausible output.
- Auditing extraction spreadsheets for formula errors, inconsistent ranges, and standard errors entered as standard deviations.
- Translating model output into plain language so you can sanity-check your own interpretation.
Prompt for analysis code:
“Write R code using the metafor package to run a random-effects meta-analysis on this dataset: [paste CSV]. Use REML, report I-squared, tau-squared, and the prediction interval, and generate a forest plot with study weights shown. Comment every line.”
Prompt for output interpretation:
“I ran this model and got these results: [paste]. Explain what an I-squared of 78% and this prediction interval imply about whether a single pooled estimate is meaningful here.”
Prompt for spreadsheet auditing:
“Audit this extraction sheet for formula errors, inconsistent ranges, and cells where a standard deviation may have been entered as a standard error.”
Step 6: Manuscript Writing and Reporting
Where AI genuinely helps at the writing stage:
- Outlining sections against PRISMA 2020 and checking coverage item by item.
- Tightening lengthy paragraphs
- Converting statistical output into readable results text in past tense.
- Compressing an abstract to a strict word limit without dropping reported numbers.
- Drafting plain-language summaries for policy and patient audiences.
Where it does damage:
- Writing the discussion section, where it may exaggerate claims or make logical gaps.
- Generating limitations you have not actually verified against your own data.
- Producing citations from memory rather than from your reference manager.
- Rewriting results text so fluently that a hallucinated number stops looking wrong.
Sample prompt for results text:
“Convert this metafor output into a PRISMA-compliant results paragraph. Report the pooled estimate, confidence interval, I-squared, prediction interval, and the number of studies and participants. Use past tense and no interpretive language.”
AI Hallucinations in Meta-Analysis and How to Prevent Them
The joint position statement from Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence describes AI as characterized by opaque decision making and black-box predictions, susceptible to overfitting, potentially embedded with algorithmic biases, and at risk of fabricated outputs and hallucinations. This is not a fringe concern.
What Hallucination Looks Like in Evidence Synthesis
| Issue | What it looks like | How to catch it |
| Fabricated citations | Plausible authors, journal, year, and DOI that resolve to nothing or to a different paper | Resolve every DOI and PMID rather than eyeballing plausibility |
| Altered numbers | A sample size of 248 becomes a subgroup sample size as well; a confidence interval quietly widens | Require a quoted source sentence and locator for every value |
| Misattributed findings | A real study credited with a conclusion it never drew | Read the full text of every included study |
| Phantom extraction | Confident values drawn from a section that does not exist in the PDF | Instruct the model to output NOT REPORTED and permit gaps |
| False consensus | Genuinely conflicting literature smoothed into one tidy claim | Ask explicitly for contradictions and disconfirming cases |
| Reference drift | Citations mutate during redrafting, especially across long sessions | Re-verify the reference list against sources before submission |
How to reduce AI hallucinations in a meta-analysis
- Use tools that read your uploaded PDFs or a real indexed corpus. Never ask a general chatbot to find studies from memory.
- Demand source anchoring from your AI tool. Every extracted value needs a quoted sentence and a page or table locator.
- Permit gaps explicitly. Authorize the model to write NOT REPORTED, because most fabrication is gap-filling under pressure.
- Verify every identifier. Resolve DOIs and PMIDs; a well-formed DOI is not evidence that the paper exists.
- Extract twice. Run 2 independent passes, human plus AI and log every discrepancy and its resolution.
- Long, detailed extraction leads to more hallucinations, so process 1 study at a time.
- Never ask AI to run any statistical analysis. All pooling must be done in proper statistical software like R, Stata, RevMan, or CMA, with code you can share.
- Run an adversarial pass, where you prompt AI to only job is to find errors in the first pass. This is not foolproof or a substitute for human verification, though.
- Keep an audit trail, including model, version, date, prompt, parameters, refinement history, and every human correction.
- Read the full text for any study you cite or are pooling data from. No exceptions here.
Reporting Standards: PRISMA-trAIce and RAISE
What is PRISMA-trAIce?
PRISMA-trAIce is a proposed extension to the PRISMA 2020 statement that establishes uniform standards for reporting the use of AI as a methodological tool in systematic reviews. It is not yet EQUATOR-registered or Delphi-validated, so it is not formally mandatory. Following it is still the safest choice, because its items closely match what publishers and reviewers are starting to ask for.
(Note that earlier guidance such as PRISMA-AI addresses AI as the subject of research, not as a tool used to produce the review.)
The PRISMA-trAIce Checklist, Section by Section
| Manuscript section | Items | What you report |
| Title and Abstract | T1, A1 (optional) | Flag AI in the title if it played a substantial role; summarize tools, stages, and role in the abstract |
| Introduction | I1 (recommended) | Rationale for using AI for these specific tasks |
| Methods | M1 to M5 (mandatory) | Protocol pre-specification and deviations; tool name, version, provider, and access URL; the stage and precise task; input data; output format and post-processing |
| Methods | M6 (mandatory), M7 | Full prompts, key parameters such as temperature, and the prompt-refinement process; operational settings for non-LLM tools |
| Methods | M8 to M10 | Human oversight including reviewer numbers, verification, and discrepancy resolution; performance evaluation methods; ethics, privacy, copyright, and terms-of-service compliance |
| Results | R1, R2 (mandatory) | A flow diagram distinguishing records handled by AI from those handled by humans; agreement rates, recall, and precision |
| Discussion | D1, D2 | Limitations including technical issues, biases, and hallucinations, with their potential impact; reflections on benefits and workload |
Item M6 is the one most authors will struggle with if they look at the checklist only at the time of writing or submission. Prompts, parameters, and refinement history cannot be reconstructed after the fact. Start a prompt log on day 1 of the protocol, not at submission.
What are the RAISE recommendations?
RAISE stands for Responsible use of AI in evidence SynthEsis. It offers tailored recommendations for roles across the evidence synthesis ecosystem: synthesists, methodologists, AI development teams, and the organizations and publishers involved in synthesis. It is published in 3 parts covering recommendations for practice, building and evaluating tools, and selecting and using tools.
Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence all support RAISE.
What AI Use You Must Declare According to RAISE
Run your own workflow against the table below.
| What you used AI for | Declaration expected? | Reason |
| Study eligibility decisions | Yes | AI is making or suggesting a judgment about inclusion |
| Risk-of-bias and quality appraisal | Yes | Appraisal is explicitly named as a judgment |
| Extraction of bibliographic, numerical, or qualitative data | Yes | Extraction shapes the pooled result |
| Synthesis of data from 2 or more studies | Yes | This is the core analytic act |
| Certainty assessments including GRADE domains | Yes | Directly affects the strength of conclusions |
| Drafting text summarizing strength of evidence or implications | Yes | Interpretation presented to readers |
| Plain language summaries | Yes | Interpretation presented to lay audiences and funders |
| Spelling, grammar, and formatting only | Generally no | Does not make or suggest a judgment. Declare only if your target journal requires it. |
Limitations, Ethics, and Wider Costs of AI-Assisted Meta-Analysis
- Accountability rests with named human authors. No publisher policy allows AI as an author.
- AI can be biased: training corpora and search indexes skew toward English-language and high-income-country research.
- Reproducibility can suffer because AI model versions change without notice, which is why archived prompts and parameters matter more than you think they do.
- Confidentiality obligations apply if you’re uploading into AI unpublished manuscripts, embargoed material, or participant-level data.
Every policy cited in this guide keeps changing, because AI is advancing at a fast pace. Re-check them while planning your meta-analysis protocol, before drafting your manuscript, and just before submitting your paper.
Frequently Asked Questions
Can AI do a meta-analysis for me?
No. AI can accelerate searching, screening, extraction, code writing, and drafting, but it cannot select a model, judge whether pooling is defensible, or take responsibility for conclusions. Every major publisher and every synthesis organization places accountability with named human authors. Treat AI as a fast, error-prone assistant whose output you verify.
What is the best AI tool for meta-analysis for beginners?
There is no single best tool, because the workflow spans very different tasks. A common starting combination is a screening and extraction platform such as Rayyan; Zotero for references; and R with the metafor package for pooling. Beginners should check institutional licensing first, since many paid tiers are already covered.
Do I need to declare AI-assisted screening in my systematic review?
Yes. Screening decisions are judgments about study eligibility, and the joint position statement names eligibility decisions explicitly as requiring declaration. You should report the tool name and version, the stage, the human oversight arrangement, and any performance metrics such as agreement rate or recall.
Do journals allow AI in systematic reviews and meta-analyses?
Yes, with conditions. Taylor and Francis explicitly permits AI in ethically conducted literature reviews, systematic reviews, meta-analyses, and bibliometric studies, provided full methodological detail and a declaration are supplied. Other publishers permit it under similar transparency requirements. What is universally prohibited is AI authorship and undisclosed use for important tasks like screening and data extraction.
How do I stop AI from inventing citations?
Never ask a general chatbot to find studies from memory. To prevent fabricated citations, use only tools grounded in a real indexed corpus or in PDFs you have uploaded, require a quoted source sentence for every claim, and resolve every DOI and PMID before it enters your reference list.
Is ChatGPT accurate for statistical pooling in meta-analysis?
Not reliably, and it should not be used that way. Large language models like ChatGPT and Gemini produce plausible-looking arithmetic without any guarantee of correctness. Use AI to write and debug analysis code, then run that code in R, Stata, RevMan, or CMA and report the software output.
Can AI replace a second reviewer in title and abstract screening?
No body like Cochrane nor any journal has agreed that AI can fully replace having a second human reviewer as required in a systematic review. AI can only speed up certain tasks rather than replace a person. In a rapid review, where the design doesn’t require two human reviewers, you can supplement single-reviewer screening with an additional layer of AI-assisted screening. But even in a rapid review, you still need documented human oversight of all AI output, a discrepancy resolution process, and reported performance metrics.
How long does an AI-assisted meta-analysis take?
Expect months, not minutes. AI reduces the amount of time you spend in searching, screening, extraction, and formatting. Protocol development, full-text reading, methodological judgment, and peer review are essentially unchanged. Verification of AI output and documentation of AI use adds new tasks to your workflow.
Key Sources
- Holst D, et al. (2025). Transparent Reporting of AI in Systematic Literature Reviews: Development of the PRISMA-trAIce Checklist. JMIR AI. doi: 2196/80247
- Flemyng E, et al. (2025). Position statement on artificial intelligence use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. Cochrane Database of Systematic Reviews, ED000178. https://doi.org/10.1002/14651858.ED000178
- Thomas J, et al. Responsible use of AI in evidence SynthEsis (RAISE), parts 1 to 3. Open Science Framework. Project DOI: 17605/OSF.IO/FWAUD
- Gartlehner G, et al. (2020). Single-reviewer abstract screening missed 13 percent of relevant studies. Journal of Clinical Epidemiology: https://doi.org/10.1016/j.jclinepi.2020.01.005
- Cook N, et al. (2026). Reporting Guidelines for Meta-Analysis in Economics, Updated for AI. Journal of Economic Surveys. https://doi.org/10.1111/joes.70116
- Page MJ, et al. (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. DOI: 1136/bmj.n71
- Nature Portfolio editorial policy on artificial intelligence. https://www.nature.com/nature-portfolio/editorial-policies/ai
- Taylor and Francis editorial policy on using AI in research and manuscript preparation. https://authorservices.taylorandfrancis.com/editorial-policies/using-ai-in-your-research-and-manuscript-preparations/
- Wiley artificial intelligence guidelines for authors. https://www.wiley.com/en-in/publish/article/ai-guidelines/
- APA Journals policy on generative AI: additional guidance. https://www.apa.org/pubs/journals/resources/publishing-tips/policy-generative-ai


Comment