Key Takeaways
- Hallucination is inherent in generative AI. Large language models generate statistically plausible text; they do not retrieve verified facts. A fabricated citation and a real one are produced by exactly the same process.
- Citations are the highest-risk output. Peer-reviewed studies report fabrication rates from 18% to 69%, varying by model, discipline, and prompt design.
- Plausibility is the danger. The typical fabricated reference names real authors, a real journal, and a working DOI that points to an unrelated paper.
- Only humans can verify AI output. Asking a chatbot whether its own output is real is not a check; it is a second chance to hallucinate.
What Is an AI Hallucination in Research?
An AI hallucination is output that reads as authoritative but has no basis in the model’s training data, in the sources you supplied, or in reality. In a research context, this rarely looks like nonsense. It looks like a citation, a statistic, a quotation, or a summary that is formatted correctly and fits the argument perfectly.
The critical point is mechanical: a language model does not distinguish between generating a true sentence and generating a plausible one. Both come out of the same process. Fluency is not evidence of accuracy.
AI hallucination vs. factual error vs. bias
The following 4 issues get conflated, but they need different fixes. Mislabeling a hallucination as “bias” or “just an error” leads researchers to the wrong remedy.
| Failure mode | What it looks like | What actually fixes it |
| Hallucination | A confident claim, citation, or quote with no real referent anywhere. | External verification against primary sources. |
| Factual error | A real entity described incorrectly: wrong date, wrong figure, wrong attribution. | Fact-checking; sometimes better retrieval. |
| Bias | Systematic skew in framing, coverage, or which perspectives get represented. | Diverse sourcing; explicit counter-prompting. |
| Sycophancy | The model agreeing with your stated hypothesis regardless of evidence. | Neutral prompt phrasing; asking for the opposing case. |
Why the term “hallucination” is contested
Many researchers object to the word, and the objection is worth understanding before you use it in a paper.
- It implies perception. The model is not misperceiving a real world; it has no perceptual access to one at all.
- It implies deviation from normal function. Generating plausible text is normal function; accuracy is the incidental outcome when the plausible text happens to be true.
- It borrows a clinical term for a computational phenomenon, which some argue trivializes psychiatric symptoms.
- Proposed alternatives include “confabulation,” “fabrication,” and simply “bullshit” in the technical, Frankfurtian sense of indifference to truth.
“Hallucination” remains the dominant term in the literature, so this article uses it. Just be aware that it names a symptom, not a malfunction.
Where hallucinations enter the research workflow
Risk is not evenly distributed across tasks. It concentrates wherever the model must produce specific, verifiable particulars that it cannot look up.
| Workflow stage | Typical task | Hallucination risk |
| Topic scoping | Explaining a concept, mapping a debate | Low to moderate: broad claims are usually well-represented in training data. |
| Literature search | Asking for relevant papers on a topic | Very high: this is the single most dangerous use. |
| Summarizing a supplied paper | Condensing a PDF you uploaded | Moderate: the source is real, but conclusions get overstated or reversed. |
| Data extraction | Pulling figures from tables or transcripts | High: numbers get transposed, averaged, or invented to fill gaps. |
| Statistical interpretation | Explaining what a result means | High: models invent plausible p values and effect sizes. |
| Drafting and editing | Improving prose you wrote | Low if you have written the bulk of the text, high if you ask the AI tool to create sentences, paragraphs, or sections |
| Citation formatting | Converting an existing reference to APA, MLA, etc. | Low if the reference is real; the model may still “correct” real details. |
Sample prompt: scoping, which is low risk
Explain the main theoretical disagreements between the capability approach and standard welfare economics. Do not cite specific papers. I will find sources myself.
This works because it asks for conceptual structure, which models handle well, and it explicitly blocks the output type that fails most often.
Why Do AI Hallucinate? The Core Causes
There is no single cause. Hallucination emerges from at least 5 separate pressures, which is why no single fix eliminates it. Understanding which pressure applies to your task tells you which mitigation will help.
Why do AI hallucinate when asked for citations specifically?
Citations are the worst case for a text predictor, and the reason is structural rather than accidental.
- The format is extremely regular. Author, year, title, journal, volume, pages. A model learns this shape perfectly from millions of examples, so it can generate flawless-looking references indefinitely.
- The content is arbitrary. There is no pattern connecting a topic to the exact page numbers of a specific 2019 article. That mapping must be memorized, and most of it never was.
- Individual references are long-tail. Any given paper appears rarely in training data. The format appears constantly. The model has learned the container far better than the contents.
- Numeric fields fail first. In one study of 636 citations, volume, issue, page, and year errors were the most common defects even in otherwise-real references.
- Book chapters are the worst category. In the same study, 70% of GPT-4’s cited book chapters were fabricated, and often the containing book did not exist either.
Sample prompts: the risky ask and the safer ask
RISKY
Write a literature review on gut microbiome and depression with 10 peer-reviewed citations.
SAFER
I have attached 8 PDFs. Write a literature review on the link between gut microbiota alterations and antidepressant use, using only these papers. Cite by the filename and page number. If a claim I need is not supported by these 8 papers, say “not covered by supplied sources” instead of citing anything else.
Prediction optimizes for plausibility, not truth
A language model is trained to make text likely, not to make it correct. When the training objective is “predict what comes next,” a false-but-typical continuation scores as well as a true one, provided both look like the kind of thing that gets written.
Researchers at OpenAI formalized this in 2025, showing that hallucinations arise as ordinary statistical errors during pretraining rather than as a mysterious glitch. When a fact cannot be reliably distinguished from a plausible non-fact given the training signal, some rate of fabrication is mathematically expected.
Training data gaps, staleness, and long-tail topics
- Every model has a training cutoff. Anything after it is unknown, and the model may generate confident content about it anyway.
- Niche subfields, non-English scholarship, gray literature, and paywalled corpora are thinly represented, so the model interpolates.
- Recently retracted or superseded findings often remain in the model’s training data and get repeated as current.
- Contradictory sources in training data produce blended outputs that match no actual source.
- Rare proper nouns are especially fragile: obscure author names, small journals, and regional institutions get merged or invented.
Evaluation and training that reward guessing over abstention
This cause is the least intuitive and arguably the most important. Most benchmarks score a wrong answer and an “I don’t know” identically, at zero. Under that scoring, guessing is always the optimal strategy, exactly as it is for a student on a multiple-choice exam with no penalty for wrong answers.
Models are therefore trained and selected to behave like confident test-takers. The proposed fix is to change the scoring so that incorrect answers are penalized more than abstentions, which would make calibrated uncertainty the winning strategy. Until benchmarks change, confident guessing remains rewarded.
There is a second, related pressure. Reinforcement learning from human feedback trains on human ratings, and humans reliably prefer confident answers to hedged ones. Confidence gets reinforced whether or not it is warranted.
Ambiguous prompts and leading questions
Users generate a large share of their own hallucinations. A prompt that presupposes a fact will usually get that fact confirmed, complete with supporting evidence that does not exist.
Sample prompts: leading vs. neutral
LEADING, PRESUPPOSES THE FINDING
Summarize the studies showing that remote work reduces employee productivity.
NEUTRAL, ALLOWS A NULL RESULT
What does the empirical literature find about remote work and employee productivity? Report findings in both directions, note where evidence is mixed or weak, and state explicitly if you are uncertain whether specific studies exist.
Tips for framing your own prompts
- Do not name a conclusion you want supported.
- Do not ask “find studies proving X.” Ask “what does the evidence say about X.”
- Avoid asking for a fixed number of sources. “Give me 10 citations” pressures the model to fill quota with inventions.
- Do not push back repeatedly until the model changes its answer; that trains the conversation, not the truth.
How Hallucination in AI Models Works in Practice
Knowing the mechanism helps you predict where a given tool will fail, which is more useful than a general warning to “be careful.”
Hallucination in AI models: intrinsic vs. extrinsic
The standard taxonomy splits hallucinations by their relationship to a provided source. This distinction determines whether grounding will help you.
| Type | Definition | Example |
| Intrinsic | The output contradicts a source that was supplied to the model. | You upload a paper reporting no significant effect; the summary reports a significant effect. |
| Extrinsic | The output makes claims that cannot be checked against any supplied source. | The model adds a citation to a 2021 meta-analysis that was never in your document set. |
| Grounding helps? | Partially: contradictions become detectable by comparison. | Yes for detection, no for prevention: the model can still add unsupported material. |
Does retrieval-augmented generation (RAG) reduce AI hallucinations?
RAG reduces hallucination substantially. It does not remove it, and researchers who assume otherwise get caught by a narrower but subtler set of failures.
- Retrieval can miss. If the search step returns nothing relevant, most systems answer anyway from parametric memory rather than declining.
- Retrieval can return the wrong passage. Semantic search matches topic, not claim. A passage about the same subject can support the opposite conclusion.
- Synthesis still hallucinates. Combining 5 real passages into one summary is a generative step, and the connective tissue between them is invented.
- Citations can be misattached. The tool cites a real, retrieved document for a sentence that document does not actually support. This is the hardest failure to catch, because the link works.
- The corpus itself may be contaminated. Preprint servers and repositories now contain AI-generated text with fabricated references, which retrieval will faithfully surface.
Settings and conditions that change the risk profile
| Condition | Effect on hallucination risk | What to do |
| High temperature | Increases variability and invention. | Use the lowest setting available for factual work. |
| Long context or long conversation | Earlier instructions and sources lose influence. | Start a fresh session per document or per question. |
| Obscure or non-English topic | Sharply increases risk. | Supply sources yourself; do not rely on recall. |
| Reasoning or extended-thinking mode | Generally reduces factual errors. | Enable it for anything analytic. |
| Web search enabled | Reduces fabrication of sources; does not eliminate misreading. | Still open every link and confirm the claim. |
| Requests for a fixed quantity | Pressures the model to fill quota. | Ask for “however many are well-supported.” |
Examples of AI Hallucination Across Subject Areas
Hallucination is not a quirk of one discipline. What changes across fields is the form the fabrication takes and who gets hurt when it survives review.
Examples of AI hallucination in law: the Mata v. Avianca sanction
This is the reference case. In a personal injury suit against the airline Avianca, an attorney used ChatGPT to research a brief and submitted 6 court decisions that did not exist. Judge P. Kevin Castel of the Southern District of New York sanctioned the attorneys and their firm $5,000 in June 2023.
The instructive detail is not that the model invented cases. It is what happened next.
- The fabricated opinions named real, sitting federal judges as their authors, and cited docket numbers belonging to unrelated real cases.
- Each fake opinion contained internal citations to further cases that also did not exist.
- When the attorney asked ChatGPT whether the cases were real, it confirmed that they were and named Westlaw and LexisNexis as places to find them. That confirmation was itself a hallucination.
- The court was explicit that using AI is not improper. Failing to verify is.
The pattern has repeated many times since across multiple jurisdictions, and later courts have treated the passage of time since this ruling as removing any excuse of novelty.
Medicine and biomedical research
Medicine has the largest body of measured evidence, partly because reference accuracy is verifiable through PubMed.
- An observational study of 30 short medical papers generated by ChatGPT found that of 115 generated references, 47% were fabricated, 46% were real but inaccurate, and only 7% were both real and accurate.
- In the same study, an incorrect PubMed ID appeared in 93% of papers, making the identifier itself an unreliable signal of authenticity.
- A study of ChatGPT responses to medical questions found 41 of 59 references, or 69%, were fabricated while appearing entirely credible.
- Another multidisciplinary study found 64% of 343 citations generated by ChatGPT-3.5 and ChatGPT-4 could not be located in PubMed or on the open web.
- Fabricated medical references typically pair real, topically appropriate authors with a real journal and an invented title, which defeats casual plausibility checks.
Psychology and mental health research
Linardon et al. (2025) asked a model to produce 6 literature reviews on mental health topics and audited all 176 resulting citations.
| Outcome | Share of 176 citations | Note |
| Real and accurate | About 44% | Fewer than half were usable as given. |
| Contained errors | About 45% | Real work, wrong details. |
| Fully fabricated | About 20% | No such work exists. |
| Fabricated but with a DOI | Over 94% of fabrications | The identifier looked legitimate. |
| Fake DOI resolving to a real, unrelated paper | About 64% of fake DOIs | Clicking through was the only way to catch it. |
That last row is the finding researchers should internalize. A DOI that resolves is not proof of anything. You have to read what it resolves to.
Another important finding was that AI accuracy was higher on well-known, well-researched topics, compared to niche, emerging, or under-researched topics.
Economics, statistics, and quantitative social science
Here the fabrication moves from the reference list into the results section, which is considerably more dangerous.
- In an audit of 84 model-generated papers across policy, economics, and other social sciences, 5 of the “literature reviews” were structured as empirical studies with invented methods and results.
- Among them was a paper containing fabricated correlation coefficients, regression coefficients, and p values, all internally consistent and plausible in magnitude.
- Models routinely invent survey sample sizes, response rates, and confidence intervals when asked to summarize a study they cannot access.
- Asked to interpret a real table, models transpose columns, misread footnotes, and silently convert nominal to real figures.
- Aggregate statistics such as “roughly 40% of firms” are generated with no underlying source and are difficult to trace precisely because they sound like common knowledge.
Sample prompt: extracting numbers with a hallucination trap built in
From the attached PDF, extract every figure I asked about into a table with 3 columns: the value, the exact page number, and the verbatim sentence it came from. If a value does not appear in the document, write NOT PRESENT. Do not calculate, infer, or supply any number from your own knowledge.
Requiring the verbatim source sentence makes fabrication far more difficult for the model and instantly checkable for you.
Computer science and software engineering
Code is where hallucination becomes an attack surface rather than merely an error.
- A large study of 16 code-generating models across 576,000 samples found 19.6% of recommended software packages did not exist, with roughly 5% for commercial models and 22% for open-source ones.
- The study catalogued 205,474 unique hallucinated package names, and only 0.17% corresponded to real packages that had been deleted. The rest were pure invention.
- Hallucinated names repeat deterministically, which means an attacker can farm them, register the popular ones, and wait. This attack is now called slopsquatting.
- Models also invent function signatures, command-line flags, and configuration keys for real libraries, which is harder to spot than a missing package because the library is genuine.
- More recent measurements on frontier models show that hallucination rates have reduced to roughly 5% to 6%, which is lower but not resolved.
History and the humanities
Humanities research has less quantified evidence, mostly because archival claims are expensive to audit. The failure modes are nonetheless well documented anecdotally and follow directly from the mechanism.
- Invented archival references: a plausible collection name, box number, and folder at a real repository.
- Misattributed quotations, where a real aphorism is assigned to a more famous figure, mirroring the same error already common in training data.
- Fabricated translations of passages from texts the model has not memorized, often stylistically convincing.
- Confident dating of undated manuscripts, and invented provenance chains for objects.
- Blended historiography, where the positions of 2 or 3 real scholars are merged into a single school of thought that no one actually holds.
ChatGPT Hallucination Rates: What the Studies Actually Measure
ChatGPT is measured more than other systems because of its adoption, not because it is uniquely unreliable. Every model in this class hallucinates. Treat the numbers below as evidence about a category, not a verdict on one product.
ChatGPT hallucination rates reported in citation studies
|
The trend across model generations is real and substantial. In one study, fabrication fell from 55% to 18% between GPT-3.5 and GPT-4, and substantive errors in the remaining real citations fell from 43% to 24%. Improvement is not the same as reliability: 18% is still roughly 1 fabricated source in every 5.
Why published rates will not predict your rate
- Discipline matters. Fabrication rates ran higher in geography than in medicine in the same comparison.
- Obscurity matters more than discipline. Well-covered topics produce more real citations; niche or emerging ones produce more fabricated citations.
- Prompt design shifts results measurably. Asking for a set number of sources raises fabrication.
- Tool configuration dominates. A model with live search and one without are effectively different systems.
- Studies age quickly. Most published figures test models that are now 2 or 3 generations old.
- Definitions differ. Some studies count a wrong volume number as an error, others as a fabrication.
Sample prompt: measure your own rate before trusting a workflow
Give me 5 peer-reviewed sources on [your exact research topic], with full bibliographic details and DOIs.
- Then check all 5 yourself in Scopus, Web of Science, or PubMed.
- Repeat across 4 topics for a 20-citation sample.
- The fabrication rate you find, not a published benchmark, is your real risk number.
The Risks of AI Hallucination for Researchers and Institutions
The consequences scale outward: from an individual embarrassment, to a corrupted literature, to institutional liability.
Personal and professional risk
- Retraction, which is permanent and publicly indexed.
- Desk rejection and reviewer distrust that carries into future submissions.
- Academic misconduct proceedings, since fabricated citations can be treated as fabrication regardless of intent.
- Loss of grant funding and, for students, thesis failure or degree revocation.
- Reputational damage that outlasts the correction, as the sanctioned attorneys in the Avianca case discovered.
Contamination of the scholarly record
The systemic risk is worse than the individual one, and it compounds.
- Fabricated references are more likely to survive in preprints, theses, and institutional repositories than in fully copyedited journal articles.
- Once a fake reference is cited, later authors may cite it secondhand without ever attempting retrieval, creating a citation cascade.
- Some fabricated references point to journals that have been identified as predatory, laundering low-quality journals into legitimate ones.
- Contaminated repositories then become retrieval corpora for the next generation of AI tools, closing the loop.
- Systematic reviews and meta-analyses are especially exposed, because they aggregate at scale and rarely re-verify every included reference.
Legal, regulatory, and research-integrity exposure
| Setting | Exposure |
| Litigation and legal filings | Sanctions, fee-shifting, dismissal, and mandatory disclosure to affected judges. |
| Clinical and public health guidance | Patient harm, professional liability, and regulatory action. |
| Regulatory and compliance submissions | Findings of misrepresentation; penalties independent of intent. |
| Grant applications and reporting | Findings of research misconduct; funder debarment. |
| Journalism and expert testimony | Defamation exposure and loss of expert credibility. |
Automation bias: the reason plausible errors survive
We do not manually recheck the output of statistical software, and that habit of trust transfers to AI tools even though the underlying task is fundamentally different. A statistics package computes; a language model generates.
Hallucinations are unusually hard to catch for 3 reasons. They are fluent, so nothing prompts a second look. They are confident, and people prefer confident answers, which is why models are trained toward confidence. And they are topically apt, so they confirm rather than disrupt your expectations.
The practical implication: your attention is drawn to output that looks wrong, but the dangerous output looks right. When you’re verifying AI output, you need to follow a defined workflow rather than “eyeball” or trust your intuition.
How to Detect and Reduce Hallucinations in Your Workflow
None of this requires abandoning AI tools. It requires treating their output as an unverified draft rather than as evidence.
A verification checklist for every AI-assisted claim
| Step | What to do | Why it catches things |
| 1. Resolve the DOI | Paste it into doi.org and open the target. | Fake DOIs frequently resolve to a real but unrelated paper. |
| 2. Search title and author | Query Scopus, Web of Science, or PubMed independently. | Confirms the work exists as described, not just that something similar does. |
| 3. Open the source | Retrieve the full text, never just the abstract | Real paper, invented finding is a common failure. |
| 4. Locate the claim | Find the specific sentence, table, or figure cited. | Catches misattached citations, the hardest RAG failure. |
| 5. Check the numbers | Verify volume, issue, pages, and year separately. | Numeric fields fail most often even in real references. |
| 6. Check the chapter | For book chapters, confirm the containing book exists. | Chapters have the highest fabrication rate of any type. |
| 7. Log what was AI-assisted | Keep a record of which claims came from a model. | Makes disclosure and later audit possible. |
The one step that does not belong on this list is asking the model to check itself. In the Avianca case that step produced a confident confirmation of 6 nonexistent cases, and studies consistently find models defend their own fabrications.
Prompting practices that measurably lower risk
| Goal | Sample prompt |
| Force abstention | If you are not confident a specific source exists, write “I do not know of a specific source” instead of producing a citation. I would rather have 2 real references than 10 uncertain ones. |
| Ground in your own corpus | Answer using only the attached documents. For each claim, quote the sentence you relied on and give its page number. If the documents do not address the question, say so. |
| Surface uncertainty | After your answer, list every factual claim you made that you are less than 90% confident about, and say what would need to be checked. |
| Get the opposing case | Now argue the opposite position as strongly as you can, and identify the strongest evidence against what you just told me. |
| Separate recall from reasoning | Do not cite anything. Explain the mechanisms and the main debates. I will supply the sources and come back to you. |
| Cross-check a suspect claim | Here is a claim I found. Without assuming it is true, tell me what evidence would confirm or disconfirm it and where I would look. |
Tools that help, and where each stops helping
| Tool type | What it catches | What it misses |
| Reference managers | DOIs that fail to resolve; metadata mismatches. | Real DOIs attached to the wrong claim. |
| Search-enabled AI | Wholesale invention of sources. | Misreading of retrieved sources; misattached citations. |
| Citation audit scripts | Bulk existence checks across a bibliography. | Whether the source supports what you said it does. |
| AI text detectors | Little that is reliable; false positives are common. | Not a verification method. Do not depend on these. |
| A second model | Some self-inconsistent fabrications. | Shared errors, since models share training data and biases. |
| A professional human librarian | Nearly everything above, plus database strategy. | Underused. Subject librarians are the highest-value check available. |
Disclosure obligations
- Most major publishers now require disclosure of generative AI use in the methods or acknowledgments.
- AI tools cannot be listed as authors, since they cannot take responsibility for the work. This is now near-universal policy.
- Requirements differ between using AI for language editing and using it for content generation or analysis. Check the specific journal.
- Funders and institutions increasingly have separate policies from publishers. Both apply.
- Policies are changing quickly, so verify against the current author guidelines rather than what applied to your last submission.
Frequently Asked Questions
How do I check if an AI-generated citation is real?
Search the exact title in Google Scholar, Scopus, Web of Science, or PubMed, then resolve the DOI separately and confirm the target matches the title and authors given. Both steps are necessary: a real-looking DOI can resolve to an unrelated paper, and a real title can be paired with fabricated publication details. Finally, open the source and locate the specific claim. Never ask the model to verify itself.
Why does ChatGPT make up DOIs that link to real papers?
A DOI is a short, highly patterned string, so the model generates one that looks structurally correct without any lookup. Some of those strings happen to be live identifiers belonging to other papers. In one 2025 audit, more than 94% of fabricated citations carried a DOI, and about 64% of those resolved to real but unrelated work. A link that opens is not evidence; only the content at the other end is.
Can I use ChatGPT for a literature review?
Use it for scoping, terminology, and structure, not for finding sources. It is genuinely good at explaining a concept, generating search terms, and polishing language. It is unreliable at producing the reference list, which is where measured fabrication rates run from 18% to 69%. Find sources in a real database always.
Do newer AI models still hallucinate?
Yes, at lower rates. Fabrication in one multidisciplinary study fell from 55% to 18% across a single model generation, and reasoning modes reduce factual errors further. But the cause is structural rather than a fixable defect, and OpenAI’s own researchers describe hallucination as an expected statistical outcome of current training and evaluation practice. Assume a reduced rate, not a zero rate.
What is the difference between an AI hallucination and a lie?
A lie requires knowing the truth and choosing to misstate it. A language model has no internal representation of truth to depart from; it produces the most plausible continuation, and truth is incidental. This is why some philosophers argue the accurate label is indifference to truth rather than deception. The distinction matters practically: there is no dishonesty to detect, so behavioral cues will not help you.
Do I need to disclose AI use when submitting to a journal?
Almost certainly yes, though the threshold varies. Most major publishers require disclosure of generative AI use in the manuscript, and none permit AI tools to be listed as authors. Some distinguish between language polishing and substantive content generation. Check the specific journal’s current author guidelines, along with your funder and institutional policies, since these are separate and all apply.
Can AI hallucinations be eliminated completely?
Not with current architectures. The behavior follows from generating text by prediction rather than retrieval, and from evaluation practices that score confident guessing above admitting uncertainty. Grounding, retrieval, and reasoning modes reduce the rate substantially. Changing benchmark scoring to reward calibrated abstention would help further. Complete elimination would require a fundamentally different design.
Which AI tool hallucinates the least for academic research?
The honest answer is that published rankings go stale within months, and configuration matters more than brand. A model with live search and grounding in your own documents will outperform a more capable model answering from memory. Rather than choosing on reputation, run a test on your own topics and manually verify 20 citations. You’ll then be able to gauge the actual likelihood of citation hallucination if you continue using AI in your research.


Comment