AI hallucinations: what they really are
Why ChatGPT and other models invent facts, citations, or links — what “hallucination” means, and how to cut the risk in everyday use.
The short version
An AI hallucination isn’t the model “seeing things.” It’s when it produces text that is confident, fluent, and wrong: a fake citation, a bad date, a scientific paper that doesn’t exist, a summary that warps the source. Large language models (ChatGPT, Gemini, Claude, and the rest) predict likely next words — not verified truth. That’s the gap between “sounds right” and “is right.”
Where the word comes from (and why it confuses)
Researchers use “hallucination” when a system invents details missing from data or context. The metaphor is vivid and misleading. The AI has no sensory experience. It also has no intent to lie. It generates.
OpenAI and other labs have documented the pattern: even with guardrails, models invent sources, names, and numbers. This isn’t a rare glitch — it’s a property of how they work, something teams try to shrink without fully erasing.
Why it happens
Several mechanisms overlap:
- Prediction, not faithful memory. The model picks words by probability. If “2019 study” fits the sentence, it may emit it without a real study.
- Knowledge gaps. On obscure or very recent topics, it fills blanks with a plausible pattern.
- Pressure to answer. Many interfaces reward a complete reply. “I don’t know” happens — less often than you’d hope.
- Mashed sources. Two true facts badly combined become a third false claim.
- Better ≠ perfect. Newer versions often hallucinate less on common tasks — and still a lot on precise details (references, case law, dosages, quotes).
What it looks like in practice
Typical cases:
- A bibliography entry with DOI and authors… invented
- A legal summary that blends several texts
- A bio with the wrong job title or date
- A link that 404s or points at the wrong page
- Code or a procedure that’s almost right except one critical detail
The danger isn’t style. It’s surface credibility: clean sentences, expert tone, tidy structure.
What it is not
- Not “the AI woke up and started delirating”
- Not only a small-model problem: big models do it too, differently
- Not proof that “AI is useless” — rather that it’s a weak solo source for high-stakes facts
- Not always obvious at a glance: subtle errors are the worst ones
NIST’s AI risk framing is useful here: reliability, transparency, use context. A hallucination in brainstorming isn’t the same cost as one in a medical or legal file.
Before / after: a verification example
Before (confident, unchecked output)
“According to a 2019 WHO study (DOI 10.xxxx/fake), drinking 3 liters of water a day cuts migraine risk by 40%.”
Red flags: round number + very specific DOI / study + assertive tone. Ask: “cite the exact source and URL” — then open the link.
After (same topic, sane method)
- You search yourself for “WHO migraine hydration” / Cochrane / a national health page.
- You may find general hydration advice — without that −40% figure or DOI.
- You rewrite: “Hydration may matter for some people with migraine; I couldn’t find a 2019 WHO study with that number. Confirm with a clinician.”
The AI was a hypothesis draft. Verification stopped a fake reference from going live. Same move for case law, a price, a biography, or an “official docs” link.
How to cut the risk (without paranoia)
- Ask for checkable sources — then open them yourself.
- Split drafting and verification. AI can write; you (or a dedicated tool) check facts.
- Push for uncertainty: “If you’re not sure, say so.”
- Distrust neat number lists and quoted “citations.”
- For law, health, money: treat every output as a draft, never as evidence.
- Cross-check with search, official docs, or a competent human.
A simple 2026 rule
Use AI as a fast assistant (structure, rewrite, ideas, first draft), not as a sworn encyclopedia. Hallucinations shrink with better tools (built-in web search, citations, “reasoning” modes) — they don’t vanish. The final filter is still you.
Teams that use AI well often add a boring step: “source check” before anything leaves the chat. That single habit — open the link, confirm the date, verify the quote — prevents most embarrassing errors. Fluency is cheap now. Trust still has to be earned per claim.
To compare assistants without drowning in marketing: /en/tools/ai.
Going further
- OpenAI — Why language models hallucinate: a lab’s view of the causes.
- NIST — AI Risk Management Framework: how to think about reliability and use.
- Google — AI principles: stated responsibilities from a major provider.
- Stanford HAI — AI Index: data and trends on model capabilities and limits.
Sources
Found an error? Email us — we correct factual mistakes and note significant updates on the article. Contact us
Keep exploring
AI slop: why the internet is filling with generic content
What “AI slop” means, why blogs, ads, and social feeds are full of interchangeable text and images, and how to spot (and dodge) the noise.
Read the article →ChatGPT vs Claude vs Perplexity: which for what
A practical comparison of ChatGPT, Claude, and Perplexity — real strengths, concrete scenarios, limits, and how to choose by task instead of by hype.
Read the article →Is a paid AI subscription worth it in 2026?
ChatGPT Plus, Claude Pro, Gemini Advanced, Copilot Pro — what paid plans actually change, who benefits, and when free is enough.
Read the article →