Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder
AI Fundamentals

Aliases: Confabulation

Hallucination

When an LLM confidently generates false information — a statistical property of how these models work, not a bug that will be patched out.

Last reviewed: July 25, 2026

What is hallucination?

Hallucination is an LLM producing fluent, confident, wrong output: invented citations, nonexistent API parameters, fabricated case law, plausible-but-false statistics. It happens because LLMs are trained to predict likely text, not to verify truth — when the model lacks knowledge, the most probable-sounding answer is still generated with the same confident tone.

Why it won’t fully go away

Training objectives historically rewarded guessing over admitting uncertainty (a test-taker who never leaves blanks scores better). Newer models hallucinate less and abstain more, but the failure mode is inherent to next-token prediction. Engineering conclusion: design systems assuming some rate of confident falsehood, the way distributed systems assume some rate of network failure.

The mitigation stack (in order of impact)

  1. Grounding — RAG or tool calls put verifiable facts in context; instruct the model to answer only from them and say “not found” otherwise.
  2. Verification — check claims against sources (guardrails), validate generated code by running it, validate citations by resolving them.
  3. Abstention prompting — explicitly permitting “I don’t know” measurably reduces fabrication.
  4. Measurement — track hallucination rate on your own domain with evals; rates vary wildly by topic and model.

What people get wrong

  • Treating low hallucination benchmarks as safety. A 2% rate at a million queries/day is 20,000 wrong answers daily; what matters is consequence per error in your domain.
  • Thinking fine-tuning fixes it. Tuning shapes style; ungrounded factual recall stays unreliable.
  • Trusting confidence. Fluency and certainty of tone carry no signal about correctness — that’s precisely what makes hallucination dangerous.

Why Hallucination Is Structural, Not a Bug

Language models are trained to predict plausible next tokens based on patterns in their training data, not to verify factual accuracy against a ground truth source — when a model doesn’t have reliable information about something, it will still generate a fluent, confident-sounding continuation, because fluency and confidence in phrasing are exactly what its training objective optimizes for, independent of whether the underlying claim is true. This is why hallucination rates tend to increase for more obscure facts, more recent events past a model’s training cutoff, and highly specific details (exact dates, statistics, citations) that the model may have seen too rarely during training to have reliably encoded.

Mitigation Strategies

Retrieval-augmented generation reduces hallucination by grounding responses in retrieved source documents rather than relying purely on the model’s internal, sometimes unreliable, memorized knowledge. Lower sampling temperature reduces some forms of hallucination by making output more conservative, though it doesn’t eliminate the underlying issue. Explicitly instructing a model to express uncertainty or decline to answer when it lacks confidence, and having it cite sources it can be checked against, are common practical techniques — but no current technique eliminates hallucination entirely, which is why human review remains standard practice for any high-stakes application built on LLM output.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.