Confidently Wrong: Hallucination and Uncertainty
By the end you'll be able to
- Explain hallucination as the machine doing its normal job, on a question where plausible and true have pulled apart.
- Predict high-risk conditions: thin coverage, specific checkable details, questions it "shouldn't" be able to answer.
- Apply a verification habit that's matched to the stakes.
- Predict
- Judge
- Decide
- Field note
⚡ 30-second warm-up — From Lesson 1: what's the difference between predictable and true?
They're two separate things. The most plausible-sounding continuation isn't automatically the true one, it's just the one the training diet made most likely. Today that gap gets a name, and a defence.
Ask an AI for five real court cases supporting your argument. It may return five perfectly formatted citations, two of which don't exist. It didn't "lie".
Start with a real professional's mistake
Lawyers have been sanctioned, including in Australian matters, for filing briefs full of AI-invented case citations. The names looked real. The years looked real. The court reporter numbers looked real. Nothing about the format looked wrong, and nothing was checked.
Commit to a prediction
Using what you already know from Act 1, what actually happened when it invented those citations?
This lesson changes one thing: you'll stop hunting for a "tell" in the tone, a giveaway that reveals a fabrication. There isn't one. Instead, you'll start predicting where fabrication is likely, because that follows a pattern.
Discover the mechanism
The raffle always draws a ticket (Lesson 4). When evidence is thick, the winning ticket is usually true. When evidence is thin, a ticket still wins, and the machine announces it in the same fluent, confident voice. That voice was learned from confident, specific human writing (Lesson 2), not from a truth-detector. So Hallucination isn't a malfunction. It's the normal mode, running without the evidence to back it up.
Where it bites hardest: specific checkable facts, like names, numbers, quotes and references. Thin-coverage topics, like your suburb, your niche, or anything recent. And questions from after training day, when there's no retrieval to fall back on (Lesson 13). What helps: sources you can actually click through, and instructions that let the model say "I don't know" without penalty (Lesson 11). Above all, run your own three-question test: is this checkable? would the training diet plausibly cover it? what happens if it's wrong? Match your verification effort to that last answer.
Reported hallucination rates are falling, and retrieval-grounding helps. But the tendency itself is structural to prediction machines, not a bug waiting for a patch, so treat any specific rate figure as a snapshot in time. Lesson 18 asks a different question: what happens when the answer is accurate to the data, but the data carries the world's tilt?Ask them anything, and their character answers. Instantly, specifically, in a confident voice, because the show must go on. That's marvellous for scenes. It's catastrophic if you mistake the performance for testimony.
- Improv skill = fluent pattern continuation (Lesson 1). Never breaking character = the model's default: always produce an answer.
- The confident, specific voice = a tone learned from confident, specific training text (Lesson 2). It tells you about style, not evidence.
- Where it breaks: a real improv actor knows they're inventing, and could stop. The model has no such awareness. It draws no inner line between its true statements and its invented ones, which is exactly why it can't reliably flag its own fabrications.
Judge the gauntlet yourself
A conceptual illustration: six fixed, pre-verified answers, not a live model. Vote on each one before you reveal the verdict, and watch your own calibration build.
Your calibration:
Final station: the same thin question, asked three separate times:
"How many artisan cheese producers operate in Tasmania's Huon Valley?"
- Run 1: "Approximately 12 producers."
- Run 2: "Around 18 boutique producers."
- Run 3: "There are 7 well-known producers."
Same question, same kind of model, three confidently different numbers. None of these was verified. All three were generated. The variation itself is the evidence: this is thin-coverage territory, and no single run's confidence would have told you that.
Where this bites at work
Kerry asks for recommended screening intervals for a client cohort. General prevention patterns are thick territory, but the specific interval is exactly the kind of checkable detail that trips people up. So she clicks through to the current Australian clinical guideline rather than trusting the number outright. That's retrieval-with-verification, the habit from Lesson 13, put to work here.
In health work, question three of the test almost always answers itself: verify, fully, every time.
"Hallucinations are rare glitches that better models will eventually eliminate" is the misconception to retire. Rates are falling, and grounding helps. But plausible-without-true is structural to how these systems work. The honest framing is managed, never cured. A citation or a confident tone is never a substitute for checking, not when the stakes are real.
Now make the call
A client asks Kerry for the recommended interval between screenings for their specific risk profile. The assistant gives a confident, specific answer. What does she do?
Try asking it to confirm its own answer too. It's worth seeing that its confidence about itself is generated the same way as the original claim.
Apply it to your work
Run the three-question test on the next specific, checkable fact an AI gives you: is it checkable? Would the training diet plausibly cover it? What happens if it's wrong? Let that last answer set your verification effort.
Prove it to yourself
Two quick questions, untimed, and you can retry them. They count toward ★ Mastered, along with completing the lesson and the decision. This is what stops "that made sense" evaporating by next week.
Explain it in your own words
Write two or three sentences for a colleague. Then compare them against the rubric below, which checks ideas, not exact wording. Only you see your note.
✓ Lesson 17 stamped: Calibrated Checker
You judged six answers, watched your own calibration take shape, and saw three confident, different answers to one thin question. Lesson 18 asks a new question: what happens when the answer is accurate to the data, but the data carries the world's fingerprints?
Evidence & review — how we know what this lesson claims
| Claim | Type | Basis | Review risk |
|---|---|---|---|
| Hallucination is the normal plausible-continuation process operating on thin or absent evidence | How it works | Primary research literature | Low |
| Confident, fluent tone is a learned stylistic feature uncorrelated with factual correctness | How it works | Primary research literature | Low |
| Reported hallucination rates and the effectiveness of grounding/retrieval mitigations | Product behaviour | Provider / research benchmarks | High — reviewed six-monthly |
| Real incidents of AI-fabricated legal citations submitted in court filings, including Australian matters | Documented incident | External reporting | Medium — reviewed six-monthly |
| The fact-or-fabrication gauntlet | Teaching device | Fixed, pre-verified illustrative items | Medium |
| The improv-actor analogy | Useful mental model | Limits stated in the lesson | Low |
Concept review due: January 2027. Product-behaviour and incident claims: six-monthly.
Where you are in the machine
You're now in the top layer: human judgement, applied to the machine's most convincing failure. Select any layer to see its role, its lessons, and your live progress.
The machine, bottom to top: training data (2) → token pieces (3) → prediction engine (1, 4–9) → the desk (10–11) → senses (12) → live knowledge (13) → tools (14) → memory (15) → agent loops (16) → human judgement (17–21).
Next: the world's fingerprints
Invention is one failure family. The second is subtler: answers that are accurate to the data, but the data carries the world's fingerprints. Next: bias and blind spots.