Confidently Wrong: Hallucination and Uncertainty
By the end you'll be able to
- Explain hallucination as the machine doing its normal job on questions where plausible and true diverge.
- Predict high-risk conditions: thin coverage, specific checkable details, questions it "shouldn't" be able to answer.
- Apply a verification habit proportionate to stakes.
- Predict
- Judge
- Decide
- Field note
⚡ 30-second warm-up — From Lesson 1: what's the difference between predictable and true?
Two separate properties. The most plausible-sounding continuation is not automatically the correct one — it's just the one the diet made most likely. Today that gap gets a name and a defence.
Ask an AI for five real court cases supporting your argument. It may return five perfectly formatted citations — two of which don't exist. It didn't "lie".
Start with a real professional's mistake
Lawyers have been sanctioned — including in Australian matters — for filing briefs with AI-invented case citations: real-looking names, real-looking years, real-looking court reporter numbers. Nothing about the format looked wrong. Nothing was checked.
Commit to a prediction
Using what you already know from Act 1, what actually happened when it invented those citations?
This lesson changes one thing: you'll stop hunting for a "tell" in the tone that reveals a fabrication, and start predicting where fabrication is likely — because there is no tell, and there is a pattern.
Discover the mechanism
The raffle always draws a ticket (Lesson 4): when evidence is thick, the winning ticket is usually true; when evidence is thin, a ticket still wins — and the machine announces it in the same fluent, confident voice, because that voice was learned from confident, specific human writing (Lesson 2), not from a truth-detector. Hallucination is not a malfunction mode; it is the normal mode, operating outside its evidence.
Where it bites hardest: specific checkable facts (names, numbers, quotes, references), thin-coverage topics (your suburb, your niche, anything recent), and questions after training day without retrieval (Lesson 13). What helps: sources you can actually click through, instructions that make "I don't know" an acceptable answer (Lesson 11), and your own three-question test — is this checkable? would the training diet plausibly cover it? what happens if it's wrong? Match your verification effort to the third answer.
reported hallucination rates are falling and retrieval-grounding helps — but the tendency is structural to prediction machines, not a bug awaiting a patch; treat any specific rate figure as a snapshot. Lesson 18 asks what happens when the answer is accurate to the data, but the data carries the world's tilt.Ask them anything — their character answers, instantly, specifically, in a confident voice, because the show must go on. Marvellous for scenes; catastrophic if you mistake the performance for testimony.
- Improv skill = fluent pattern continuation (Lesson 1). Never breaking character = the structural default to always produce output.
- The confident, specific voice = tone learned from confident, specific training text (Lesson 2) — its style, not its evidence.
- Where it breaks: a real improv actor knows they're inventing and could stop; the model has no such awareness — no inner distinction between its true and invented statements, which is precisely why it can't reliably flag its own fabrications.
Judge the gauntlet yourself
Conceptual illustration — six fixed, pre-verified answers, not a live model. Vote on each before revealing, and watch your own calibration build.
Your calibration:
Final station — the same thin question, asked three separate times:
"How many artisan cheese producers operate in Tasmania's Huon Valley?"
- Run 1: "Approximately 12 producers."
- Run 2: "Around 18 boutique producers."
- Run 3: "There are 7 well-known producers."
Same question, same kind of model, three confidently different numbers. None of these was verified — all three were generated. The variation itself is the evidence: this is thin-coverage territory, and no single run's confidence tells you that.
Where this bites at work
Kerry asks for recommended screening intervals for a client cohort: general prevention patterns are thick territory, but the specific interval is exactly the checkable-detail category. She clicks through to the current Australian clinical guideline rather than trusting the number outright (retrieval-with-verification, Lesson 13's habit, applied here).
In health work, question three of the test almost always answers itself: verify fully.
"Hallucinations are rare glitches that better models will eventually eliminate" is the misconception to retire. Rates are falling and grounding helps, but plausible-without-true is structural to how these systems work — the honest framing is managed, never cured. A citation or a confident tone never substitutes for checking when the stakes are real.
Now make the call
A client asks Kerry for the recommended interval between screenings for their specific risk profile. The assistant gives a confident, specific answer. What does she do?
Try asking it to confirm its own answer too — worth seeing that its confidence about itself is generated the same way as the original claim.
Apply it to your work
Run the three-question test on the next specific, checkable AI-supplied fact you're about to use: is it checkable? would the diet plausibly cover it? what happens if it's wrong? Let the third answer set your verification effort.
Prove it to yourself
Two quick questions — untimed, retryable, and they count toward ★ Mastered (completed + decision + self-check all correct). This is what defeats the feeling of "that made sense" evaporating by next week.
Explain it in your own words
Two or three sentences for a colleague. Then compare against the rubric — it checks ideas, not wording, and only you see your note.
✓ Lesson 17 stamped: Calibrated Checker
You judged six answers, watched your own calibration form, and saw three confident, different answers to one thin question. Lesson 18 asks: what happens when the answer is accurate to the data — and the data carries the world's fingerprints?
Evidence & review — how we know what this lesson claims
| Claim | Type | Basis | Review risk |
|---|---|---|---|
| Hallucination is the normal plausible-continuation process operating on thin or absent evidence | How it works | Primary research literature | Low |
| Confident, fluent tone is a learned stylistic feature uncorrelated with factual correctness | How it works | Primary research literature | Low |
| Reported hallucination rates and the effectiveness of grounding/retrieval mitigations | Product behaviour | Provider / research benchmarks | High — reviewed six-monthly |
| Real incidents of AI-fabricated legal citations submitted in court filings, including Australian matters | Documented incident | External reporting | Medium — reviewed six-monthly |
| The fact-or-fabrication gauntlet | Teaching device | Fixed, pre-verified illustrative items | Medium |
| The improv-actor analogy | Useful mental model | Limits stated in the lesson | Low |
Concept review due: January 2027. Product-behaviour and incident claims: six-monthly.
Where you are in the machine
You're now in the top layer — human judgement, applied to the machine's most convincing failure. Select any layer for its role, lessons and your live progress.
The machine, bottom to top: training data (2) → token pieces (3) → prediction engine (1, 4–9) → the desk (10–11) → senses (12) → live knowledge (13) → tools (14) → memory (15) → agent loops (16) → human judgement (17–21).
Next: the world's fingerprints
Invention is one failure family. The second is subtler: answers that are accurate to the data — and the data carries the world's fingerprints. Next: bias and blind spots.