🧐

Evaluating & Reducing Hallucination

Building guardrails for a system that will confidently make things up.

Part 1

The Appraiser's First Look

what hallucination actually is in a language model, and why evaluating and reducing it deserves the same rigor an appraiser brings to authenticating an artifact.

Part 2

When Confidence Isn't the Same as Authenticity

why a language model's confident, fluent tone provides no genuine signal about whether its claims are actually true.

Part 3

Before There Was Anyone Checking Provenance

how hallucination was handled, largely informally, before systematic detection and mitigation techniques matured into a genuine discipline.

Part 4

Reading the Object, Not the Story Told About It

how grounding a model's claims in actual retrieved source material is the single most effective way to reduce hallucination.

Part 5

The Difference Between a Forgery and a Guess

the meaningfully different types of hallucination — outright fabrication versus reasonable but incorrect extrapolation — and why the distinction matters for mitigation.

Part 6

Tracing the Chain of Ownership

why requiring a model to cite its actual sources, and verifying those citations genuinely exist, is essential to trustworthy grounded output.

Part 7

When Two Experts Disagree

how self-consistency checks — generating multiple independent answers and comparing them — help catch hallucination that a single answer alone would hide.

Part 8

The Appraiser Who Says 'I'm Not Sure'

why calibrated uncertainty expression — a model genuinely communicating when it doesn't know something — is a mark of real trustworthiness, not weakness.

Part 9

Testing Under Different Light

the evaluation methodology needed to rigorously measure a model's hallucination rate, beyond just spot-checking a few examples.

Part 10

The Cost of a Bad Appraisal

the genuine, real-world consequences of hallucination going undetected, and why this risk deserves serious, deliberate attention.

Part 11

A Second Opinion Before the Sale

how a dedicated verification layer, checking a model's output before it reaches a user, catches hallucination that generation-time mitigation alone would miss.

Part 12

Training the Eye to Spot a Fake

how fine-tuning and alignment techniques reduce a model's underlying tendency to hallucinate, beyond just catching errors after the fact.

Part 13

The Appraiser's Reference Library

why the quality and completeness of a retrieval system's underlying knowledge base directly determines how effective grounding can actually be.

Part 14

When the Provenance Paper Itself Is Forged

the specific, especially dangerous risk of a model fabricating a citation or source that looks entirely legitimate but doesn't actually exist.

Part 15

A Checklist for Every Appraisal

practical, prompting-level techniques anyone can apply immediately to reduce hallucination risk, without needing to fine-tune or rebuild infrastructure.

Part 16

The Appraiser in the Auction House

why hallucination risk compounds across a multi-step agentic task, and why a single wrong step early on can cascade into a genuinely bad final outcome.

Part 17

Insuring Against a Bad Call

how human-in-the-loop review, calibrated to genuine stakes, functions as a deliberate risk management layer against hallucination's worst-case consequences.

Part 18

The Appraiser's Growing Case File

how continuous monitoring and feedback loops let an organization track hallucination rates over time and catch emerging patterns before they cause real harm.

Part 19

Certifying the Whole Practice

how governance and audit trails formalize hallucination risk management at an organizational level, beyond individual system-level mitigation.

Part 20

The Appraiser's Reputation

reassembling every piece covered across this series into the complete picture of how genuine trust in AI output actually gets earned and maintained.