The Genie's Rulebook

October 15, 2026 · Part 11 of 20

Opening Scene

A wisher who’s learned from experience doesn’t just describe the good outcome they want — they also explicitly rule out the specific bad interpretations a clever, literal-minded genie might otherwise reach for. “Make me wealthy, but don’t take anything from anyone I love, and don’t shorten anyone’s life to do it.” The positive wish alone leaves dangerous room open; the explicit exclusions close it. Prompts benefit from exactly the same discipline.

In Plain English

Negative instructions (or explicit constraints) state what a model should avoid, alongside what it should do — “don’t include any information you’re not confident is accurate,” “don’t use technical jargon,” “don’t exceed 200 words.” These are genuinely easy to forget, since it’s natural to focus entirely on describing the desired outcome, but they close off exactly the kind of unwanted interpretation that Article 4’s ambiguity discussion warned about, often more effectively than positive instructions alone.

The Old Way

Before negative instructions were recognized as a distinct, valuable technique, prompts often relied entirely on positive description:

  • Early prompts commonly described only the desired outcome, leaving unwanted approaches to that outcome entirely unaddressed and open.
  • A model would sometimes technically satisfy a positive instruction through an approach the prompt author never anticipated and wouldn’t have wanted.
  • The specific discipline of explicitly ruling out known problematic responses, not just describing the ideal one, wasn’t yet a standard, widely practiced technique.

Recognizing negative instructions as their own distinct, valuable category represents real, accumulated practical experience about where positive instructions alone tend to fall short.

What’s Changing (and Why AI Is the Reason)

  1. Accumulated practical experience has identified common categories worth explicitly constraining against — unverified claims, off-topic tangents, undesired formatting, tone mismatches — giving practitioners a genuine, reusable checklist.
  2. Negative instructions have become a standard complement to positive instructions in well-engineered prompts, rather than an afterthought added only after a problem is discovered in production.
  3. This connects directly to the failure-mode analysis covered in Article 12 — a genuinely useful negative instruction often comes directly from having previously observed a specific, unwanted response pattern.

The Metaphor, Fully Extended

The Genie’s LampNegative Instructions Concept
A wish describing only the good outcome desiredA prompt describing only the desired positive outcome
Explicitly ruling out the specific bad interpretations a clever genie might reach forExplicitly ruling out specific unwanted model responses
A wisher who’s learned from painful experience what to explicitly excludeA practitioner who’s learned from observed failures what to explicitly constrain against
A rulebook that closes off dangerous, unwanted loopholesA prompt that closes off dangerous, unwanted response patterns

For Beginners: What to Actually Do

  • Practice adding explicit negative instructions to a prompt after observing even one unwanted response pattern, rather than only relying on positive description.
  • Build a small, personal checklist of common negative instruction categories — accuracy caveats, tone, formatting, scope — to consider for any new prompt.
  • Get comfortable revisiting and updating negative instructions as new, unexpected response patterns get discovered over time.

For Practitioners and Leaders: The Deeper Layer

  • Build negative instructions into standard prompt templates as a matter of default practice, not just as reactive patches after a problem occurs.
  • Maintain a shared, organizational log of common failure patterns and their corresponding negative instructions, connecting directly to Article 12’s failure-mode analysis.
  • Treat well-crafted negative instructions as genuine risk mitigation, particularly for prompts used in consequential or public-facing applications.

Quick Recap

  • Negative instructions explicitly state what a model should avoid, complementing positive instructions about what it should do.
  • These are easy to forget but often close off unwanted interpretations more effectively than positive instructions alone.
  • Common categories include accuracy caveats, tone, formatting constraints, and scope limitations.
  • Well-crafted negative instructions often come directly from observing specific, previously unwanted response patterns.

Where This Fits in the Series

Article 11 covered explicitly ruling out unwanted responses. Article 12 covers what to do when, despite everything, the genie still misunderstands the wish.