Judging by a Good-Looking Season

August 20, 2026 · Part 3 of 20

Opening Scene

Before Ronald Fisher formalized randomized experimental design in agricultural field trials in the 1920s, farmers judged a new technique largely by feel: this year’s crop looked healthier, so the new approach must be working. It’s an understandable instinct, and it’s also exactly the trap that produced generations of confidently wrong agricultural advice — advice that looked reasonable in the moment and simply didn’t hold up when tested rigorously.

In Plain English

Before formal A/B testing, and still today in many organizations, decisions get made through informal judgment: launching a change and watching whether things seem to improve, without a genuine control group or randomization to isolate the actual cause. This isn’t a strawman — it remains the default decision-making method in a great many businesses, even ones with access to plenty of data.

The Old Way

The pre-experimental instinct shows up wherever a change and an outcome happen to coincide:

  • A business launching a new marketing campaign right before a seasonal sales bump, crediting the campaign for what the season would have produced anyway.
  • A manager attributing a good quarter to a specific new process, without accounting for other factors that shifted at the same time.
  • A farmer attributing a good harvest to a new technique, without a control plot to show what that season would have yielded regardless.

In each case, a genuinely good outcome and a recent change happened to coincide, and the coincidence got mistaken for proof.

What’s Changing (and Why AI Is the Reason)

  1. Fisher’s core insight — randomize the assignment of treatment, and use a genuine control group — has become the accepted gold standard, moving experimentation from an academic specialty into standard business and product practice.
  2. As digital systems make randomized assignment technically trivial to implement, the cost of running a genuine experiment instead of relying on informal judgment has dropped dramatically compared to Fisher’s era.
  3. AI-driven products, which make many small, frequent decisions, have made rigorous testing of those decisions a genuine necessity — informal judgment simply doesn’t scale to the volume of decisions modern systems make.

The Metaphor, Fully Extended

The Field TrialInformal Judgment Concept
A good-looking season credited to a new technique, without a control plotA good business result credited to a recent change, without a control group
Fisher’s insight: randomize, and compare against a genuine controlThe core principle every modern A/B test still relies on
Decades of confidently wrong agricultural advice, later corrected by rigorous trialsConfidently wrong business decisions, correctable by rigorous testing
A discipline moving from academic specialty to standard practiceA/B testing moving from rare to routine in digital products

For Beginners: What to Actually Do

  • Learn the basic history of Fisher’s agricultural field trials — the origin story of nearly every experimental design principle this series covers.
  • Practice noticing informal judgment in your own organization’s decision-making, and ask what a genuine test would have looked like instead.
  • Get comfortable proposing a real experiment as an alternative to “let’s just try it and see,” even when that feels like it’s adding friction.

For Practitioners and Leaders: The Deeper Layer

  • Audit how many significant decisions in your organization are currently made through informal judgment rather than genuine controlled testing.
  • Build a cultural expectation that testable changes get tested, rather than shipped and judged by feel.
  • Recognize that the discipline this series covers isn’t new or exotic — it’s a century-old, extremely well-validated approach being applied to a new domain.

Quick Recap

  • Informal judgment — launching a change and watching for improvement — remains the default decision method in many organizations.
  • Ronald Fisher’s agricultural field trials established randomization and genuine control groups as the antidote, a century ago.
  • Digital systems have made randomized experimentation dramatically cheaper to run than in Fisher’s era.
  • AI-driven products, making many frequent decisions, have made rigorous testing a genuine operational necessity.

Where This Fits in the Series

Article 3 covered the trap randomization was invented to escape. Article 4 covers randomization itself, the core mechanism that makes a real experiment trustworthy.