Different Soil, Different Result

November 5, 2026 · Part 14 of 20

Opening Scene

A fertilizer trial’s overall average result might show a modest, genuinely positive effect across the whole farm. But average across enough different soil types, and that modest positive number can hide something more interesting underneath: a strong positive effect in sandy soil, and an actual negative effect in dense clay. The average, taken alone, tells a true but genuinely incomplete story.

In Plain English

Heterogeneous treatment effects occur when a treatment’s real effect differs meaningfully across subgroups of the population — by user segment, by region, by device type, by soil type — rather than being uniform. Reporting only an overall average effect, without examining whether it varies meaningfully across relevant subgroups, can mask both genuinely strong opportunities and genuinely important risks hiding within an otherwise unremarkable average.

The Old Way

Before formal subgroup analysis was standard experimentation practice, average effects were often reported and acted on without this deeper look:

  • A medical treatment’s overall average effect reported without examining whether it worked differently across age groups or pre-existing conditions, potentially masking real risks for specific patients.
  • A marketing campaign’s average lift reported without examining whether it actually varied by customer segment, potentially missing that it only worked for one segment and actively hurt another.
  • A fertilizer’s average yield effect reported without examining variation across different soil types, exactly the scenario opening this article.

In each case, an average, reported alone, could paper over genuinely important variation underneath.

What’s Changing (and Why AI Is the Reason)

  1. Modern experimentation platforms increasingly support automated subgroup analysis, surfacing whether an effect varies meaningfully across pre-specified segments, rather than requiring painstaking manual investigation.
  2. Machine learning methods for estimating heterogeneous treatment effects — sometimes called uplift modeling or causal forests — have matured specifically to identify which subgroups genuinely benefit most (or least) from a treatment, connecting directly to this content library’s causal inference series.
  3. This has enabled more targeted rollout strategies: applying a treatment specifically to the subgroups it genuinely helps, rather than uniformly to everyone based on a single, potentially misleading average.

The Metaphor, Fully Extended

The Field TrialHeterogeneous Effects Concept
A modest average effect across the whole farmA modest overall average treatment effect
A strong positive effect in sandy soil, hidden inside that averageA strong positive effect in one subgroup, hidden inside the overall average
A negative effect in clay soil, also hidden inside that averageA negative effect in another subgroup, also hidden inside the overall average
Applying the fertilizer specifically where the soil actually benefitsApplying a treatment specifically to the segments it genuinely helps

For Beginners: What to Actually Do

  • Practice examining an experiment’s results across a few pre-specified, meaningful subgroups, not just the overall average.
  • Learn the important caveat: exploring many subgroups after the fact risks the multiple testing problem covered in Article 18 — pre-specify subgroups of genuine interest before running the test where possible.
  • Get comfortable with the idea that “no significant overall effect” doesn’t rule out a real, important effect within a specific subgroup.

For Practitioners and Leaders: The Deeper Layer

  • Pre-specify subgroups of genuine business interest before launching a major experiment, rather than fishing through results after the fact.
  • Invest in heterogeneous treatment effect estimation methods for major decisions, where a uniform rollout might genuinely help some segments and hurt others.
  • Build targeted rollout capability so that treatments found to help specific subgroups can actually be deployed selectively, rather than forced into an all-or-nothing decision.

Quick Recap

  • Heterogeneous treatment effects occur when a treatment’s real effect differs meaningfully across subgroups.
  • An overall average effect can mask both genuine opportunities and genuine risks hiding within specific subgroups.
  • Modern platforms and causal machine learning methods increasingly support systematic subgroup analysis.
  • Pre-specifying subgroups of interest avoids the multiple testing risk of exploring too many subgroups after the fact.

Where This Fits in the Series

Article 14 covered the variation an average can hide. Article 15 covers how to make legitimate early decisions about a trial without falling into the peeking trap from Article 8.