This is not an exotic edge case. It is the default state of organizational intervention.
Randomized controlled trials are the gold standard of causal inference. When you randomize treatment, you sever every backdoor path—every common cause that could confound your estimate. The method is elegant and powerful. But there is a hidden assumption buried inside every RCT: SUTVA, the Stable Unit Treatment Value Assumption. SUTVA has two components. First, no interference—the treatment of one unit does not affect the outcomes of other units. Second, no hidden variations—there is only one version of the treatment, not multiple versions that different units receive. In observational data and simple laboratory settings, these assumptions often hold. But in organizations, in policy rollouts, in technology deployments, they almost never do. A training program travels through an instructor who learns from running it. A pricing change is visible to competitors who adjust their own prices. An employee technology rollout creates discussion, imitation, and social pressure. The very properties that make interventions scalable and deployable—integration with existing infrastructure, visibility across social networks, spillover from early adopters to late ones—are the properties that violate SUTVA. Understanding this gap is not an attack on randomization. It is what lets you use randomization correctly and communicate what a trial result does and does not guarantee.
All three methods fail when a confounder exists that is unmeasured. Sensitivity analysis tells you how large the unmeasured confounder would have to be to change your conclusion, but it does not tell you the confounder does not exist.
Chapter 10 ended in a hard place. Observational methods have become sophisticated. Backdoor adjustment identifies causal effects when you condition on the right set of confounders. Instrumental variables exploit sources of random variation that affect treatment but not outcomes directly. Double machine learning combines machine learning with causal inference to handle high-dimensional settings. Yet all of these methods share a fundamental vulnerability: they require that you measure every confounder. If a variable affects both treatment and outcome, and you cannot measure it, no amount of methodological sophistication eliminates the bias it introduces. Sensitivity analysis quantifies how robust your conclusions are to unmeasured confounding—how large an unmeasured confounder would have to be to reverse your findings. But this is not the same as knowing the confounder does not exist. The randomized controlled trial offers an escape. It does not require measuring confounders. It does not require identifying a sufficient adjustment set. It addresses confounding structurally, at the design stage, by making treatment independent of every potential confounder, measured or unmeasured.
When backdoor paths exist, the observed association P(Y|X) conflates the causal effect with false association created by shared causation. These two paths cannot be separated in observational data.
A confounder is a variable that causally affects both treatment and outcome. In causal graphs, confounders generate a specific structure: they appear as common causes that create two distinct types of dependence between treatment and outcome. The first path runs forward from treatment to outcome—this is the causal mechanism you want to study. The second path runs backward through the confounder, creating spurious association. This backward path is called a backdoor path because it does not represent the true causal effect. It is false association created by shared causation. When you observe the relationship between treatment and outcome in data where randomization has not occurred, you observe both paths simultaneously. The observational association P(Y|X) includes both the genuine causal effect and the spurious effect from the backdoor path. You cannot separate them. The backdoor criterion, developed in prior chapters, tells you when you can eliminate this problem by conditioning on a measured adjustment set—variables that block all backdoor paths. When no such set exists because the confounder is unmeasured, identification from observational data fails completely.
The backdoor criterion is powerful when confounders are measured. When they are not, no statistical sophistication overcomes the problem.
The backdoor criterion provides a precise formal rule for identifying which variables you need to condition on to block all backdoor paths and recover the causal effect from observational data. The criterion is elegant: if a sufficient adjustment set exists and you measure it, you can identify the causal effect without special assumptions. But the criterion has a hidden cost. You must measure every variable in the adjustment set. You cannot approximate a confounder with a proxy. You cannot estimate it from other variables. It must be directly observed in your data. When a true confounder is unmeasured—when it exists in reality but does not appear in your dataset—no adjustment set can block the backdoor path it creates. Your treatment effect estimate remains biased. This is not a limitation of the statistical method or the analyst's skill. It is a limitation imposed by the data and the problem structure itself. No amount of advanced methodology can overcome an unmeasured confounder. This is the hard limit that Chapter 10 reached, and it is the problem that randomization is designed to solve.
Randomization applies the do-operator formally: it severs every arrow pointing into the treatment node, eliminating backdoor paths at the source, not merely blocking them by conditioning.
Randomization works through formal structural intervention. When you randomly assign treatment to units, you are mathematically applying the do-operator, which severs the causal mechanism producing treatment in the original observational system. In observational data, treatment arises from many sources—some measured, some unmeasured. These sources create arrows pointing into the treatment node in the causal graph. These arrows are the foundation of backdoor paths. Randomization removes these arrows entirely. The treatment is now caused only by the randomization device, not by any of the original determinants of treatment. The modified causal graph has no backdoor paths. X becomes independent of all confounders. As a result, the observed difference in outcomes between treated and control groups estimates the average treatment effect—the true causal effect—without requiring any further adjustment or conditioning. This is why randomization is so powerful and why it is the gold standard. It does not require measuring confounders. It does not require identifying a sufficient adjustment set. It does not depend on assumptions about unmeasured variables. It addresses confounding structurally, before any data is collected.
SUTVA stands for Stable Unit Treatment Value Assumption. When SUTVA holds, each unit's potential outcome depends only on that unit's own treatment status, not on others' treatments or which version of the treatment they receive.
Randomization severs backdoor paths and eliminates confounding, but it is built on an assumption that is often unstated and frequently violated: SUTVA, the Stable Unit Treatment Value Assumption. SUTVA decomposes into two separate, independent components. The first is no interference, also called the assumption of no spillover. The potential outcome for unit i must not depend on the treatment assigned to unit j. What happens to one unit does not causally affect what happens to another unit. The second component is no hidden variations of treatment. What we call the treatment is a single, well-defined intervention. Every unit assigned to treatment receives exactly the same treatment. There are not multiple versions that different units experience, and treatment is not divided into hidden attributes that are allocated separately from the explicit randomization. In a simple laboratory experiment—subjects in a room receiving a pill or a placebo—both components of SUTVA are often reasonable. But in organizational settings, in policy rollouts, in any real-world context where interventions are embedded in social and institutional infrastructure, SUTVA is fragile and frequently violated.
When unit i receives treatment and unit j observes it, learns from it, or adapts to it, interference has occurred. Unit j's potential outcome becomes dependent on unit i's treatment.
No interference means that the treatment assigned to one unit does not causally affect the outcomes of any other unit. But in organizational contexts, interference is not exceptional—it is endemic. A training program does not remain isolated to the person trained. The instructor learns from running it. Trained employees teach untrained colleagues. Knowledge diffuses through networks. A pricing change is not invisible. Competitors observe it and adjust their own prices. Customers notice and compare. A technology rollout is not silent. Early adopters use it in ways that later users observe and imitate. Peer effects amplify adoption. Network effects multiply impact. In the work of Esther Duflo on organizational interventions at scale, these forms of interference—knowledge spillover, behavioral adaptation, competitive response, peer learning—are not exceptions. They are the dominant mechanisms through which interventions create impact in real organizations. But each instance of interference violates SUTVA. When unit i's treatment affects unit j's outcome, and unit j is part of your trial, you cannot separate the effect of unit j's own treatment from the spillover effect of unit i's treatment. Your treatment effect estimate conflates direct effects with indirect effects caused by interference.
The second SUTVA component requires not just one treatment definition, but uniform implementation across all treated units. When implementations differ systematically, there are hidden variations, and SUTVA fails.
The second component of SUTVA requires that there is one version of the treatment and no hidden variations. But in organizational interventions, the gap between the protocol and the implementation is vast. A training program is designed to be delivered in a specific way, but instructors interpret it differently, emphasize different modules, and adapt it to their local context. A new policy is written clearly, but managers implement it with varying strictness and adapt it to their own circumstances. A technology is deployed with documentation, but different teams onboard with different levels of effort and commitment. Some employees resist adoption; others become local champions. These variations in how treatment is actually delivered and received create fundamentally different versions of the treatment across units. A unit assigned to treatment does not receive the same treatment as another unit assigned to treatment. This violates the no hidden variations assumption. The problem is acute because implementation variations are often correlated with outcomes. The team that implements most enthusiastically is likely the team that would have performed well anyway. The employee who adopts the technology most fully is the employee most predisposed to change. When you estimate the treatment effect, you are not isolating the effect of the treatment itself; you are estimating the effect of treatment plus the effect of the kind of unit that receives it most thoroughly.
Each design trades statistical power for credibility under SUTVA violations. The choice depends on which component is most likely to fail in your context.
When you recognize that SUTVA will be violated in your study, you can adjust the trial design to make the violation less damaging to your causal inference. Cluster randomization assigns intact groups—teams, departments, locations, schools—to treatment or control rather than randomizing individuals. By isolating treated clusters from untreated clusters, you reduce the direct interference between individual units. A team's treatment does not easily travel to a different team. The cost is statistical power. You have fewer independent units. Cluster sizes are small relative to the number of clusters, and estimating treatment effects requires substantially larger sample sizes. Time-staggered rollout deploys treatment in phases rather than all at once. You randomize the timing of deployment. Early cohorts receive treatment first; later cohorts serve as controls initially. You measure effects before widespread spillover and imitation spread the treatment effects through untreated units. As time passes, later cohorts observe earlier cohorts, contaminating the control group, but by then you have your primary estimates. The third approach is encouragement design. Instead of randomizing treatment itself, you randomize the offer, the incentive, or the encouragement to take it up. Some units are offered treatment; others are not. Compliance is not forced or ensured. This design allows you to separate the effect of being offered treatment from the effect of actually taking it up.
A training program shows a 15% productivity gain in a randomized trial of fifty workers in one plant. When deployed across five thousand workers across twenty plants, gains are smaller and uneven. The trial measured effects in isolation. Deployment happens inside a connected system where spillover is the mechanism of impact.
The randomized controlled trial is extraordinarily powerful at identifying causal effects when its assumptions hold. When SUTVA is satisfied, the average treatment effect estimated from a trial is unbiased. The estimate translates directly to the real-world effect. But the moment you move from the trial to organizational deployment, SUTVA violations become dominant forces. In the trial, you carefully isolated treated and control units. You prevented cross-group communication and observation. You ensured uniform implementation through strict protocols. You measured outcomes before significant spillover could accumulate. In the deployment, none of that structure persists. Treated and untreated units work together. Managers see effects in early sites and adjust how they implement in later sites. Employees learn from peers and adopt from social proof. Competitors observe and respond. Spillover is not an exception; it is the deployment itself. A common empirical finding is that trial effects exceed deployment effects. The trial showed what treatment can do when carefully isolated. The deployment showed what treatment does when embedded in a realistic organizational system with interference, adaptation, and heterogeneous uptake. This is not a failure of randomization or a limitation of the trial method. It is a natural consequence of SUTVA violations. The trial measured something true about the treatment, but something narrower than what most organizations need to know.
When you report trial results, you are reporting the causal effect of treatment in an isolated, controlled context. When you deploy that same treatment in an organization, you are implementing it in a connected, adaptive, social context. The gap between these two estimates is where most real organizational interventions live.
Randomization is the most powerful causal inference tool available. It works by severing all arrows pointing into the treatment node, eliminating every possible backdoor path and making treatment independent of every confounder, measured or unmeasured. This is a profound achievement. But randomization does not solve every causal inference problem. It does not prevent one unit's treatment from affecting another unit's outcomes. It does not eliminate variations in how treatment is implemented or received across units. These are separate problems with separate solutions. Cluster randomization, time-staggered rollout, and encouragement designs are tools for auditing and mitigating SUTVA violations. But they do not eliminate the core tension between trials and deployments. The trial, by definition, is an isolated, controlled, carefully managed environment. The deployment, by definition, is an embedded, adaptive, interconnected system. A training program that works in a trial with careful implementation and no spillover may not produce identical effects when rolled out across an organization where instructors adapt, employees imitate each other, and managers adjust based on what they observe. This is not a criticism of the trial. The trial identified something real and valuable: the causal effect of the treatment under controlled conditions. But organizations need to know the effect under realistic conditions. Bridging this gap requires not just a trial, but a theory of how trial results translate to deployment contexts.
Living Models · Chapter 11 · Treatments
The randomized controlled trial is the gold standard of causal inference because randomization eliminates confounding structurally. But this strength is conditional on SUTVA. When SUTVA is violated—when interference and implementation variations are present—the trial result does not directly translate to organizational deployment. The strongest practice is not to choose between randomization and SUTVA auditing. It is to do both. Run the trial. Randomize. But simultaneously audit SUTVA rigorously. Identify which component—no interference or no hidden variations—is most likely to fail in your context. Choose a design that mitigates that failure. Cluster randomization for direct interference, time-staggered rollout for learning and spillover, encouragement designs for compliance heterogeneity. Measure spillover effects empirically. Document implementation variations. Collect data on how the treatment spreads and how it changes as it spreads. Build a model of the trial-deployment gap. This approach transforms SUTVA from a hidden assumption into an auditable, visible part of the study design. It makes the strengths of randomization explicit—you know exactly what confounding bias it eliminates. It makes the limitations explicit—you know exactly which spillover and variation problems remain. That clarity is what transforms a randomized trial from an isolated result into a credible, actionable input for organizational decision-making.