An equation is symmetric. A causal claim is asymmetric. Statistics can capture the first but not the second.
Suppose you fit a linear regression and get Y equals 0.6 times X plus 1.2. The slope is 0.6, the fit is good, and you want to make a decision about how to intervene. So you say: X causes Y with a coefficient of 0.6. This is where we go wrong. The regression equation does not know which variable causes which. If you had regressed X on Y instead—fit the equation backward—you would get a different slope and different intercept. Both equations describe the same correlation in the same data. Neither tells us anything about causation. The equation is symmetric; causation is not. This fundamental asymmetry is the entire problem of causal inference. Statistics knows correlation. Statistics has never known cause.
If you adjust X because you think it causes Y, but you are wrong about the direction, your decision will backfire. The ability to distinguish cause from correlation determines whether your actions have the intended effect.
People who want to make a decision almost never stop at the equation. They want to say: X causes Y, so if I change X, I can change Y. But the data do not tell them this. The equation is silent on causation. Suppose hot weather causes both ice cream sales and drowning deaths. If you observe the correlation and infer that ice cream causes drowning, you might ban ice cream sales at the beach. The correlation is real. The causal claim is wrong. Your intervention fails. This is not a statistical curiosity. It is the difference between effective action and useless action. To know what will happen when you intervene, you must know the structure of causation. The equation cannot give you this knowledge. You need something else. You need a diagram that makes the causal theory explicit.
The same data produce both quantities. But they answer different questions: one asks about association, the other about causation.
In the 1920s, Sewall Wright, a geneticist at the U.S. Department of Agriculture, made the distinction mathematical. He pointed out that the regression coefficient and what he called the path coefficient are different objects, even when computed from the same data. The regression coefficient does not care which variable is the cause. It is symmetric. The path coefficient is asymmetric. It points from cause to effect. It can be different in different directions. Wright invented a formal notation for this asymmetry. He drew a diagram with arrows. X points to Y means X is a cause of Y. The diagram cannot be reduced to an equation, because the equation is symmetric and the diagram is not. The arrows carry information the equation cannot. This notation—the causal diagram, the directed acyclic graph—is the foundation of everything in causal inference.
The diagram is a language for encoding theories about causal structure. Two researchers can disagree about which diagram is correct—that is a scientific disagreement made precise.
A causal diagram is a set of nodes connected by arrows. Each node is a variable. Each arrow points from a cause to an effect. Acyclic means there are no closed loops. You cannot start at a node, follow the arrows, and return to where you started. This is a minimal structure, but it encodes a tremendous amount of information. The graph is a language. Each diagram is a sentence in that language, encoding a specific theory about the structure of cause and effect in the system. It is precise enough that two researchers can disagree about which sentence is true. That is a substantive scientific disagreement. What the graph buys us is the ability to make the disagreement precise. Each researcher can draw their diagram, and we can examine what each implies about the data we should observe. Before the diagram, the disagreement was buried in prose. After the diagram, it is visible.
If the effect of X on Y flows entirely through another variable Z, we draw X→Z→Y instead of a direct arrow. The diagram must match the mechanism, not just the correlation.
When we draw an arrow from X to Y, we make a structural commitment. We claim that X is a direct cause of Y in the system we are modeling. Direct has a specific meaning: not mediated by any other variable in the diagram. If the effect of X on Y flows entirely through another variable Z, we do not draw a direct arrow. Instead, we draw two arrows: X points to Z, Z points to Y. The diagram must represent the actual mechanism. This is not a choice about convenience. It is a claim about what is actually happening in the system. Different diagrams imply different conclusions about interventions, about confounding, about which correlations are spurious. Get the diagram wrong and your causal inference will be wrong. Get it right and you can read off what the data should tell you.
A diagram with arrows everywhere claims nothing. A diagram with most arrows missing claims a great deal. Every line you do not draw is a hypothesis.
Missing arrows are claims too. When we do not draw an arrow from X to Y, we commit to the claim that there is no direct causal effect of X on Y. This is often the more important commitment. Consider two diagrams. In the first, every variable connects to every other variable—a complete graph. This diagram claims nothing. It is consistent with any data pattern. In the second, most arrows are missing. This diagram is specific. It rules out patterns. It makes falsifiable predictions. The second diagram is scientifically useful. The first is not. Every line you do not draw in your diagram is a hypothesis. You are claiming that removing that causal pathway would not change the structure of the system. This claim may be wrong. Your theory may be incomplete. But until you specify which arrows are absent, you cannot reason clearly about causation.
The diagram X→Z→Y is very different from X→Y. In the first, the effect is entirely mediated by Z. In the second, X has a direct effect not explained by any variable in the model.
When we draw an arrow directly from X to Y, we claim there is a direct causal effect. But causation can also flow indirectly. Suppose X causes Z, and Z causes Y. Then X affects Y, but the effect is mediated by Z. The path X→Z→Y is an indirect causal path. If we draw a direct arrow from X to Y in the diagram, we commit to the claim that X affects Y in a way not captured by Z. This distinction matters enormously for inference. If we condition on Z—that is, if we look only at cases where Z has a fixed value—the indirect path is blocked. Only the direct effect remains. If we draw X→Y directly when in truth there is only an indirect path through Z, we will make mistakes when we condition on Z. We will see the effect vanish and conclude that the diagram is wrong. But the diagram was wrong from the start. It did not match the mechanism.
The data are equally consistent with all three. Only causal knowledge from outside the data can tell us which story is true. This is exactly what a DAG encodes.
The classic example of a spurious correlation is the relationship between ice cream sales and drowning deaths. Both rise in summer; both fall in winter; the correlation is real and quite strong. No competent statistician would conclude that ice cream causes drownings. But the data, on their own, do not rule it out. There are at least three causal stories that produce the same correlation. In story one, ice cream sales cause drownings—perhaps because cold ice cream gives swimmers cramps. In story two, drownings cause ice cream sales—perhaps because tragic news prompts comfort eating. In story three, some third variable causes both. Hot weather causes more people to swim and therefore drown, and also causes more people to buy ice cream. The first two stories are absurd. The third is obviously correct. But how do we know? The data are equally consistent with all three. Our knowledge that hot weather affects both swimming and ice cream consumption is causal knowledge. It comes from outside the data. It is exactly the kind of knowledge a DAG encodes.
Heat is a confounder of the ice cream–drowning relationship. Once we account for heat, the correlation between ice cream and drowning should vanish. The diagram predicts this; the data alone do not.
In the ice cream example, heat is a confounder. It causes both swimming and ice cream consumption. When a third variable causes both the exposure and the outcome, they become correlated even if there is no causal connection between them. This is confounding. Once we draw the diagram and see that heat points to both variables, we can reason about what to do. If we intervene on ice cream sales—handing out free cones at the beach—would we increase drownings? The diagram says no. Ice cream is not a cause of drowning. Heat is. The correlation between ice cream and drowning is spurious. Remove the confounder, or condition on it, and the spurious correlation should disappear. The data alone cannot tell us this. The regression coefficient for ice cream predicting drowning will be positive with or without the confounder. But the interpretation changes. With heat in the picture, we know that coefficient is not a causal effect.
The DAG, on its own, does not tell us how strong the causal effects are. To get magnitudes, we add functions or path coefficients. That is the move from DAG to structural causal model.
A DAG tells us the structure of causation but not its strength. It says X causes Y, but not by how much. To get magnitudes, we must add something more. In the linear case, we add path coefficients—numbers that quantify the effect of one variable on another. More generally, we add functional relationships. This is the move from a directed acyclic graph to a structural causal model, or SCM. The SCM combines the graph with the functions. It specifies both the structure and the magnitudes. With an SCM, we can predict what will happen when we intervene. We can say not just that X affects Y, but by how much. The graph is the language of causation. The SCM is the language of causation with numbers. Together, they give us what statistics alone cannot: the ability to reason about what we would observe under different interventions, not just what we have observed in the past.
Statistics captures correlation and association. It has no method for detecting causation. You must bring causal knowledge—from science, from mechanism, from domain expertise—to the data. The DAG is how you encode that knowledge precisely. Once encoded, you can examine what the data should tell you under each possible causal model. The graph makes your theory explicit. The data either confirm it or refute it. Without the graph, your theory is hidden in prose. With it, your assumptions are visible, your disagreements precise, your inferences sound.
This is the central insight of this chapter and the foundation of causal inference. Statistical methods are symmetric. They look at correlation and association. They cannot, by themselves, distinguish between X causing Y and Y causing X. They cannot detect confounding. They cannot tell you what will happen when you intervene. All of these things require knowledge from outside the data. You must know the causal structure before you can ask what the data mean. Sewall Wright showed that this knowledge cannot be extracted from the correlation alone. Causation requires asymmetry, and asymmetry requires a language that the equation cannot provide. The directed acyclic graph is that language. It encodes structural commitments—claims about which variables cause which others, and which variables do not directly cause which others. Once you draw the diagram, you can reason about what you would observe under intervention. You can detect confounding. You can trace indirect effects through mediators. You can make causal inferences. Without the diagram, you cannot.
Living Models · Chapter 6 · Graphs That Think
You now understand the distinction between correlation and causation, and you have the notation to make that distinction precise. You know what a directed acyclic graph encodes: the structural claims about which variables cause which others. You know that regression coefficients are symmetric and path coefficients are asymmetric. You know that missing arrows are claims too, and that confounding arises when a third variable points to both the exposure and the outcome. You know that interventions change the data in ways that correlations alone cannot predict. You know that causal knowledge is always external to the data. And you know the notation—the DAG—that lets you make your causal theory explicit. This is the language on which the rest of causal inference rests. Everything that follows builds on this foundation.