The visual fluency trap arrives when an image model produces a figure that looks like professional work but cannot be traced through by its own author.
A designer types a prompt into an image generator asking for a diagram of a pipeline. The result arrives in seconds: gradient-filled boxes, soft drop shadows, fourteen labeled nodes, three legends, decorative gears, a looping arrow. It is polished. It looks professional. But when the author tries to trace the actual path a manuscript takes through the pipeline, they cannot. The arrows go everywhere. Two labels contradict the source text. One word is misspelled. The figure is not wrong, exactly. It is full. And a figure that is full has made no decisions. This is the visual fluency trap at the level of the picture: something that reads as correct but means nothing, and carries more authority than it deserves because of its polish.
The image carries more authority, which makes imprecision in figures more dangerous than imprecision in text. The stakes of deciding what to show are higher because readers assume the picture is right.
A reader will distrust a clumsy paragraph and give a polished diagram the benefit of the doubt. This asymmetry is the danger. Images carry more authority than sentences. The visual polish that marks a professionally generated figure becomes permission to stop thinking critically. When a figure is syntactically correct—all the arrows connect, all the labels are spelled right—a reader assumes it is semantically correct too. But syntax and semantics are independent. A diagram can be perfectly formatted and completely wrong about the relations it claims to show. Understanding this stakes the entire problem: generating a figure is now free, and free is the whole problem. The real work is not how to make a machine produce a figure. The work is deciding what a single figure is allowed to contain, and how to make a machine that wants to give you everything give you only what matters.
George Miller's 1956 paper became the most cited work in the history of psychology, establishing a foundational constraint on human cognition.
In 1956 George Miller published what would become the most cited paper in the entire history of psychology. Its title was "The Magical Number Seven, Plus or Minus Two," and it fixed in everyone's mind a simple claim: working memory holds about seven items. That number became canonical. It appears in textbooks, in design guidelines, in software architecture decisions. Seven became the figure everyone knew. But this number, while foundational, turned out to be optimistic. Miller's original experiments were looser than they appeared. His successors would tighten the methodology and revise the estimate downward. Still, Miller's core insight was correct: the human mind has a hard ceiling on how many independent things it can hold in active thought at once. That ceiling is not a soft guideline or a design preference. It is a fact about neurology.
Nelson Cowan's 2001 review of decades of tighter experiments revised the true working memory capacity downward, establishing four independent chunks as the realistic cognitive limit.
Nelson Cowan spent decades reviewing tighter and more carefully controlled experiments on working memory. His 2001 synthesis found that Miller's estimate was too generous. The real capacity, measured under conditions that eliminate confounds, is closer to four chunks. Not four thousand. Four. Four independent things a mind can hold in active thought and relate to each other simultaneously before something falls out. This is a crucial distinction: Miller was measuring the number of items in a longer-term memory buffer, where rehearsal could extend capacity. Cowan was measuring the number of chunks that can be mentally manipulated at once without conscious effort. The difference matters because a figure is not something a reader memorizes. A figure is something they parse in real time while looking at it. They must hold the pieces in working memory and see how they relate. Cowan's four is the constraint that governs what a figure can contain.
This is not about displaying information. This is about enabling a specific cognitive action: the simultaneous awareness of how parts connect.
What does a figure do? Not display. Not illustrate. Not make complex things look simple. A figure makes a cognitive commitment. It asks the reader to hold its parts in working memory all at once and see how they relate. That is the entire job. That is what a figure is for. If the figure shows four interacting components, a reader can hold them. They can see the relationships between them. If it shows fourteen components, the reader holds the first four, loses them while fetching the next four, and walks away with an impression of having understood something—which is worse than confusion because confusion at least knows itself. The impression is false confidence. The reader thinks they have learned the relations when they have only seen a mosaic of disconnected pieces. This is why the constraint is not cosmetic. It is not a design preference or a stylistic choice. It comes directly from the architecture of cognition. A figure that exceeds working memory capacity fails its only function, which is to show relation.
This theory moves the constraint from a static fact to a dynamic principle: every pixel that doesn't serve learning actively harms it.
John Sweller built an entire instructional theory on this principle: cognitive load theory. The core insight is simple and devastating. The load a learner can carry is fixed and small. That load is working memory. Everything that competes for attention inside working memory while the learner is trying to understand something is either teaching or it is harming. There is no neutral. A gradient background is not free. It costs capacity. The decorative gears that someone added to make the diagram look more professional did not merely fail to help. They consumed capacity that the reader needed to understand the arrows. Sweller's framework turned this from a psychological curiosity into an engineering principle: every element in a figure must justify its existence by teaching something. If it is there for completeness, for showing how much you understand, for making the work look professional, then it is stealing from the reader. This is the mechanism that connects working memory capacity to figure design.
Every element you add to show how complete your understanding is, you subtract from the reader's ability to follow it. The figure containing your whole mental model teaches nothing.
This is the trade-off named plainly, because every figure decision in your career will be made against this choice. Comprehensiveness versus comprehension. On one side: the drive to show everything you know, to document the full complexity, to prove you understand. On the other side: the reader's need to understand one thing clearly. These are in direct opposition. The more complete your figure becomes, the less comprehensible. The more you add in the name of accuracy and completeness, the more you subtract from the reader's working memory, which is already running at capacity just tracking the main relations. The figure that contains your entire mental model teaches nothing. The figure that contains one relation, shown cleanly and completely, teaches that relation. This is not a weakness of the figure or a limitation of the reader. It is the structure of how humans learn. Understanding happens through focus. Mastery happens through building from clear pieces.
When a concept genuinely has more than eight moving parts, the answer is not a smaller font and a bigger canvas. The answer is two figures.
Here is the operative rule across the entire craft. It is a ceiling. Six to eight labeled components per figure, and never more. This is not a stylistic preference. This is not a starting point that you can negotiate or exceed if the subject is complex. It is a working-memory budget grounded in cognitive science. Cowan's four-chunk limit would suggest the ceiling should be four. But figures operate in a domain where readers have external support: they can look back at the figure while reading, refresh their working memory, trace a path multiple times. This extended interaction slightly loosens the constraint. Six to eight emerges as the practical ceiling that acknowledges Cowan's research while accounting for the realities of how people actually read. Every element above that number increases cognitive load without increasing comprehension. You are competing for eight slots of working memory. Every component you add takes one. At eight, all slots are full. At nine, something gets pushed out. The slot that gets pushed out is usually the element the reader most needed to hold—the relation itself.
The cleanest criterion for splitting is Cowan's number applied directly: past four interacting components, and you need a second figure.
Knowing when to split is itself a decision with criteria. The cleanest test applies Cowan's four-chunk limit directly to the problem at hand. When you find yourself designing a figure with more than four distinct interacting components, you are probably past one figure. The components do not need to be visual elements. They are the conceptual units the reader must track: agents, forces, outcomes, scales, timeframes, anything that plays an independent role in the relation. Count them honestly. Can a reader hold all four in mind while understanding how they relate? If yes, one figure works. If you are at five, split. The cleanest way to split is to ask: can I make two figures such that each one teaches a single relation clearly, and together they build understanding? If the answer is yes, that is the division. This is where the reductive instinct proves correct. The instinct to simplify, to separate, to force clarity—these are not compromises. They are the path to teaching. When in doubt about whether a figure should be one or two, the figure wants to be two figures.
Both patterns multiply the relations a reader must track. Branching creates path multiplication. Crossing scales requires continuous translation. Both demand separate figures.
There are two structural patterns that always demand splitting, because they multiply the cognitive load in specific ways. The first is branching: parallel paths, competing outcomes, a policy that hits multiple systems at once (economic and legal and social). When a figure has branches, it is not showing one relation. It is showing multiple relations that diverge from a common point. Each branch is its own path the reader must track. At the moment of the split, the reader's working memory splits too. They must hold the context of both branches simultaneously to compare them. This doubles the load. The second pattern is crossing scales. When a concept spans from the individual level to the institutional level to the societal level, or from the short run to the long run to the very long run, a single figure forces the reader to translate continuously between frames. Individual agents have one set of constraints. Institutions have different constraints. Society has its own. Holding all three in mind while watching how they interact is a load that exceeds capacity. The answer is two figures, each at one scale, with explicit text showing how they connect.
CAJAL is not a tool for generating figures. It is a tool for deciding what to leave out—for enforcing working memory constraints at the level of design.
The skill that governs all of this decision-making has a name in the toolchain used for this book. It is called CAJAL. CAJAL is not tied to this textbook. It is a general method that works the same way on a signaling cascade as it does on a treaty structure, a monetary transmission mechanism, or a pipeline diagram. The fundamental problem CAJAL solves is this: how do you make a machine that wants to give you everything give you only that which matters? CAJAL provides a process for deciding, at each step, what a figure is allowed to contain. The full command set and the SVG Style Guide are documented in Appendix I. The point here is that CAJAL embodies all the constraints we have discussed—the cognitive load theory, Cowan's four-chunk limit, the working-memory budget—and operationalizes them into a design procedure. Learn CAJAL once and it transfers. It transfers to every book in the series. It transfers to any figure you ever commission again. It is the craft that sits beneath figure generation.
Figure generation is free now. What is scarce is clarity. Clarity comes from constraint, from ruthlessly eliminating anything that doesn't serve the one relation you are trying to teach. The figure that looks most professional is often the one that has made the fewest decisions.
The epigraph of this chapter asks: why is the hard part of a figure deciding what to leave out? The answer is now complete. Generation is free. A designer can type a sentence into an image model and have a complete diagram in seconds. The model will give you everything—all the arrows, all the labels, all the ornaments. It will give you completeness. What it will not give you automatically is clarity. Clarity requires decision. It requires the authorial act of saying: this component matters, this one does not. This relation is central, this one is secondary. We will show four things here and save the rest for a second figure. The cognitive science is unambiguous. Cowan's four. The working-memory budget. The trade-off between comprehensiveness and comprehension. These are not opinions. They are facts about how minds learn. The figure that has made these decisions ruthlessly is the figure that will teach. The figure that has made no decisions—that has tried to show everything—has made the mind's job harder. The hard part is not the generation. The hard part is the exclusion. And exclusion, in a culture that mistakes polish for mastery, requires judgment.
AI 1 · Chapter 11 · Creating Figures