Same chapter. Same research prompt. What changed between the first pantry file and the second was forty-five minutes of human evaluation.
Open the pantry file for a chapter on AI's impact on graphic design. The research file claims to have prioritized primary sources. Section One cites a Medium post with three followers — the only primary source it names. The second bullet references a study showing seventy-eight percent of design firms integrating AI: no study named, no year, no methodology. A McKinsey reference exists but with no specific report, no page number, no figure you could find. Section Three, which is supposed to contain examples of graphic design, has none. Marketing examples, a small business logo, educator lesson planning — everything except graphic design. Section Nine asserts that primary sources were prioritized. They were not. The pantry file you are about to hand to your chapter author is authoritative-sounding and false. Your author will read it, draft the chapter, and pass forward something that sounds credible but has no real basis in the domain.
The pantry file will be read by someone who does not know graphic design and has no reason to distrust what the Gatherer produced. They will write the chapter based on it.
Cowork's writing pipeline depends on trust. The Chapter Research Gatherer is a prompt — not separate software, not a person, a prompt — that runs on a long-context model with retrieval. It is genuinely capable. The problem is not what it does. The problem is the transfer of authority. When the Gatherer writes a pantry file, the writer who comes after reads it as gospel. They cite it. They draft the chapter based on it. If the pantry file contains a vague citation, the chapter will cite it vaguely. If it has examples from the wrong domain, the chapter will use those examples. If it claims seventy-eight percent and cites no study, that percentage will appear in the final chapter with the authority of the book behind it. The gap between what the Gatherer produces and what actually stands up to scrutiny is structural. This is what happens when you skip the evaluation stage.
The Gatherer produces a nine-section research synthesis file: Primary Sources, State of the Field, Application Domain Examples, Book's Thesis Connection, AI Wayback Machine Candidates, Pedagogical Delivery Research, Representation and Display Research, Open Questions, Sourcing Notes.
The Chapter Research Gatherer is embedded in the Cowork workflow. You give it the TIKTOC.md file and a chapter specification. For every chapter, it executes a three-step process. First, it reads the capability statement, learning outcomes, bridge question, and application domain from the specification. Second, it consults any shared library files in the pantry directory — files with the library prefix that have already been evaluated. Third, it runs web research using a long-context model with retrieval capability and compiles a nine-section notes file saved as pantry slash research dash chapter XX. The process takes two to three hours per chapter. The nine sections follow Harris Cooper's framework for research synthesis, adapted to the needs of a chapter draft. One file per chapter. The output looks comprehensive. It is the input to what comes next.
Cooper's 1982 framework argued that compressing any stage was the move that hid the work. The Gatherer compresses three stages into one automated prompt. Evaluation and analysis cannot be compressed without losing the judgment that makes research trustworthy.
Harris Cooper published his framework for integrative research reviews in the Review of Educational Research in nineteen eighty-two. Problem formulation: define what you are asking and what counts as evidence. Data collection: find all sources within scope. Evaluation: decide which sources are credible, primary, or derivative. Analysis: identify patterns, contradictions, what is novel versus consensus. Presentation: communicate the synthesis. Cooper's argument was that compressing any stage hid the work. It made a review look authoritative without showing its reasoning. The Gatherer compresses the first three stages — problem formulation, data collection, and evaluation — into one prompt. This is a gain in speed. It is a loss in transparency. The two stages Cooper said cannot be compressed are evaluation and analysis. Those require judgment that the language model cannot be trusted with alone. That is where the human re-enters the process.
All of formulation, collection, and evaluation happen inside the Gatherer prompt. The language model does not expose its reasoning at any intermediate step. You receive the output: a pantry file that looks complete.
A process flow shows problem formulation, data collection, and presentation grouped inside a machine band. A vertical seam marks the machine-to-human hand-off. Evaluation and analysis sit inside a human band. The final output is the draft-ready pantry file. Machine stages are tinted one color. Human stages are another. The implication is that you cannot see what happened inside the machine. You cannot observe the Gatherer's intermediate reasoning. You cannot ask it to show you the search queries it tried first and rejected. You receive only the output — nine sections, dozens of citations, dozens of examples. All of that material emerged from a single prompt run. None of it shows its working. This is not a limitation of the Gatherer. It is a structural property of language models. The model generates one token at a time until it stops.
Retrieval-augmented generation moved citation fabrication from roughly half the time toward something lower and harder to measure. The improvement is real. The risk is not gone.
Since the release of deep research agents between twenty twenty-four and twenty twenty-five, the rate of citation fabrication has declined measurably. Previously, language models would generate citations that sounded plausible but named non-existent papers, misquoted real studies, or invented sources entirely. Retrieval-augmented generation mitigates this. The model is now grounded in actual search results. It is not generating citations from statistical patterns in its training data; it is generating them from documents it has actually retrieved. The improvement is substantial. The risk is not eliminated. There is a structural reason why fabrication persists: language models learn the surface form of citation — the author-year-title shape — without modeling the act of verifying whether the source is real or whether the quote is accurate. Citation-shaped text is cheap to generate. Verified citation is not. Someone must do the verification. The model can generate citation-shaped text grounded in a retrieval result, but it cannot verify that the retrieval result actually supports the claim being made.
Language models can produce citation-shaped text without ever checking whether the source is real or the quote is correct.
A language model trained on millions of books learns what citations look like. It learns the statistical patterns of author names, publication years, journal titles, the syntax of how a quote is embedded in prose. That is surface form — the appearance of a citation. What it does not learn is the semantic act of verifying. The model does not ask: Does this author exist? Did they write this paper? Did the paper contain this quote? Is the quote accurate? Verification requires checking a source you know to be real, comparing text against it, and making a judgment about whether the citation supports the claim. Language models do not do this. They generate tokens. When you ask a model to write research synthesis, it produces text statistically likely given its training data. Citation-shaped text has high probability. Verified citation has the same probability from the model's perspective as fabricated citation, as long as both are citation-shaped. Retrieval-augmented generation constrains the model to match results from actual search. But retrieval grounding does not mean verification. The model can generate a correct citation to a retrieved source while misquoting it or using it to support a claim it does not support.
Same chapter. Same research prompt. The only difference: human evaluation and re-sourcing of credible primary sources.
Now read the same chapter's pantry file after a forty-five-minute evaluation pass. Section One now names the Adobe Firefly Three release notes from March twenty twenty-four with version-specific features sourced from Adobe's own release page. A peer-reviewed paper from the Journal of Design Research on AI-augmented studio workflows. The AIGA twenty twenty-four Design Census with methodology and sample size explicitly stated. Section Three now has five examples: a brand identity system for a regional law firm, an editorial layout for a quarterly magazine, a motion design project, a packaging redesign, a product lockup system. All graphic design. All specific. All names, not abstractions. Section Nine now names what was filtered out: Medium trend pieces, LinkedIn posts, listicles. The Studies-show bullet is gone. It was fabricated. Someone noticed. Someone removed it. Same chapter. Same TIKTOC.md. Same Gatherer run. The difference is what happened after the script finished.
The evaluation pass filters aggregator sources and vague citations, replacing them with primary sources and domain-specific examples that a skeptical reader could actually verify.
What makes the evaluation pass different from the initial Gatherer run is that it requires domain knowledge and skeptical judgment. When you read the Gatherer's pantry file for a graphic design chapter, you must know what graphic design is. You must recognize that marketing examples are not graphic design examples. You must know who in the field would be trustworthy. You must be willing to say: this section claims to have sourced primary sources, but it has not. I will find primary sources. This means you might search for the Adobe Firefly Three release notes yourself. You might know of the AIGA Design Census and think to check whether it exists. You might recognize that a Medium post with three followers is not a primary source. You might notice that the section says Studies show seventy-eight percent and then provides no study. You will cross-check. You will filter. You will replace. This is labor. It is also the labor that makes the pantry file worth passing forward to the chapter author.
The Gatherer produces a draft pantry file. Your evaluation pass produces one that is trustworthy. The difference between those two files is not cosmetic editing. It is the core work of research synthesis itself.
Harris Cooper described his framework as if evaluation and analysis were sequential steps that came after data collection. In practice, evaluation begins the moment you start reading sources. You notice that one source is a peer-reviewed journal and another is a blog. You notice that a claim has a citation and another claim has Studies show with no study attached. You notice that one author is cited in dozens of subsequent papers and another has published once. This noticing is not separate from collection; it is contemporaneous with it. The Gatherer compresses all of this into one black-box process. The evaluation pass reverses that compression. You slow down. You read each source. You ask whether it is primary or derived. You ask whether the claim the Gatherer is making is actually what the source says. You ask whether the example is in scope for the chapter. This is slow work. It is also the work that prevents an author from confidently drafting a chapter on graphic design with no graphic designers cited.
Language models have real power for research synthesis. But power without verification is hallucination wearing a bibliography. The pantry file is not a draft and not a citation list. It is the only thing standing between the chapter author and an authoritative-sounding lie.
The temptation is to see the evaluation pass as an editing phase that comes after the real work is done. It is not. The real work is evaluation and analysis. The Gatherer does the fast parts — formulation and collection — and gives you a summary that looks complete. That summary looks like research. But the summary is, by design, a compressed version of a process that should have stopping points for judgment. The stopping points are where you re-enter. You read the pantry file. You check whether the sources are real. You check whether they are primary. You check whether they actually support the claims. You check whether the examples are in domain. You replace sources that fail these checks. You remove claims that cannot be verified. You add examples that your domain knowledge tells you should be there. This is not cosmetic. This is not optional. This is the research synthesis. The Gatherer produced the raw materials. You are producing the synthesis itself.
Series · Chapter 6 · Research Pass: Pantry Population
The pantry is not a draft and it is not a citation list. It is the only thing standing between Cowork and an authoritative-sounding lie.