The team in Australia's Goulburn-Broken Catchment built a 30-node Bayesian network through structured expert elicitation alone — then used it to guide real management decisions.
In the early 2000s, an ecology team faced a management problem that modern data science handles poorly. Fish populations are expensive to monitor, historical records had gaps measured in decades, and the interventions they were considering had never been attempted at the scale under consideration. There was no outcome data from these interventions. What the team had instead was expertise: six ecologists with collective decades of fieldwork and academic study. Over twelve weeks, a facilitator worked with these six experts through a structured elicitation process. They identified roughly thirty variables — not every variable the experts initially wanted, but the smaller set that mattered most for the management question. They debated directionality, specified probabilities not as point estimates but as ranges with confidence levels, and tested the model against scenarios the experts knew deeply. What emerged was a working probabilistic model that could run forward to predict consequences of interventions, identify relative importance of drivers, and update as new monitoring data arrived. The catchment authority used it to prioritize among candidate interventions. Twelve weeks. Six experts. No machine learning. One causal model that supported consequential decisions.
This is the operating environment for most complex systems in medicine, ecology, and policy. Data science alone cannot handle it. Expert knowledge becomes mathematically necessary.
The Goulburn-Broken Catchment case exemplifies a category of decision-making problem that data-driven approaches alone cannot solve. The data were thin — not missing or absent, but genuinely sparse. Fish populations are difficult and expensive to monitor. The historical record had gaps measured in decades. The interactions among variables were poorly understood: habitat quality, flow regime, predation, water temperature, and dozens of other factors all feedback into one another in ways nobody had fully mapped. Most critically, the interventions under consideration had never been tried at the scale being contemplated. There was no historical precedent, no outcome data from large-scale environmental flow adjustments or riparian restoration at catchment scale. Yet decisions had to be made. The catchment authority could not simply wait for better data. Resources were finite. Management choices would have real consequences for native fish populations and ecosystem health. Stakeholders needed explanations they could understand and believe. Under these constraints, traditional data-driven analytics offers no clear path forward.
This is not a practical limitation that better algorithms will eventually overcome. It is a mathematical fact that applies to any finite dataset in any domain.
Chapter 7 established the mathematical foundation: Markov equivalence. Different causal graph structures can be consistent with identical observational data. The correlations, conditional independencies, and statistical patterns may be identical even when the underlying causal structures are fundamentally different. Two variables X and Y may be correlated because X causes Y, or because Y causes X, or because a third variable Z causes both. All three scenarios can produce the same statistical footprint in a finite dataset. Machine learning algorithms cannot resolve this ambiguity by finding patterns in the data, because the patterns are genuinely indistinguishable. No amount of additional data makes one structure more obvious than another — Markov equivalence holds regardless of sample size. This is what makes expert judgment mathematically necessary rather than merely helpful. The expert brings knowledge about mechanism: which causes which. This knowledge exists outside the statistical patterns in the data. It comes from domain understanding, prior research, physical constraints, and repeated observation of how the system actually operates. The expert supplies the structural commitments that data alone cannot make.
This is what makes Bayesian network elicitation a knowledge engineering problem, not a machine learning problem.
In the Goulburn-Broken Catchment case, the team's six experts debated which edges should appear in the network and in which directions they should point. Does water temperature directly influence fish survival? Does habitat quality influence water temperature? Does predation directly limit populations, or does it operate primarily through other mechanisms? These are not questions that emerge from statistical analysis. They emerge from ecological knowledge: how the system actually functions, what mechanisms connect variables, which interactions matter for the management question at hand. The experts identified roughly thirty variables — not the dozens they initially wanted. This curation itself required expertise. Most variables in a complex system have some influence on some outcome, but most of that influence is mediated through other variables in the network. Identifying the subset where direct influence is both mechanistically real and relevant to the decision-making problem is an expert judgment. Once variables are identified, edges must be drawn: which variables directly influence which others. Each arrow is a claim about causation, not correlation. Data cannot settle these claims. Domain knowledge must.
This is the Knowledge Engineering with Bayesian Networks (KEBN) discipline, refined over twenty years of practice.
The Goulburn-Broken case followed a structured elicitation process: a facilitator working with the six experts through a disciplined workflow. First, variable identification — ruthless prioritization toward the management question. What variables directly affect fish population outcomes? What environmental and management factors influence those drivers? The smaller set that passed these filters became the nodes of the network. Second, edge elicitation: the experts drew arrows, debating directionality. Does this variable influence that one? Is the influence direct or mediated through other variables? Third, consistency checking: the emerging structure was tested against scenarios the experts knew deeply from fieldwork. If the model predicted differently than what they had observed, where did the structure need refinement? Fourth, probability assessment: at each node, the conditional probability distribution given the node's parents must be specified — not as single point estimates, but as ranges reflecting genuine uncertainty. Fifth, confidence calibration: confidence levels attached to all estimates, acknowledging that some relationships are well-established while others depend on inference or extrapolation.
In the catchment case, the experts initially wanted dozens of variables. The disciplined process compressed this to about thirty — a set small enough to elicit probabilities for in twelve weeks, but complete enough to support decision-making.
The Goulburn-Broken team illustrates a critical insight: variable identification is not a data science task, it is an expert judgment task. The six ecologists could name many variables that influence native fish populations — water temperature, habitat quality, flow regime, predation, food availability, disease, pollution, invasive species competition, riparian vegetation, substrate composition, water chemistry, channel morphology. The list could extend indefinitely, because in a complex system, most variables have some causal influence on most other variables. Yet including every possible variable defeats the purpose of model-building. A model with hundreds of nodes becomes unwieldy. Probabilities cannot be elicited reliably for that many relationships in a reasonable timeframe. The expert judgment required is not whether a variable influences the outcome — it probably does, in some path — but whether that influence is direct, whether it is mechanistically understood, and whether it is relevant to the management question. The team needed to understand what interventions would change fish populations. Habitat restoration directly changes populations. Riparian vegetation influences habitat quality, which influences populations. These direct or near-direct influences mattered. Detailed nutrient chemistry effects on zooplankton — probably important at some level, but not a direct lever for management.
The Goulburn-Broken team did not settle on point estimates. They specified ranges — and attached explicit confidence levels to those ranges.
One of the crucial refinements in KEBN practice over two decades is the move away from false precision in probability elicitation. If a facilitator asks an expert for a single probability, the expert will give an answer — but that answer often reflects discomfort with admitting uncertainty rather than true confidence. They may anchor on round numbers or be overconfident based on vivid examples. The Goulburn-Broken process used ranges instead. What is a plausible range for the effect size if we restore this riparian section? The team might specify: somewhere between a five percent and twenty-five percent increase in fish survival. Then: how confident are we in this range? Very confident, based on multiple peer-reviewed studies. Or: moderately confident, because most data comes from temperate systems and we're in a Mediterranean climate. This structure — range plus confidence level — captures epistemic reality. Some relationships are well-established, with narrow ranges and high confidence. Others depend on extrapolation, so ranges are wider and confidence is lower. The model itself tracks this uncertainty. When the model is run forward to predict consequences of interventions, uncertainty from high-confidence relationships has less impact on the output than uncertainty from low-confidence relationships.
This is not curve-fitting or parameter tuning. It is a validation method for ensuring the model actually captures the experts' domain knowledge.
The Goulburn-Broken team did not build the network structure and probabilities, then call it done. They tested the model against scenarios the experts knew deeply from years of fieldwork. In this scenario, we knew from observation that X happened. Does the model predict X? If not, where is the mismatch? This kind of testing is not about validating the model against new empirical data. It was about validating that the model captures what the experts actually know about how the system works. If the model's predictions surprised an expert — if the expert had expected the system to behave differently than the model predicted — then the surprise pointed to a gap. Maybe a variable was defined incorrectly. Maybe an edge was missing or misoriented. Maybe a probability at a node was wrong. The surprise forced the team to go deeper into the implicit knowledge. Why did you expect the system to behave that way? What mechanism were you modeling in your head? Can we make that mechanism explicit? Can we represent it correctly in the network structure and probabilities? This iterative refinement continued until the model's behavior aligned with the experts' understanding of what they had actually observed in the field.
The choice between them depends on which cognitive biases pose the greatest risk for the particular elicitation task.
Over two decades of Knowledge Engineering with Bayesian Networks, practitioners have developed structured protocols to suppress different cognitive biases that arise when experts are asked to elicit knowledge. The Delphi protocol uses anonymous rounds of estimation with controlled feedback. Experts provide estimates without seeing what others have said. The facilitator aggregates results and provides feedback on the distribution: the median, the quartiles, how much disagreement exists. The anonymity matters. It suppresses authority bias — the tendency to defer to the highest-ranking expert — because no expert knows whose estimate is whose. The Delphi protocol suppresses conformity pressure and groupthink. The IDEA protocol takes a different approach, designed for different biases. Experts assess individually first, then come together for structured group deliberation. The goal is to surface legitimate dissent and prevent false consensus. In group settings, the loudest voice often dominates. The minority view gets abandoned not because it is wrong, but because it is unpopular. Individual assessment first preserves that minority view on the record. Group deliberation then happens with the knowledge that genuine disagreement exists, and the facilitator works to understand why those disagreements occur and whether they reflect real uncertainty.
A Bayesian network is not a data analysis technique. It is a knowledge representation and inference engine — a way to make implicit expertise actionable.
The Goulburn-Broken Catchment case demonstrates what becomes possible when expert knowledge is made explicit and structured. The six ecologists knew their system deeply — they understood fish populations, habitat, flow, predation, vegetation. But that knowledge lived in individual heads, expressed through conversation and accumulated experience. It could not be easily transferred to others. It could not be queried systematically to ask: what is the relative importance of this driver versus that one? What would change if we intervened here? How much improvement might we expect? A Bayesian network transforms implicit expertise into explicit structure: variables, edges, conditional probabilities. That structure can be run forward as a simulator: given these environmental conditions and management interventions, what outcomes would we predict? It can be queried: if we want to increase fish survival, which variables have the most leverage? It can be updated: as new monitoring data arrives, the network can be revised to reflect what we are learning. It can be communicated: stakeholders can see the underlying causal logic that leads from management decisions to outcomes. The twelve weeks of structured elicitation were not a data collection process. They were a knowledge engineering process: taking what the experts knew and making it computable, testable, and actionable.
Knowledge Engineering with Bayesian Networks has documented this principle and refined the elicitation discipline over twenty years across environmental science, clinical medicine, and intelligence analysis. Yet the approach has not crossed into corporate settings where the need is urgent. The barrier is not technical. It is structural.
The mathematical foundation is clear: Markov equivalence proves that observational data cannot uniquely determine causal structure. Multiple causal graphs can fit identical statistical patterns. This is not a limitation that better algorithms will overcome. It applies to any finite dataset in any domain. Therefore, when a causal model is required — when decisions depend on understanding intervention consequences, not just predicting under current conditions — expert knowledge is mathematically necessary. The expert supplies what data cannot: direction of causality, mechanism, structural understanding. The Goulburn-Broken case is not an anomaly. It is an exemplar. Over two decades, the discipline of Knowledge Engineering with Bayesian Networks has documented successes in environmental management, clinical diagnosis and treatment planning, epidemiology, system reliability, and intelligence analysis. The elicitation protocols have been refined through repeated application. Delphi and IDEA protocols exist specifically to suppress the cognitive biases that arise when experts are asked to articulate implicit knowledge. Yet despite this success history, the approach remains outside the corporate mainstream. Organizations continue to build causal models from data alone, or they avoid causal modeling entirely. The barrier is not mathematical. It is structural. Six categories of barrier exist: organizational, epistemological, incentive-driven, resource-driven, skill-driven, and tool-driven. Understanding which barrier is binding in any given context is the key to whether an architectural intervention can bring KEBN into the corporate boardroom.
Living Models · Chapter 14 · The Expert in the Room
Expert knowledge is not optional. It is mathematically required when causal understanding matters. The next chapters of Part Three describe the architectural moves that bring expert-elicitation discipline into environments where the corporate mainstream still expects causal models to emerge from data alone. Those moves rest on the foundation established here: the mathematics that proves data's limitations, the track record of Knowledge Engineering with Bayesian Networks showing what structured elicitation can accomplish, and the diagnosis of which structural barriers are binding in any given corporate context. Understanding that diagnosis is what makes those architectural choices coherent rather than arbitrary.