A conceptual framework – three genuinely different ways a formalized generalization can miss reality
A formalization can fail to match observed reality, and when it does, the first question is never how large the gap is. The first question is where the gap is coming from – because a formalization can fail to match observed reality in three genuinely different ways, and telling them apart determines whether the actual cause gets addressed or the wrong thing gets adjusted.
Epistemic error
The formalization is faithful to the generalization. The true value is a fixed number. It simply cannot be computed exactly. The Taylor remainder is the cleanest example of this – a single fixed real number, not a random quantity, that formalization itself gives a computable bound for, via calculus and inequalities. Floating-point truncation is the same failure in a different guise, tracing back to the fact that an irrational number cannot be written down completely, only approximated to some finite precision.
What distinguishes epistemic error from the other two is that it is bounded deterministically. There is no uncertainty about how large it can be – there is only the work of computing that bound precisely. A model with only epistemic error is a model that is, in principle, exactly right, and only practically approximate.
Aleatory error
The phenomenon itself has genuine random variability. No amount of additional precision in the formalization removes this, because the deviation is not a computational shortfall – it is a property of the thing being modeled. Deviation of this kind is modeled as a random variable with a distribution, and quantified through variance and confidence intervals, not through a tighter bound.
A model with aleatory error is not wrong to leave some deviation unexplained. Leaving it unexplained, and quantifying how much of it there is, is the correct response to a phenomenon that is inherently variable. The error here is not a defect to be engineered away – it is a fact about the world the model is honest about.
Structural error
The generalization itself was wrong. The class was misjudged, or the pattern claimed to hold does not actually hold in the form assumed – fitting a straight line to a phenomenon that is genuinely curved, or applying Newtonian mechanics near the speed of light, where the underlying assumptions the formalization rested on simply stop being true.
This is a failure in the claim, not in the machinery built to express it. No amount of better formalization fixes a generalization that was claimed too broadly, or claimed for the wrong class of cases to begin with. If epistemic and aleatory error are questions the formal structure can answer about itself, structural error is a question the formal structure cannot answer about itself – it requires stepping back outside the formalization and re-examining the claim that came before it.
Lenses, not compartments
These three are analytic lenses for asking where a deviation is coming from, not compartments a given deviation must sit inside cleanly. A real gap between a model and an observation is usually some blend of all three at once, not a single labeled cause waiting to be identified.
A diagnostic score is a clear case of this. It carries epistemic error in its own internal computation – the arithmetic of summing points or updating odds is itself an approximation of some underlying exact quantity. It carries aleatory error from the imprecision of any single test – a blood pressure reading taken twice in the same minute will not read identically both times. And it may carry structural error on top of both, if the population the score is currently being applied to has drifted from the population it was originally built and validated on. All three can be present in the same number, at the same time, contributing to the same observed deviation.
Separating the three is a way of thinking about a failure – a discipline for asking the right question first – not a claim that the failure has exactly one address. Treating them as strict, mutually exclusive compartments risks the opposite error: assigning a deviation to the wrong cause simply because it had to be assigned to one of the three, when the honest answer was that it belonged to more than one.
This is also the reason governance and monitoring of a deployed model cannot substitute for getting the underlying generalization right in the first place – a model can be watched carefully for output drift and still fail if the class it was built for has moved out from under it, which is a structural failure no amount of monitoring, by itself, repairs.
The Constant
Telling epistemic, aleatory, and structural error apart is not an exercise in classification for its own sake. It determines whether the actual cause of a deviation gets addressed, or whether the wrong thing – a tighter bound, a larger sample, a re-tuned parameter – gets adjusted instead, while the real problem, a claim that was never true of the class it was applied to, goes untouched.