
A conceptual framework – the third leg, and the one place the earlier framework needs an extra clause
Mathematics gives a generalization the language to be stated precisely. Statistics gives a generalization a way to be estimated, with honest uncertainty, when it cannot be derived outright. Neither, on its own, turns a formalized generalization into something that runs – something that can be computed at scale, simulated under many conditions, optimized, or deployed to act on new cases as they arrive. That is computation’s job, and it is the third leg of the same axis: mathematics enables formulation, statistics makes it credible, computation makes it practical.
What computation actually adds
A formalized generalization is, at the moment it is written down, still just an expression on a page. Computation is what implements it numerically, runs simulation and optimization against it, and scales it from one case to real-world size. A model that predicts one house’s price by hand is a proof of concept. The same model, implemented and run against a hundred thousand houses, an evolving market, and a live pipeline of new listings, is a working system – and the distance between those two things is entirely computation’s contribution, not mathematics’ or statistics’.
This is not a minor, purely engineering afterthought bolted onto the real work of modeling. Simulation lets a formalized generalization be tested under conditions that were never directly observed. Optimization lets a model be used not just to predict an outcome but to choose an action. Scale lets a model built and validated on a modest sample be applied to a volume of cases no person could work through by hand. All three are computation doing something mathematics and statistics alone cannot do.
Machine learning is the same act, scaled
Machine learning and artificial intelligence do not introduce a kind of modeling outside this framework. They scale the same generalize-then-formalize act, using the same building blocks already described – number systems, vector and matrix structures, the arithmetic that comes with them – and producing the same three error types already discussed, now estimated from far larger samples with far more computation behind the estimation.
A trained model is a formalized generalization. Its training data is the evidence a generalization is estimated from. Its predictions are the same inference-and-prediction split statistics already draws, carried out at scale and largely automated. What changes with machine learning is the volume of data and the automation of the estimation step – not the underlying act itself. Theorem-proving and formal generalization bounds, such as those built on VC dimension or PAC-learning theory, are the epistemic-error machinery for this setting, doing for a trained model exactly what a Taylor remainder does for a truncated series: bounding how far a sample-estimated pattern can be trusted to extend to the class it was claimed to hold over.
Where the class is no longer stated – only sampled
One clause is worth adding here, because it marks a real asymmetry rather than a minor wrinkle in an otherwise complete picture. In physics, the class a model is claimed to hold over is fixed by the modeler before formalization, and it does not move afterward – Newtonian mechanics did not drift into applying to a different range of speeds after it was written down.
In machine learning, the class is not asserted. It is sampled – encoded jointly in the architecture, the training distribution, and the loss function – and a sampled class is an empirical object that can silently shift after the model is deployed, with nothing in the training process itself positioned to notice. This means structural error is not, for a deployed machine learning system, the rare failure mode it is in a physical model that was correctly derived once. It is the default failure mode, because the class was never derived or postulated in the first place, only estimated from whatever sample happened to be available at training time.
A credit-scoring model estimated on one economic cycle and applied through the next carries this risk directly – the formalization, the trained weights, has not changed and is not wrong on its own terms; the population it was fit to has moved out from under it. A diagnostic model validated on one demographic and applied to another carries the same risk in a higher-stakes setting: the individual test results and symptoms feeding it are tangible and precisely measured, but the disease state itself is not directly observable at the time of diagnosis, and a model trained on one age range, sex, or ethnic group may simply encode a generalization that does not hold for a patient outside that population. A spam filter trained on one era of language and applied to the next shows the same pattern in a lower-stakes setting: the words themselves have moved.
In every one of these cases, epistemic and aleatory error are properties of the formal system itself – fixed once training ends, and checkable from the system alone, by examining its own internal computation and the imprecision of the individual measurements feeding it. Structural error, in this setting, cannot be checked from the system alone. It requires an ongoing comparison between the class the model was fit to and the class it is currently being applied to.
Why governance is not the same question
This is the concrete content behind calling monitoring and periodic re-validation a foundational requirement for machine learning, rather than an operational nicety attended to after the real modeling work is done. A model can be monitored for output drift, retrained on new data, and still fail if the original generalization was claimed too broadly for its intended use, or if the population it was validated on has shifted – the same risk already named for banking and medicine above.
Governance answers whether a model is being watched. It does not answer whether the generalization behind the model still holds. That second question can only be answered by returning to the two acts this framework opened with: checking the generalization’s claimed class against the case currently at hand, and checking whether the formalization built to express it is still faithful to that claim. Monitoring and re-validation are the only mechanism available for catching a form of error that, for this one kind of model, never announces itself from the inside.
The Constant
Mathematics enables formulation, statistics makes it credible, and computation makes it practical – including, at the far end of that axis, machine learning, which does not stand outside the earlier framework so much as expose the one clause it was always going to need: a class that was only ever sampled has to be watched, because unlike a class that was asserted once and fixed, it can move.