
A conceptual framework – from a rough sense of a pattern to a structure you can compute with
Modeling is not, at its root, a collection of formulas, structures, or error bounds. Those are its ingredients and its machinery. Modeling itself is a single deliberate act made of two distinct moves: generalization and formalization.
Everything that gets built on top of a model – the number system it uses, the quantities it names, the vectors and matrices that hold its variables, the question of whether it behaves deterministically or randomly, the ways it can fail – exists to serve that one act. Before any of that machinery is worth discussing, the act itself has to be understood on its own terms.
Before a Generalization Is Even Stated
A generalization does not arrive fully formed. Before it is stated as a claim about a class of cases, it usually exists in a rougher form, and telling that rougher form apart from the claim itself is the first discipline this framework asks for.
An intuition is an unstated sense that some pattern is there – formed from experience, but not yet examined. It typically carries an ambiguity: the class the pattern might hold over has not yet been decided, and the same intuition could be read narrowly or broadly. A researcher who has “a feeling” that two things are related has an intuition, not yet a generalization. This is the raw material a generalization is built from, not the generalization itself.
Once that raw sense is committed to a direction, it takes one of two forms. A postulation is a starting claim adopted without derivation and accepted as a foundation for what follows – the laws Newton’s mechanics rests on are postulations, not conclusions reached from something prior. A hypothesis is a candidate pattern proposed from observed cases and offered for testing, not accepted outright – it is the form an inductive generalization takes before evidence has confirmed or revised it.
Postulation and hypothesis are two different starting points for the same act. Which one applies to a given generalization is the deductive-versus-inductive distinction – a question of whether the claim was derived from principles already accepted, or proposed from a finite set of observed cases and still awaiting evidence. Both are legitimate ways to begin. What is not legitimate is skipping the question of which one you are doing, and proceeding as though a hypothesis had the certainty of a postulation, or as though a postulation still needed the kind of testing a hypothesis requires.
1. Generalization – the move from instance to class
Generalization is the act of taking something true of a specific case and asserting that a pattern holds across an entire class of cases, not just the one observed.
Consider a single concrete fact: this particular stone, dropped from this particular height, hit the ground after 2.3 seconds. That sentence, by itself, is not a model – it is a record of one event. It becomes the seed of a model only when someone makes the leap: any object, dropped from any height, near the Earth’s surface, will fall according to the same underlying pattern. That leap – from “this stone, this time” to “any object, any time, under these conditions” – is generalization. It is a claim about a class, made on the evidence of a member, or a few members, of that class.
Generalization has degrees, and being explicit about the degree claimed is itself part of doing it responsibly.
- Narrow generalization: “objects of this shape and density, dropped in air, near sea level” – a claim restricted to a tightly specified class.
- Broad generalization: “any object with mass, anywhere in a gravitational field” – a claim stretched to a much wider class, and correspondingly a much bigger claim to defend.
The width of the generalization is a choice, and it is the first place a model can go wrong – not through faulty arithmetic, but through over-generalization: claiming a pattern holds for a class wider than the evidence actually supports. A drug that worked in forty healthy adult volunteers has not been shown to be safe for everyone; the generalization that would say so was never earned by the evidence that produced it.
2. Formalization – the move from pattern to precise structure
Formalization is a separate act from generalization, and collapsing the two is the second discipline this framework insists on. Once a pattern is believed to hold across a class, formalization is the act of expressing that pattern in exact, symbolic, manipulable language – turning an intuitive or verbal claim into a structure that can be computed with, checked, and used to derive new consequences.
Continuing the falling-stone example: the generalization “objects fall according to a consistent pattern” is still just a sentence – useful as an idea, but not yet usable for calculation. Formalization turns it into a precise relation between distance fallen and time elapsed:
$$\text{d} = \frac{1}{2} \text{gt}^2$$
Now the pattern has become a function relating a variable, d, distance fallen, to another variable, t, time elapsed, through a constant, g, gravitational acceleration, by way of a precise arithmetic operation. This formal object can now be manipulated – solved for t, differentiated to get velocity, integrated, compared against a measurement – in ways the verbal generalization alone never could be.
Two distinctions are worth stating precisely at this point, because they are easy to blur and the blurring is where a great deal of bad modeling begins.
- Generalization without formalization is common and often useful on its own – “practice makes performance more reliable,” “prices tend to rise with demand.” These are real generalizations, asserted across a class of cases, but stated only in natural language. They guide intuition but cannot yet be computed with, tested numerically, or used to derive a precise prediction.
- Formalization without a genuine generalization behind it carries its own risk: writing down an equation that appears precise and settled, when the underlying claim about what class this holds for was never examined. A formula can be well-formed mathematics and still rest on a generalization that was not verified.
Two examples make this concrete. Phrenology took skull-shape measurements and turned them into precise numerical charts correlating bump location with character traits – the formalization was exact and computable, but the generalization behind it, that skull shape determines character, was never actually established. A subtler case: a regression fit connecting ice-cream sales to drowning deaths is a perfectly valid formal object – a coefficient, a line, a p-value – but the generalization it implicitly asserts, that one causes the other, was never made or defended; both are driven by a third factor, summer heat. The equation is real; the claim it appears to formalize was not.
A second worked example, drawn from calculus, shows the same two-step structure applied to approximation. The generalization: near a chosen point a, a sufficiently smooth function behaves like a polynomial that matches its value and derivatives at that point, and the match stays close through a neighborhood of the point. Formalization gives this claim an exact structure – the Taylor expansion of $f(x)$ around a, to degree n:
$$f(x) = f(a) + f′(a)(x−a) + f″(a)\frac{(x−a)^2}{2!} + \cdots + f^{(n)}(a)\frac{(x−a)^n}{n!} + R_n(x)$$
The generalization – “the function behaves like a matching polynomial nearby” – has now been formalized into a specific expression, with a remainder term $R_n(x)$ that formalization also gives a fixed, computable bound, via Lagrange’s form:
$$R_n(x) = \frac{f^{(n+1)}(c)(x−a)^{(n+1)}}{(n+1)!}$$ for some c between a and x
$R_n(x)$ is a single fixed real number, not a random quantity – the polynomial is an approximation of $f(x)$, and the bound on how far that approximation can be from the true value follows directly from how precisely the generalization was formalized. This is why a later error can often be traced back to a decision made here, at the formalization step, rather than to anything that happens downstream.
Modeling, then, is the disciplined pairing of the two: generalize deliberately – state the class the pattern is claimed to hold over – then formalize precisely – give that pattern an exact structure – while keeping track of the boundary of the class the whole time. That boundary, the domain of validity, is what a model’s later machinery exists to help specify, test, and, when the formal structure fails, diagnose.
3. Why this ordering matters
Generalization comes first conceptually because it is the claim being made – it is a statement about reality, and it can be right or wrong, wide or narrow, careful or casual or arbitrary, entirely independent of any mathematics. Formalization comes second because it is the expression of that claim in a form precise enough to compute with, test, and extend.
A model is only as good as its generalization is honest and its formalization is faithful to that generalization. No amount of mathematical sophistication applied afterward compensates for a generalization that was incorrect from the outset. This is the ordering that the rest of this framework is built to protect: get the claim right first, then build the structure that expresses it.
Mathematics, statistics, and computation – the machinery that turns a formalized generalization into something usable, testable, and scalable – are each their own piece of this framework, and are taken up in turn.
The Constant
Modeling is not a collection of formulas, structures, or error bounds. It is the deliberate act of generalization and formalization – and everything else exists only to serve that act.