The Machinery – Mathematics as Structure and Logic

Generalization makes a claim. Formalization gives that claim an exact structure. But formalization cannot do its work with nothing to build from – it needs numbers, it needs a vocabulary for the pieces of a relation, and once a claim involves more than one quantity at a time, it needs containers built for that job. None of what follows in this piece is modeling’s purpose. It is modeling’s toolkit – the machinery formalization requires, and nothing more than that.

Formalization needs numbers to work with, and numbers come in a hierarchy, each containing the one before it. The naturals, used for counting, sit inside the integers, which add negative numbers and zero. The integers sit inside the rationals – numbers whose decimal expansion terminates or repeats. And the rationals sit inside the reals, which also contain the irrationals: numbers whose decimal expansion is non-terminating and non-repeating, such as $\pi$, $e$, and $\sqrt2$.

The moment a formalized model involves such a number, exact computation becomes structurally impossible – not a flaw in the modeler’s method, but a property of the number itself. This is the first, most primitive place where a formalized generalization runs into an unavoidable gap between the exact mathematical object and anything that can actually be written down or computed. It is worth sitting with this fact rather than passing over it: the gap is not a limitation of current tools that better tools will someday close. It is built into what an irrational number is. Every computation involving $\pi$ or $e$ or $\sqrt2$ that a computer ever performs is, at some finite decimal place, a deliberate act of stopping short of the exact value – and that gap resurfaces, in a more developed form, wherever a model’s error is later discussed.

To formalize the relation $\text{d} = \frac{1}{2} \text{gt}^2$ , the pieces of that relation first had to be sorted into roles. This sorting is not decoration – it is the vocabulary that makes the formalization legible and reusable.

  • Constant: fixed everywhere the model applies. In a simpler treatment of the falling-stone relation, this could be g itself, or the $\frac{1}{2}$
  • Variable: changes within one application of the model. Distance fallen and time elapsed are both variables here – each application of the model produces a different pair of values for them.
  • Parameter: fixed within one instance of the model but free to differ across instances of the same model family. This is what g actually is – 9.8 meters per second squared near Earth’s surface, a different value on the Moon. The form $\text{d} = \frac{1}{2} \text{gt}^2$ generalizes across any gravitational field, with g as the dial that adapts it to a specific one.
  • Function: the formal relation itself, binding these roles together into a single computable object.

Recognizing g as a parameter rather than a fixed constant is not a pedantic distinction – it is the immediate payoff of doing Part I’s discipline correctly. It is what lets the same formalized generalization extend to the Moon without being rebuilt from scratch. The generalization was broader than the first formalization made evident, and the vocabulary of constant, variable, and parameter is precisely what exposes that breadth.

A generalization can involve one varying quantity, two, or many. A relation involving a single variable is univariate – the falling-stone model, expressing distance as a function of time alone, is univariate in time. A relation involving two is bivariate. A relation involving many at once is multivariate.

A model of house prices built from square footage, location, age, and room count is multivariate – and this is not merely a description of how many inputs happen to be listed. The generalization being claimed, that price depends jointly on these factors, is inherently about several quantities interacting at once, and the formalization that follows has to represent that jointness, not simply list the factors separately as though each acted alone.

Once a formalization is multivariate, it needs containers built for that job. A vector is an ordered list – a single point in a multivariate space. A matrix is a two-dimensional grid – a dataset, or a linear transformation. A tensor is the n-dimensional generalization of both.

Return to the house-price case. The generalization was that price depends jointly on square footage, location score, age, and room count. Formalized as a relation for a single observation, this is price equal to a weighted sum of the four factors plus a constant offset – each weight a parameter attached to one factor, the constant offset a parameter of its own:

price $= w_1x_1 + w_2x_2 + w_3x_3 + w_4x_4 + b$

Written this way, four observations would require four separate equations, each repeating the same four parameters. The vector and matrix structures exist to remove that repetition. Let $X$ be one house, written as a vector of its four factors, and w be the parameters, written as a vector of the same length:

$X = (x_1, x_2, x_3, x_4) ~~         W= (w_1, w_2, w_3, w_4)$

Then price becomes the dot product of w and x, plus the constant:

price $= W\cdot X + b$

For $m$ houses at once, stack each house’s $x$ as a row, giving an $m$-by-4 matrix $X$.

The single dot product becomes one matrix-vector product, producing all m predictions in one operation instead of m separate equations:

prices $= XW + b$

This is the point worth pausing on: the formalization did not change. It is the same generalization, the same parameters w, the same constant b. Only the container changed, from a single equation to a matrix equation, so that one operation applies uniformly whether there is one house or one hundred thousand.

A tensor extends this the same way once a third dimension enters – tracking the same m houses across T time periods gives an m-by-4-by-T tensor, and the operation that generalizes matrix multiplication to this case is contraction: summing over one shared dimension while keeping the others intact.

Each structure comes with its own arithmetic – vector addition and the dot product, matrix multiplication and the transpose, tensor contraction – and these are what let a formalization of many interacting quantities be computed with as a single object, rather than as an unwieldy list of separate variables that happen to share a name.

This is a property of the finished formal structure, not of the generalization behind it. Given a fixed input, does the function return the same output every time – deterministic, as with $\text{d} = \frac{1}{2} \text{gt}^2$ – or does it return a distribution of possible outputs because randomness is built into the structure itself – stochastic, as with a Markov chain or a queuing model?

Both are legitimate formalizations of a legitimate generalization. Neither is more rigorous than the other in the abstract. The choice depends entirely on whether the phenomenon being generalized over is itself inherently variable or not – forcing a deterministic structure onto a phenomenon with genuine randomness does not make the model more precise, it only hides the randomness inside a false appearance of certainty.

With the vocabulary of numbers, quantities, dimensionality, and structure in place, the next question is where a formalized generalization can fail to match observed reality – and, crucially, why the three ways it can fail are not the same failure wearing different names.

None of this machinery is modeling’s purpose. It exists only because formalizing a generalization requires a vocabulary and a set of structures precise enough to carry it. The number system, the roles of constant, variable, and parameter, dimensionality, and the containers of vector, matrix, and tensor – all of it is in service of the act described before it, never a substitute for that act.

Next>

Scroll to Top