Statistics and Data – Uncertainty and Inference

With generalization and formalization established as modeling’s core act, statistics and data can be placed precisely, rather than folded into modeling as if the three were interchangeable. They are not rivals to modeling, and they are not modeling under another name. Each has a specific job, and the job only exists because of a specific limitation in how a generalization can sometimes be reached.

The relation $\text{d} = \frac{1}{2} \text{gt}^2$ was generalized from a small number of observations and derived from deeper physical principles – its class and its form could, in principle, be argued for without reference to a large dataset. Most generalizations a professional actually needs are not like this. There is no first-principles derivation available for how a drug affects blood pressure across a population, or whether a given customer will churn.

In those cases, the generalization can only be estimated from a limited, sampled record – and that is precisely the job statistics exists to do. Its mandate has not changed in three and a half centuries: describe what the sample actually shows, then infer, with stated uncertainty, what that implies about the larger population or process the sample was drawn from.

That mandate branches into two further tasks that are often run together, because both are built on the same sample, but that ask genuinely different questions. Inference asks what the sample implies about the population or process as a whole – a claim about the entire class. Prediction asks what is likely true of one new, specific instance not yet observed – a claim about a single member of that class.

A regression fit to housing data can be used to infer the population-wide relationship between square footage and price, or to predict one particular new house’s price. The same formalized generalization serves both tasks – but the uncertainty attached to a claim about the whole population is not the same uncertainty attached to a claim about one new case.Reporting one in place of the other overstates or understates confidence in what was actually asked, and this confusion is one of the most common quiet errors in applied statistical work.

Statistics, in this framing, is the specific tool a modeler reaches for when the generalization step cannot be settled by derivation and must instead be settled by evidence, with honestly quantified uncertainty. Aleatory error, discussed earlier, is exactly the error type statistics is built to quantify – it is not a coincidence that variance and confidence intervals are statistics’ own native vocabulary.

Statistics needs something to describe and infer from – that something is data. Data has an origin, a state, and a format, and it is held in some storage architecture, and it becomes valuable only once it passes from mere availability into something that actually gets used.

But data, by itself, generalizes nothing and formalizes nothing. It is inert material. A dataset sitting unused makes no claim about a class and expresses no pattern in exact structure – it becomes part of a model only once statistics, or some other method of estimation, uses it to support or refine a generalization, which is then formalized into a usable structure. Treating the presence of data as though it were itself the presence of a model is a category error, and a common one.

Modeling – generalize, then formalize – is the overarching act. It is the only one of the three that decides what claim is being made, and in what form. Statistics – describe, then infer – is the tool invoked within modeling, specifically when the generalization step cannot be derived and must instead be estimated from an incomplete record. Data is the raw substrate statistics operates on – necessary for that one step, but not itself an act of generalizing, formalizing, describing, or inferring.

Keeping these three distinct, rather than treating modeling as just a label for doing statistics on data, is what makes it possible to recognize the many models that require no data and no statistics at all – physical laws, geometric proofs, formal logical systems – and to see clearly, when a model does rely on data, exactly which step of the modeling act that data is serving.

Statistics itself has its own long history as a discipline – from an instrument of state administration to a tool now applied to a single person’s purchases and habits – and data has its own separate history of origin, scope, and format. Both are worth their own treatment; what matters here is only where each one fits into the act this framework opened with.

Describe what a sample actually shows. Infer, with stated uncertainty, what that implies about what was not observed. This two-word mandate has not changed since statistics began, and every later refinement – significance testing, Bayesian updating, sampling theory, experimental design, even machine learning’s predictive models – is an elaboration of it, not a replacement for it.

Next>

Scroll to Top