Model Validation, Selection, and Assessment

Data Set Splits in a Statistical Learning Procedure

The subset of data used to fit the model – to estimate parameters and make predictions on known data.

  • Training error is a poor estimate of true performance – it almost always underestimates test error because the model was optimized on this very data.

A subset held back during training, used to tune model choices – hyperparameters, model complexity, variable subsets, regularization parameter $\lambda$, etc.

  • The model does not learn from this data, but model selection decisions are driven by performance on it.
  • Also called the hold-out set.

A subset held back entirely until the very end, used once to report an honest, unbiased estimate of the final chosen model’s performance on new, unseen data.

  • Must never be used to guide any modeling decision.
  • Using it more than once invalidates its role as an unbiased estimator.

Training error underestimates test error. The validation set guides model selection. The test set gives the final unbiased verdict – used exactly once.


We want to estimate a scalar parameter $\theta$ using an estimator $\hat{\theta}$.

The MSE is defined as:

$MSE(\hat{\theta}) = E[(\hat{\theta} – \theta)^2]$

Add and subtract $E[\hat{\theta}]$:

$MSE(\hat{\theta}) = E[(\hat{\theta} – E[\hat{\theta}] + E[\hat{\theta}] – \theta)^2]$

Expanding:

$= E[(\hat{\theta} – E[\hat{\theta}])^2] + (E[\hat{\theta}] – \theta)^2 + 2 \cdot E[(\hat{\theta} – E[\hat{\theta}])](E[\hat{\theta}] – \theta)$

The cross term

$E[\hat{\theta} – E[\hat{\theta}]]$

$= E[\hat{\theta}] – E[E[\hat{\theta}]]$

$= E[\hat{\theta}] – E[\hat{\theta}] =0$.

Therefore:

$MSE(\hat{\theta}) = \text{Var}(\hat{\theta}) + [\text{Bias}(\hat{\theta})]^2$

where $\text{Bias}(\hat{\theta}) = E[\hat{\theta}] – \theta$.


The true model is:

$Y = f(X) + \epsilon, \quad E[\epsilon] = 0, \quad \text{Var}(\epsilon) = \sigma^2_\epsilon$

We estimate $f$ by $\hat{f}$ and predict at a new point $x_0$.

The expected test MSE at $x_0$ is $E[(y_0 – \hat{f}(x_0))^2]$

Now $y_0 = f(x_0) + \epsilon$:

$E[(f(x_0) + \epsilon – \hat{f}(x_0))^2]$

$= E[(f(x_0) – \hat{f}(x_0))^2] + E[\epsilon^2] + 2\,E[(f(x_0) – \hat{f}(x_0))\,\epsilon]$

cross term = 0

$\hat{f}(x_0)$ is trained on $(x_1,y_1),\dots,(x_n,y_n)$, so it depends only on the training noise $\epsilon_1,\dots,\epsilon_n$. The test noise $\epsilon_0$ is a fresh draw, independent of all training noise:

$$\epsilon_0 \perp \{\epsilon_1,\dots,\epsilon_n\} \implies \epsilon_0 \perp \hat{f}(x_0)$$

$\therefore E[(f(x_0) – \hat{f}(x_0))\,\epsilon_0]$

$= E[f(x_0)-\hat f(x_0)]\cdot E[\epsilon_0]$

$= E[f(x_0)-\hat f(x_0)]\cdot 0$

$ = 0$

hence,

$E[(f(x_0) + \epsilon – \hat{f}(x_0))^2]= E[(f(x_0) – \hat{f}(x_0))^2] + E[\epsilon^2] $

Decompose the first term as in Formulation 1:

$E[(f(x_0) – \hat{f}(x_0))^2] = \text{Var}(\hat{f}(x_0)) + [\text{Bias}(\hat{f}(x_0))]^2$ where $\text{Bias}(\hat{f}(x_0)) = E[\hat{f}(x_0)] – f(x_0)$.

$\therefore E[(y_0 – \hat{f}(x_0))^2] = \text{Var}(\hat{f}(x_0)) + [\text{Bias}(\hat{f}(x_0))]^2 + \sigma^2_\epsilon$

The third term $\sigma^2_\epsilon$ is the irreducible error – it cannot be reduced regardless of how well we estimate $f$, as it is noise inherent in $Y$ itself.

Recall: $\text{Cov}(X_1,X_2) = E[X_1 X_2] – E[X_1]E[X_2]$. Let $E[X_2]=0$ and $X_1, X_2$ are uncorrelated.

Then, $E[X_1 X_2] = \text{Cov}(X_1, X_2) + E[X_1]E[X_2]$

$ = 0 + E[X_1]\cdot 0$

$ = 0$

Note: independence isn’t strictly required – only $\text{Cov}(f(x_0)-\hat f(x_0),\ \epsilon_0)=0$ is needed. Independence is the stronger, more natural assumption to state, but uncorrelatedness (plus $E[\epsilon_0]=0$) is enough to prove that the cross product becomes zero


The core problem: Training error is useless for judging generalization – a model can always fit its training data better by becoming more complex, but that complexity may hurt it on new data (overfitting). We need honest estimates of test error, and we need them to compare competing models.

Randomly split data into training and a hold-out validation set. Fit on training, evaluate on validation. That validation error estimates test error.

Weaknesses:

  1. High variance – a different random split gives a different answer.
  2. The model assessed is fit on less data than the final deployed model, so validation error tends to overestimate true test error (bias).

Hold out exactly one observation, train on the rest, record error on that point. Repeat for every observation; average the errors.

  • Bias: Nearly eliminated – each training set is almost the full data.
  • Determinism: No random splits, so it gives the same answer every run.
  • Cost: Must fit the model n times (expensive for large n or slow models). A shortcut exists for linear models that makes it nearly free, but it doesn’t generalize to other methods.
  • Subtle problem: The n training sets overlap almost entirely with each other, so the n error estimates are highly correlated. Averaging correlated quantities doesn’t reduce variance as much as averaging independent ones – so LOOCV can have higher variance than expected.

Randomly divide data into K roughly equal folds. Hold out one fold as validation, train on the remaining K−1, record error on the held-out fold. Repeat through all K folds, average.

  • K = 5 or K = 10 are standard – empirically neither too biased nor too variable.
  • Compared to LOOCV: smaller training sets per step (more bias) but less correlated error estimates (less variance). This trade-off tends to favor K-fold in practice.
  • Used for choosing between models or tuning parameters: fit each candidate on the same K folds, compare CV errors, pick the lowest.

One-standard-error rule: If several models have similar CV error, prefer the simplest one whose error is within one standard error of the minimum.

Samples from the data with replacement to create many artificial data sets of the same size as the original.

  • On average, about one-third of original observations are left out of any given bootstrap sample (the “out-of-bag,” OOB, observations).
  • Fitting the model to many bootstrap samples and observing how estimates vary gives a measure of uncertainty – standard errors, confidence intervals – for quantities with no closed-form formula.
  • Not primarily a test-error estimator the way CV is, because the model is trained on samples overlapping the evaluation data. OOB observations solve this by acting as a natural internal test set.

Bagging (bootstrap aggregation): Fit the model to many bootstrap samples and average the predictions. Averaging reduces variance without increasing bias.

  • Especially powerful for high-variance procedures like decision trees – a single deep tree is unstable, but averaging hundreds of trees grown on bootstrap samples cancels out the instability.
  • OOB error: For each observation, collect predictions only from trees that didn’t use it in training, and average those. This is a valid test-error estimate, essentially equivalent to LOOCV for bagged models – without extra computation.
  • Small K: training sets much smaller than full data $\rightarrow$ weaker model evaluated $\rightarrow$ more bias.
  • Large K (toward LOOCV): less bias, but training sets nearly identical to each other $\rightarrow$ highly correlated error estimates $\rightarrow$ more variance.
  • K = 10 is the empirical sweet spot most practitioners use.

CV answers “which model predicts new data best?” – not “which fits training data best?” (the more complex model always wins that contest).

The typical pattern is a U-shape in CV error vs. complexity: error falls as useful flexibility is gained, then rises as the model starts overfitting (chasing training noise). The bottom of the U is the target.

Critical mistake to avoid: doing variable selection or any data-driven preprocessing on the full dataset before running CV. If you screen variables using all the data and then cross-validate, the validation folds have already leaked information through the screening step – CV error becomes optimistic, sometimes dramatically so.

Rule: Every data-driven step – scaling, variable selection, dimension reduction – must happen inside the CV loop, applied only to the training folds at each step.

Boosting builds trees sequentially, each correcting the errors of the previous ones (unlike bagging’s independent trees). Because the trees are dependent, boosting will eventually overfit if run with too many trees – the number of trees is a tuning parameter chosen via CV. Its CV error curve also shows the same U-shape.

Use cross-validation to select the model and its complexity, keeping the test set untouched throughout. Once all decisions are made, evaluate the final model on the test set exactly once – that number is what you report.

Bootstrap gives something different and complementary: a window into the uncertainty of estimates (how stable is a coefficient, how much would predictions change with slightly different data). Cross-validation tells you how good the model is; bootstrap tells you how reliable your estimates are.


Models range from simple/interpretable (low flexibility, higher bias, lower variance) to complex/flexible (high flexibility, lower bias, higher variance):

Parametric models (fixed functional form assumed):

  • Linear Model (LM)
  • Generalized Linear Model (GLM)
  • Generalized Additive Model (GAM)
  • Ridge / Lasso regression (regularized linear models)
  • Linear/Quadratic Discriminant Analysis (LDA/QDA)
  • Naive Bayes

Non-parametric / flexible models (form learned from data):

  • K-Nearest Neighbors (KNN)
  • Decision Trees
  • Random Forest (RF)
  • Boosting
  • Support Vector Machine (SVM)
  • Neural Networks (NN)

Cross-validation is the tool used to decide where on this spectrum to land for a given problem – by minimizing estimated test error, not training error.


Model Selection: The process of estimating the performance of several candidate models (or several complexity/tuning levels of the same model) using only training and validation data, in order to choose the single best one.

Model Assessment: The process of estimating how well the one, final, already-selected model will generalize to new data – using a test set that played no role in fitting or selecting that model.

Primary reference (introductory):

The book distinguishes evaluating a model’s performance (model assessment) from selecting the proper level of flexibility for a model (model selection).

Primary reference (detailed treatment):

In a data-rich situation, the dataset is divided into training, validation, and test parts: the training set fits the models, the validation set estimates prediction error for model selection, and the test set is reserved for assessing the generalization error of the final chosen model.

AspectModel SelectionModel Assessment
GoalChoose which model, or which level of flexibility, to useReport an honest, unbiased number for the chosen model’s real-world performance
Data usedTraining set + validation set (or resampled equivalents)Test set only, and only once
Question answered“Which of these candidates predicts best?”“How good is this one model, really?”
Typical inputsA set of candidates (LM, GLM, GAM, ridge/lasso, KNN, LDA/QDA, naive Bayes, trees, RF, boosting, SVM, NN) or one model at different tuning values (λ, K, number of trees)The single winning model produced by model selection
ToolsValidation-set approach, LOOCV, K-fold CV; sometimes AIC/BIC/Cp as analytical shortcutsA single pass of test-set evaluation
OutputOne chosen model/configuration, plus a CV error curve (often U-shaped)One number: the reported test error/accuracy
Repeatable?Yes – can be revised and reused as many times as neededNo – exactly once
Key risk if done wrongReusing the same data for both tuning and reporting error – always makes performance look artificially betterTouching the test set more than once, or using it for any decision – invalidates it as an unbiased estimate
  • Model selection always happens first and can be repeated as many times as needed – precisely why it must never touch the test set.
  • Model assessment happens exactly once, after selection is finished, and its number is what gets reported.
  • The validation/CV error from selection is a biased, optimistic estimate of true test error, because that same data helped tune the model. Only data that never influenced any decision – the untouched test set – gives an unbiased estimate.

Although both are “held out” from training, they are functionally different:

Validation SetTest Set
Used how many timesRepeatedly – every hyperparameter value, every model, every fold gets checked against itExactly once, at the very end
Role in the loopActively drives decisions: which λ, which K, which model winsPassive: reports a number, decides nothing
Does the model “see” it?Indirectly – hyperparameters are chosen specifically because they performed well on this dataNever, in any form
Statistical property of its error estimateOptimistically biased, because many configurations were searched to minimize itUnbiased, because nothing was tuned to make it look good

The tuning problem, explained: Suppose 50 values of λ (ridge/lasso) or 50 tree counts (boosting) are tried, and whichever gives the lowest validation error is kept. The model’s parameters weren’t trained on the validation set — but the validation set did influence a decision (which λ to keep). This is a subtler form of leakage than training on it directly, but it is still leakage: with enough candidates tried, one will look good on the validation set purely by chance — the same way one of many random stock pickers will look good in hindsight.

This effect is called selection bias (or “overfitting to the validation set”). The more hyperparameters/models searched, the larger this optimistic bias grows, even without ever literally training on that data. This is exactly why ESL insists on a third, completely untouched set: after the search concludes, the test set gives a number reflecting the actual final model — not the best-looking result from a multi-way comparison.

Practical consequence: Inside “model selection,” the validation set does double duty — tuning hyperparameters and choosing among model families. All of that is legitimate as long as it stays inside that boundary. To report “how good is my final random forest with these tuned hyperparameters,” one must leave the validation data behind entirely and go to the test set, because the validation number is already contaminated by the search that produced it.

In more careful setups – especially with heavy hyperparameter search (e.g., deep learning) – practitioners use nested cross-validation:

  • An outer loop handles model assessment.
  • An inner loop (run within each outer training fold) handles model selection/hyperparameter tuning.

This structure specifically prevents validation-set optimism from leaking into the final reported number, giving a more rigorous separation between “choosing the best configuration” and “reporting how good it really is” when the tuning search is extensive.


The “NN” entry in the model spectrum is really the entry point to a much larger family. Deep learning is not a different idea from the neural network – it is the same idea (layers of weighted connections, non-linear activations, trained by gradient descent) pushed further along several dimensions at once: more layers, specialized architectures for specific data types, and vastly more compute and data.

A classic (shallow) neural network has one hidden layer sitting between input and output. Deep learning simply stacks many hidden layers. Depth matters because each layer can build on the features the previous layer discovered – early layers learn simple patterns, later layers combine those into increasingly abstract ones. This wasn’t practical until three things came together:

  • Backpropagation at scale – the same chain-rule-based gradient computation used in shallow nets, but made efficient enough for many-layer networks.
  • Better activation functions (e.g., ReLU replacing sigmoid/tanh) – avoided the “vanishing gradient” problem where signals died out across many layers.
  • Compute and data – GPUs made the matrix operations feasible; internet-scale datasets gave these high-flexibility models enough examples to avoid pure overfitting.

In bias-variance terms (refer ‘MSE for Estimating f’): depth and width increase flexibility, pushing bias down – but this only pays off if there’s enough data and regularization to keep variance from exploding in return.

ArchitectureBuilt forCore idea beyond a plain NN
CNN (Convolutional Neural Network)Images, spatial dataSmall filters slide across the input, sharing weights – captures local patterns (edges, textures) regardless of position, with far fewer parameters than a fully connected layer
RNN / LSTM / GRUSequences (text, time series, audio)Connections loop back in time, so the network carries a “memory” of earlier inputs. LSTMs and GRUs add gating mechanisms to control what memory is kept or forgotten, fixing the vanishing-gradient problem in long sequences
TransformersSequences, especially languageReplace recurrence with self-attention – every element in the sequence directly looks at every other element, weighted by relevance, rather than passing information step by step. This parallelizes far better than RNNs and underlies modern large language models
AutoencodersCompression, denoising, anomaly detectionTrains the network to reconstruct its own input through a narrow bottleneck layer, forcing it to learn a compressed representation
GANs (Generative Adversarial Networks)Generating realistic synthetic dataTwo networks compete: a generator tries to produce convincing fake data, a discriminator tries to catch the fakes – both improve through the competition
Graph Neural Networks (GNNs)Data with relational/graph structureGeneralizes convolution to irregular structures (social networks, molecules) by aggregating information from a node’s neighbors

Every one of these is still, underneath, doing what (refer ‘ MSE for Estimating f ‘ – ‘ How to Assess and Choose a Model ‘) describe:

  • They’re fit by minimizing a loss function via gradient-based optimization – a fancier descendant of the same fitting principle as any parametric model.
  • They still sit at the high-flexibility, high-variance end of the model spectrum – arguably far beyond RF/SVM/boosting.
  • Model selection and model assessment still apply in principle: hyperparameters (learning rate, depth, width, regularization strength, architecture choice) are tuned on a validation set, and a genuinely held-out test set gives the final honest number.

A few validation-related habits shift once models get this large:

  • K-fold CV becomes rare. Training a single deep model can take hours to weeks; training it K times is often not affordable. A single train/validation/test split (refer ‘MSE for Estimating a Parameter θ’ ) is far more common in deep learning than repeated resampling.
  • Early stopping doubles as implicit model selection. Instead of tuning “how many training iterations,” the model is trained while continuously watching validation loss, and training stops once validation loss stops improving – this is functionally a form of the U-shaped CV curve in ‘ Choosing Between Models of Different Complexity ‘, just traced against training epochs instead of model complexity.
  • Regularization tools multiply. Alongside the ridge/lasso-style penalties already in the model spectrum, deep networks add dropout (randomly disabling neurons during training) and batch normalization, both aimed at the same bias-variance goal – keeping the highly flexible model’s variance in check.
  • The hyperparameter-leakage risk from ‘Validation Set vs. Test Set’ (section) gets bigger, not smaller. Deep learning involves searching over many more hyperparameters (architecture, learning rate schedule, regularization strength, data augmentation policy) than a classical model, so validation-set optimism compounds unless the final number always comes from a test set that was never part of that search – making the model-selection/model-assessment boundary in ‘Validation Set vs. Test Set’ (section) more important here, not less.

A practical reference for implementing everything above. Grouped to match the sections they belong to.

TaskRPython
Simple train/test splitsample() on row indices; caret::createDataPartition()sklearn.model_selection.train_test_split()
Stratified split (preserve class balance)caret::createDataPartition(y, ...)train_test_split(..., stratify=y)
Tidy/pipeline-based splittingrsample::initial_split() (tidymodels)sklearn.model_selection (pipeline-integrated)
MethodRPython
Validation-set approachcaret::createDataPartition() + manual fit/evaluatetrain_test_split() used twice (train $\rightarrow$ val split)
LOOCVboot::cv.glm(data, model, K = n); caret::trainControl(method="LOOCV")sklearn.model_selection.LeaveOneOut()
K-fold CVboot::cv.glm(data, model, K = k); caret::trainControl(method="cv", number=k); glmnet::cv.glmnet()sklearn.model_selection.KFold(), cross_val_score(), cross_validate()
Repeated K-foldcaret::trainControl(method="repeatedcv")sklearn.model_selection.RepeatedKFold()
Stratified K-fold (classification)caret::trainControl(method="cv", classProbs=TRUE)sklearn.model_selection.StratifiedKFold()
Bootstrap (general)boot::boot()sklearn.utils.resample(), scipy.stats.bootstrap()
Nested CV (refer ‘What Changed in Practice’)mlr3::AutoTuner + nested resample(); tune package (tidymodels)GridSearchCV/RandomizedSearchCV wrapped inside cross_val_score()
One-standard-error rule / CV curve plottingglmnet::cv.glmnet() (built-in lambda.1se)Manual via cross_val_score() results + numpy/matplotlib
ModelRPython
Bagging (general)ipred::bagging()sklearn.ensemble.BaggingClassifier / BaggingRegressor
Random ForestrandomForest::randomForest(); ranger::ranger() (faster)sklearn.ensemble.RandomForestClassifier / RandomForestRegressor
OOB error extractionrandomForest()$err.rate; ranger()$prediction.errorRandomForestClassifier(oob_score=True) $\rightarrow$ .oob_score_
Gradient Boostinggbm::gbm()sklearn.ensemble.GradientBoostingClassifier
XGBoostxgboost (R package)xgboost (Python package)
LightGBMlightgbm (R package)lightgbm (Python package)
CatBoostcatboost (R package)catboost (Python package)
ModelRPython
Linear Model (LM)lm()sklearn.linear_model.LinearRegression; statsmodels.api.OLS
Generalized Linear Model (GLM)glm()sklearn.linear_model.LogisticRegression (classification); statsmodels.api.GLM
Generalized Additive Model (GAM)mgcv::gam(); gam::gam()pygam.LinearGAM / LogisticGAM; statsmodels.gam.api.GLMGam
Ridge / Lasso / Elastic Netglmnet::glmnet()sklearn.linear_model.Ridge, Lasso, ElasticNet
K-Nearest Neighbors (KNN)class::knn(); caret::train(method="knn")sklearn.neighbors.KNeighborsClassifier / Regressor
LDA / QDAMASS::lda() / MASS::qda()sklearn.discriminant_analysis.LinearDiscriminantAnalysis / QuadraticDiscriminantAnalysis
Naive Bayese1071::naiveBayes()sklearn.naive_bayes.GaussianNB / MultinomialNB
Decision Treesrpart::rpart(); tree::tree()sklearn.tree.DecisionTreeClassifier / Regressor
Random ForestrandomForest; rangersklearn.ensemble.RandomForestClassifier / Regressor
Boostinggbm; xgboost; lightgbmxgboost; lightgbm; sklearn.ensemble.GradientBoostingClassifier
Support Vector Machine (SVM)e1071::svm()sklearn.svm.SVC / SVR
Neural Network (shallow)nnet::nnet()sklearn.neural_network.MLPClassifier / MLPRegressor
ArchitectureRPython
General deep learning frameworkkeras (R interface to Keras/TensorFlow); torch (R port of PyTorch)TensorFlow/Keras; PyTorch
CNNkeras (layer_conv_2d())torch.nn.Conv2d; tf.keras.layers.Conv2D
RNN / LSTM / GRUkeras (layer_lstm(), layer_gru())torch.nn.LSTM / GRU; tf.keras.layers.LSTM
Transformerstorch (R) + transformer-style custom layers (limited ecosystem)transformers (Hugging Face); torch.nn.TransformerEncoder; tf.keras
Autoencoderskeras (custom encoder/decoder models)torch/tf.keras (custom encoder/decoder models)
GANskeras (custom generator/discriminator)torch/tf.keras (custom generator/discriminator)
Graph Neural Networks (GNN)Limited native support; often via reticulate calling PythonPyTorch Geometric; DGL (Deep Graph Library)
Dropout / Batch Normkeras::layer_dropout(), layer_batch_normalization()torch.nn.Dropout, BatchNorm2d; tf.keras.layers.Dropout
Early stoppingkeras::callback_early_stopping()tf.keras.callbacks.EarlyStopping; PyTorch: manual loop or pytorch-lightning callback
PurposeRPython
Unified modeling front-end (train many models, common CV interface)caret; tidymodels (parsnip + rsample + tune + yardstick)scikit-learn (Pipeline, GridSearchCV, cross_validate)
Automated ML / large-scale model comparisonmlr3 ecosystemauto-sklearn; TPOT; PyCaret
Experiment tracking (for deep learning tuning, refer ‘What Changed in Practice’)mlflow (R interface)mlflow; Weights & Biases (wandb)

Practical note: In Python, almost everything from data splitting through classical ML (refer ‘MSE for Estimating a Parameter θ’ – refer ‘The Model Spectrum’) sits inside a single consistent library, `scikit-learn`, which is why it’s the default reference implementation in most modern courses. In R, the equivalent breadth requires combining several packages – historically `caret`, more recently the `tidymodels` collection – since R’s ecosystem grew as separate specialized packages (`glmnet`, `randomForest`, `e1071`, `mgcv`) rather than one unified library.


Model selection picks the best model using training + validation data (repeatable, biased if reported as final performance). Model assessment reports the chosen model’s true performance using the test set (done once, unbiased). Keeping them separate – and being alert to validation-set leakage through repeated hyperparameter tuning – is what keeps a performance estimate honest.

Scroll to Top