Linear Models

This notes provides the necessary mathematical, statistical and computational details of Linear Modeling (LM).

  1. k: Number of predictors
  2. p = k + 1
  3. $X$: Model Matrix of predictors including constant term  $$ X = \left(\begin{array}{cccc} x_{11} & x_{12} & \cdots & x_{1p} \\ x_{21} & x_{22} & \cdots & x_{2p} \\ \vdots & \vdots & \ddots & \vdots \\ x_{n1} & x_{n2} & \cdots & x_{np} \end{array}\right)_{n\times p} =  [x_{ij}]_{n\times p}$$ such that $x_{i1} = 1 ~~~~~~\forall i = 1,2\cdots n$
  4. $Y$: Response variable $$ Y = \left(\begin{array}{c} y_{1} \\ y_{2} \\ \vdots \\ y_{n} \end{array}\right)_{n\times 1} =\left (\begin{array}{cccc} y_1 & y_2 & \cdots &y_n \end{array}\right)^T_{n\times 1}$$
  5. $\mu:$ Mean Response $$ \mu = E[Y]= \left(\begin{array}{c} \mu_{1} \\ \mu_{2} \\ \vdots \\ \mu_{n} \end{array}\right)_{n\times 1} =\left(\begin{array}{cccc} \mu_1 & \mu_2 &\cdots & \mu_n \end{array}\right)^T_{n\times 1}$$
  6. $\beta$: Parameter matrix $$ \beta = \left(\begin{array}{c} \beta_{1} \\ \beta_{2} \\ \vdots \\ \beta_{p} \end{array}\right)_{p\times 1}= \left(\begin{array}{cccc} \beta_1 & \beta_2 &\cdots& \beta_p \end{array}\right)^T_{p\times 1}$$
  7. $\epsilon$: Random error matrix. The errors are assumed to have mean zero and unknown variance $\sigma^2$; also, it is assumed that the errors are uncorrelated. $$\epsilon = \left(\begin{array}{cc} \epsilon _{1} \\ \epsilon _{2} \\ \vdots \\ \epsilon _{n} \end{array}\right)_{n\times 1}= \left(\begin{array}{c} \epsilon_1 & \epsilon_2 & \cdots & \epsilon_n \end{array}\right)^T_{n \times 1}$$

A Linear Model is $Y=X\beta + \epsilon$. If the predictors are controlled by the researchers in that X is measured or observed with negligible error and Y is a random variable then, $$E(Y|X)=\sum \beta_jx_{ij}$$ and $$V(Y|X)=\sigma^2$$

Mean of Y is a linear function of Y. Since the errors are uncorrelated, the responses are also uncorrelated.


  1. Linearity $Y=X\beta+\epsilon$
  2. $X$ is a full rank matrix
  3. $E[\epsilon|X] = 0$; given any $X$, no information of predictors convey any information about errors
  4. $E[\epsilon\epsilon^T|X] = \sigma^2I$ where $I$ is the identity matrix of order $n$; this helps to handle constant variance and no autocorrelation between the errors
  5. $\epsilon \sim N(0,\sigma^2I)$, largely helpful for inference about $\hat\beta$

    Form of $E[\epsilon\epsilon^T|X]$

    $$E[\epsilon\epsilon^T|X] =\left (\begin{array}{cccc} V(\epsilon_{1}) & COV(\epsilon_{1}\epsilon_{2}) &   \cdots  & COV(\epsilon_{1}\epsilon_{n})\\ COV(\epsilon_{2}\epsilon_{1}) & V(\epsilon_{2}) &   \cdots  & COV(\epsilon_{2}\epsilon_{n})\\ \vdots & \vdots & \ddots &\vdots\\   COV(\epsilon_{n}\epsilon_{1}) & COV(\epsilon_{n}\epsilon_{2})&   \cdots  & V(\epsilon_{n}) \end{array} \right)_{n\times n}$$

If the assumptions are constant variance and zero auto correlation then the above matrix is $\text{diag}(\sigma^2)_{n\times n} = \sigma^2I_{n\times n}$


The notion of Least Square Method (LSM) is to minimize the value of $||Y-\hat \mu||^2$

$$||Y-\hat \mu||^2 = \sum_{i=1}^{n}(y_i-\hat\mu_i)^2= \sum_{i=1}^{n}(y_i-\sum_{j=1}^{p}\hat\beta_jx_{ij})^2$$

Similar expression can be obtained from the notion of MLE when we assume that $\forall i = 1,2,\cdots, n~~~~ y_i\sim \text{Normal}(\mu_i,\sigma^2)$.

$\Rightarrow$ the log likelihood (excluding constant) is $$l(\beta|Y) = -\frac{\sum_i (y_i-\mu_i)^2}{2\sigma^2}$$ Hence, maximizing likelihood is equivalent to minimize the $\sum_i (y_i-\mu_i)^2$

Hence $l(\beta) = \sum_{i=1}^{n}(y_i-\sum_{j=1}^{p}\beta_jx_{ij})^2$ yields p equations while we equate $\frac{\partial~ l }{\partial~ \beta_j}$ to zero $\forall j=1,2,\cdots, p$

That is $\sum_i(y_i-\mu_i)x_{ij}=0 ~~~~~\forall j=1,2,\cdots, p$

Hence the LS estimates satisfy $\sum_iy_ix_{ij}=\sum_i\hat\mu_ix_{ij}~~~~~\forall j=1,2,\cdots, p$


$l(\beta)=||Y-X\beta||^2=(Y-X\beta)^T(Y-X\beta)$

=$(Y^T-\beta^TX^T)(Y-X\beta)$

=$Y^TY-Y^TX\beta-\beta^TX^TY+\beta^TX^TX\beta$

It can be observed that $Y^TX\beta$ is a constant and hence $\beta^TX^TY=(Y^TX\beta)^T=Y^TX\beta$

Hence, $l(\beta)=Y^TY-2Y^TX\beta+\beta^TX^TX\beta$

  1. $\frac{\partial (a^T\beta)}{\partial\beta}=a$
  2. $\frac{\partial (\beta^TA\beta)}{\partial\beta}=(A+A^T)\beta=2A\beta ~~~~~\text{(if A is symmetric)}$

Using these results $\frac{\partial ~l}{\partial\beta}=-2X^TY+2X^TX\beta=-2X^T(Y-X\beta)$

Hence, $\frac{\partial ~l}{\partial\beta}=0 \Rightarrow X^TX\beta=X^TY$

$\Rightarrow \hat\beta= (X^TX)^{-1}X^TY$

Also, $\frac{\partial^2 ~l}{\partial\beta^2}=2X^TX$ which is positive definite and hence the minimum exists at $\hat\beta$


Let us define hat matrix H as $H=X(X^TX)^{-1}X^T$. The fitted values are $\hat\mu=X\hat\beta=HY$

Also, $E(\hat\beta)=\beta$ and $V(\hat\beta)=\sigma^2(X^TX)^{-1}$. These results are obtained from the expectation of matrices

  1. $E(AY)=AE(Y)$
  2. $V(AY)=AV(Y)A^T$

The estimate of $\sigma^2$ is $s^2=\frac{\sum(y_i-\hat\mu_i)^2}{n-p}$. It can be noted that $V(\hat\beta)$ is the main diagonal elements of the matrix $\hat{\sigma^2}(X^TX)^{-1}$.

Using the point estimates of $\beta$ and associated $V(\hat\beta)$ it is possible to obtain 100(1-$\alpha$)% confidence intervals for $\hat\beta$.

Let us define the following

1. Total sum of squares TSS = $\sum(y_i-\overline {Y})^2$

2. Error sum of squares SSE = $\sum(y_i-\hat \mu_i)^2$

3. Regression sum of squares SSR = $\sum(\hat \mu_i-\overline {Y})^2$

Following results from these three quantities are important

TSS = SSE + SSR

$\frac{\text{SSE}}{\sigma^2}\sim \chi^2(n-p)$

$\frac{\text{SSR}}{\sigma^2}\sim \chi^2(p-1)$

It can be noted that Error Mean square  MSE = $\frac{\text{SSE}}{n-p}$

Regression Mean square  MSR = $\frac{\text{SSR}}{p-1}$

$\text{F-Ratio}=\frac {\text{MSR}}{\text{MSE}} \sim \text{F}_{p-1,n-p}$

Using TSS,SSE,and SSR it is possible to perform the global test for statistical significance of regression;

$$\text{H}_0:\beta_1=\beta_2=\beta_3=\cdots\cdots\beta_k=0 ~~\text{vs}~~ \text{H}_1: \beta_j\ne 0$$ at least for one $j=1,2,\cdots, ~k$

If $\text{F}_0$ is calculated F- ratio  then reject $\text{H}_0$ if  $\text{F}_0>\text{F}_{p-1,n-p}$

In a similar way, a test on individual regression coefficients can be performed for their statistical significance

That is test $$\text{H}_0:\beta_j= 0 ~~\text{vs}~~ \text{H}_1: \beta_j\ne 0 ~~~ \forall j = 1,2,\cdots~p$$

Reject $\text{t}_0$ if  $\text{t}_0>\text{t}_{\alpha/2,n-p}$ where $\text{t}_0=\frac{\hat\beta_j}{\sqrt{(V(\hat\beta_j))}}$

  1. $\hat\beta=(X^TX)^{-1}X^TY=(X^TX)^{-1}X^T(X\beta+\epsilon)$ which may be reduced to $\hat\beta=\beta+(X^TX)^{-1}X^T\epsilon$
  2. This indicates $\hat\beta$ is linear in $\epsilon$
  3. Also to note $E[\hat\beta]=\beta$ (Unbiased)
  4. It can be proved that $\hat\beta$ has minimal variance among all linear unbiased estimate
  5. $COV(\hat\beta) = E[(\hat\beta-\beta)(\hat\beta-\beta)^T]$=$(X^TX)^{-1}X^TE[\epsilon\epsilon^T]X(X^TX)^{-1}$=$(X^TX)^{-1}X^T\Sigma X(X^TX)^{-1}$which is a $p\times p$ matrix
  6. Under constant variance and zero auto correlation this becomes $COV(\hat\beta)=\sigma^2(X^TX)^{-1}$
  7. Hence, with normality assumption on errors, we have $$\hat\beta\sim N(\beta, \sigma^2(X^TX)^{-1})$$
  8. Also the estimate of $\sigma^2$ is $\frac{e^Te}{n-p}$
  9. With this unbiased estimate of $\beta$, MSE($\hat\beta$)=V($\hat\beta$)

  1. Data = fit + residuals; $Y =  X \beta +\epsilon$ Fitted model is $\hat Y =  X \hat\beta$ and * The difference between the observed value $y_i$ and the fitted value $\hat y_i$ is a residual $\epsilon_i= y_i-\hat y_i$
  2. Standardized residual (or internally studentized) divides each raw residual by its standard error (For an observation $i$)
  3. Studentized residual is based on the fit of the model after excluding $i^{th}$ observation (For an observation $i$)
  4. Leverage determines the precision of estimating $\mu$. Larger the values $y_i$ may have a large influence on $\hat \mu_i$. When the leverage is relatively large, $y_i$ is highly correlated with $\hat \mu_i$.
  5. Larger leverage with high residual may be influential observation.
  6. Akaike and Bayesian Information Criterions (AIC and BIC) for comparing models. The competing models need not be nested or based on the same family of distributions for the random component
  1. Raw or response residuals are the difference between the observed and the fitted value $y_i-\hat y_i$
  2. Pearson residuals are raw residuals divided by the standard error of the observed value se($y_i$)
  3. Standardized Pearson residuals are the raw residuals divided by their standard error se($y_i-\hat y_i$)

$r_i=\frac{y_i-\hat\mu_i}{s\sqrt{1-h_{ii}}}$ where $\text{h}_{ii}$ is the $\text{i}^{th}$ main diagonal element of the hat matrix H; $\text{s}^2$ is the estimator of $\sigma^2=\text{V}(y_i)$

If the linear model holds, then $-3\le r_i \le 3$

For observation $i$, $D_i$ combines the magnitude of residual and the location in X space to understand the influence.It is based on the change in $\hat\beta$ when observation $i$ is removed from the data set.

$\text{D}_i=\frac{h_{ii}~(y_i-\hat\mu_i)^2}{ps^2(1-h_{ii})^2}$

$D_i$ more than 1 may require attention. Usually larger D ($\text{D}_i\ge1$), occurs when both Standardized Residual $r_i$ and the leverage are relatively large

Sometimes we might include irrelevant predictors; we use model selection criterions (minimizing AIC or BIC). AIC can aid in variable selection

AIC =$-2\Big[l(\hat\beta)- p_1\Big]$ where $p_1$ is the number of parameters in the model. In the case of linear model this includes the constant variance $\sigma^2$

For linear model, $l(\hat\beta)=-\frac{n}{2}~\log(s^22\pi)-\frac{1}{2s^2}\sum(y_i-\mu_i)^2$

  • Measures both leverage and the effect of large residual. If the absoluate value of $\textrm{|DFBETAS}_{j,i}|$ exceeds $\frac{2}{\sqrt n}$ then the $i^{th}$ observation may require an attention
  • Measures the effect of the deletion influence of $i^{th}$ observation on the fitted or predicted value.
  • If the absoluate value of $\textrm{|DFFITS}_{i}|$ exceeds $\frac{2}{\sqrt{p/n}}$ then the $i^{th}$ observation may require an attention

Observations with |COVRATIO-1| near to or larger than  $\frac{3p}{n}$ may require an attention


This is one of the classical procedures that has been implemented in most (or all) of the analytical software systems / platforms and in analytical programming environments. For eg, here is R implementation

Scroll to Top