Point Estimation 2 – ML

This notes confines to MLE in obtaining Point Estimators for parameters

  • Estimator
  • Estimate
  • Likelihood function
  • Method of Moments (MOM)
  • Maximum Likelihood Estimates (MLE)
  • Bias, Variance and Mean Squared Error (MSE) of an estimator
  1. Method of finding estimators
  2. Criteria to find a “best” estimator
  3. Assessing tools – goodness of estimator
  1. $\theta:$ Parameter
  2. $f(x_i|\theta):$ Probability density/mass function
  3. $L(\theta|X):$ Likelihood function
  4. $l(\theta|X)= \ln L(\theta|X):$ Log Likelihood function

An estimator of $\tau(\boldsymbol\theta)$, a function of parameter is any function $W(X_1,X_2,\cdots,X_n)$ of a sample; that is any statistic is a point estimator  

Here, $\boldsymbol{\theta}=(\theta_1, \theta_2,\cdots,\theta_k)$

The process of MLE is to maximize the likelihood function $L(\theta~|~X) = f(X~|~\theta)$, the joint pdf or pmf of the sample X

that is to find a parameter value $\hat\theta(X)$ at which $L(\theta~|~X)$ attains its maximum as a function of $\theta$

First, we consider the analytical method (calculus approach) using second derivative test to find maximum of a function.

However, second derivatives are not explicitly derived in this note and left as an exercise.

Let $X_1,X_2,\cdots,X_n \sim \text{Bernoulli} (\theta)~~~~~~~ x = 0,~1;~~~~ 0 < \theta<1$

$f(x_i|\theta) = \theta^{x_i} (1-\theta)^{1-x_i} ~~~~~ i = 0,~1$

$\Rightarrow L(\theta|X) = \prod_{i=1}^n f(x_i|\theta)$

$L(\theta|X) =\prod_{i=1}^n \theta^{x_i} (1-\theta)^{1-x_i}$

$L(\theta|X) = \theta^{\sum x_i} (1-\theta)^{n-{\sum x_i}}$

$l(\theta|X) = \sum x_i ~\ln\theta +(n-\sum x_i) ~~\ln(1-\theta)$

$\frac{dl}{d\theta} = \sum x_i \frac{1}{\theta} + \frac{(n-\sum x_i)}{1-\theta} (-1)$

$\frac{dl}{d\theta} = 0$

$\Rightarrow \frac{1}{\theta}\sum x_i – \frac{n-\sum x_i}{1-\theta} = 0$

$\Rightarrow (1-\theta) \sum x_i – (n-\sum x_i) \theta = 0$

$\Rightarrow \sum x_i – n\theta = 0$

$\hat\theta_{ML} = \frac{\sum x_i}{n} = \overline {X}$

$X_1,\cdots,X_n \sim \text{Poisson}(\theta)$

$f(x_i~|~\theta) = e^{-\theta}~ \frac{\theta~^{x_i}}{x_i~!}$

$L(\theta|X) = \prod e^{-\theta}~ \frac{\theta~^{x_i}}{x_i!}$

$=e^{-n\theta}~ \frac{\theta~^{\sum x_i}}{\prod x_i!}$

$l(\theta|X) = -n~\theta + \sum x_i \ln \theta + k$

$\frac{\partial~l}{\partial~\theta} = 0$

$\Rightarrow -n + \sum x_i~ \frac{1}{\theta} = 0$

$\Rightarrow n =\frac{\sum x_i}{\theta}$

$\Rightarrow \hat\theta_{ML} = \overline {X}$

Let $X_1,X_2,\cdots,X_n \sim \text{Geometric}~(\theta)$

$\Rightarrow f(x_i,\theta) = \theta (1-\theta)^{x_i -1 } ~~~x_i = 1,2,3,…$

$L(\theta|X) = \prod_{i=1}^{n} \theta (1-\theta)^{x_i -1}$

$L(\theta|X) = \theta^n (1-\theta)^{\sum x_i-n}$

$l(\theta|X) = n \ln \theta + (\sum x_i – n ) \ln(1-\theta)$

$\frac{dl}{d\theta} = n \frac{1}{\theta} + \frac{\sum x_i -n}{1-\theta} (-1)$

$\frac{dl}{d\theta} = 0$

$\Rightarrow \frac{n}{\theta} -\frac{\sum X_i – n}{1-\theta} = 0$

$\Rightarrow n(1-\theta) – (\sum X_i – n) \theta = 0$

$\Rightarrow n – (\sum x_i) \theta = 0$

$\hat\theta_{ML} = \frac{n}{\sum x_i} = \frac{1}{\overline{X}}$

Let $X_1,\cdots,X_n \sim \text{Gamma}~(\alpha,\beta)$ $X>0$ Assume $\alpha$ is known

$L(\theta|X) = \prod f(x_i~|~\theta)$ $\theta = \beta >0$ which is a scale parameter

$L(\theta|X) = \prod_{i=1}^n \frac{1}{\Gamma\alpha~\beta^\alpha}~ x_i~^{\alpha – 1}~ \exp\Big[-\frac{x_i}{\beta}\Big]$

$= (\frac{1}{\Gamma \alpha})^n \beta^{-\alpha~n}~ \prod x_i~^{\alpha-1} \exp\Big[-\frac{x_i}{\beta}\Big]$

$l(\theta|x) = k -~ \alpha~n~\ln\beta + (\alpha -1 )~\sum \log X_i – \frac{\sum \alpha_i}{\beta}$ where $k$ is a constant free from the parameter $\beta$

$\frac{\partial~l}{\partial~\beta} = 0$

$\Rightarrow -~\alpha~n~ \frac{1}{\beta} + \frac{\sum X_i}{\beta^2} = 0$

$\frac{\sum X_i}{\beta} = -\alpha~n$

$\Rightarrow \hat\beta_{ML} = \frac{\sum X_i}{\alpha~n}$

$\hat\beta_{ML}= \frac{\overline {X}}{\alpha}$

If $\alpha >0$ is also unknown.

$l(\theta|x) = -n~\ln \Gamma\alpha ~-~\alpha~n~\ln \beta + (\alpha -1) \sum \ln X_i – \frac{\sum X_i}{\beta}$ and $\theta= \alpha ~ \& ~ \beta$

$\frac{\partial~ l}{\partial~ \alpha} = 0$

$\Rightarrow -\frac{n~(\Gamma \alpha)’}{\Gamma \alpha} – n~\ln \overline {X} + \alpha~n~\frac{1}{\alpha} + n~\ln \alpha + \sum \ln X_i – n  = 0$

$\Rightarrow n~\log \alpha – \frac{n~(\Gamma\alpha)’}{\Gamma \alpha} = n~\ln \overline {X} – \sum \ln x_i$

$\Rightarrow \ln \alpha – \frac{(\Gamma\alpha)’}{\Gamma \alpha} = \ln \overline {X} – \frac{\sum \ln x_i}{n}$

solving this equation in $\alpha$ provides ML estimator for $\alpha$

$X_1,\cdots,X_n \sim \text{Exponential}(\theta)$

$f(x~|~\theta) = \theta~e^{-\theta x}~~~~~~~x>0;~~~~\theta>0$

$L(\theta|X) = \prod_{i=1}^{n} f(x_i~|~\theta)$

$L(\theta|X) = \prod \theta~e^{-\theta x_i}$

$L(\theta|X) = \theta^n e^{-\theta\sum x_i}$

$l(\theta|X) = n~\ln \theta – \theta \sum x_i$

$\frac{dl}{d\theta} = 0$

$\Rightarrow \frac{n}{\theta} – \sum x_i = 0$

$\Rightarrow \frac{n}{\theta} = \sum x_i$

$\Rightarrow \hat\theta_{ML} = \frac{n}{\sum x_i} = \frac{1}{\overline {X}}$

Let $X_1,\cdots,X_n \Rightarrow NBD(r;\theta)$

If x is the number of trials required to get $r$ successes in a sequence of Bernoulli experiment

$f(x_i~|~\theta)= {x_{i}- 1~ \choose r-1}~ \theta^r~ (1-\theta)^{x_i – r} ~~~ x_i = r,r+1,\cdots$

$L(\theta|X) = \prod_{i=1}^n {x_i -1 \choose x_i -r} \theta^r (1-\theta)^{x_i – r}$

$= \theta^{nr} (1-\theta)^{\sum x_i – nr} \prod {x_i -1 \choose x_i – r}$

$l(\theta|X) = n~r\ln\theta + (\sum x_i – n~r)~\ln (1-\theta) + k$

$\frac{\partial~l}{\partial~ \theta} = 0 \Rightarrow$

$\frac{n~r}{\theta} + \frac{\sum x_i – n~r}{1-\theta}(-1) = 0$

$n~r(1-\theta) – (\sum x_i – nr) \theta = 0$

$\Rightarrow n r – \theta\sum x_i = 0$

$\hat\theta_{ML}  = \frac{nr}{\sum x_i}$

$\hat\theta_{ML}  = \frac{r}{\bar X}$

If $x_i$ referes the number of failures preceeding rth successes, then

$f(x_i~|~\theta) = {x_i~+~r-1~~\choose r-1}~ \theta^r~(1-\theta)^{x_i} ~~~~x = 0,1,2,\cdots$

$L(\theta|X) = \prod {x_i+r-1~ \choose r-1} ~\theta^r ~(1-\theta^{x_i}) ~~~~ x_i = 0,1,2,\cdots$

$= \theta^{nr}~ (1-\theta)^{\sum x_i}~ \prod {x_i+r-1~ \choose x_i}$

$l(\theta|X) = n~r~\ln \theta + \sum x_i~ \ln(1-\theta) + k$

$\frac{\partial~l}{\partial~\theta} = 0 \Rightarrow$

$\frac{nr}{\theta} + \frac{\sum x_i}{1-\theta}(-1) = 0$

$\Rightarrow n~r~(1-\theta) – (\sum x_i)~\theta = 0$

$\Rightarrow n~r-n~r~\theta – (\sum x_i) ~\theta = 0$

$\Rightarrow n~r – \theta(n~r+\sum x_i) = 0$

$\Rightarrow \theta_{MLE} = \frac{n~r}{n~r+\sum x_i}$

$= \frac{n~r}{\sum(x_i + r)}$

$=\frac{n~r}{\sum y_i}$

$\hat\theta_{ML} = \frac{r}{\bar y}$

Even if $x_i$ is defined in different manner, MLE of the proportion parameter is same

Also to note that when r = 1, we get MLE for the parameter in Geometric distribution.  

Let $X_1,X_2,\cdots,X_n) \sim \text{Beta}(1,\theta)$  

$\Rightarrow f(x_i ~|~\theta) = \theta (1-x_i)^{\theta-1}~~~~~~~~~0 < x < 1;~ \theta>0$

$L(\theta|X) = \prod f(x_i ~|~\theta)$

$= \prod \theta (1-x_i)^{\theta-1}$

$=\theta^n \prod (1-x_i)^{\theta-1}$

$l(\theta|X) = n~\ln\theta + (\theta-1) \sum \ln(1-x_i)$

$\frac{\partial~l}{\partial~\theta} = 0$

$\Rightarrow \frac{n}{\theta} + \sum \ln (1-x_i) = 0$

$\Rightarrow \hat\theta_{ML} = \frac{n}{-\sum \ln (1-x_i)}$

Following situations help to understand the limitations of analytical method (calculus approach). Alternative approach is needed to maximize the likelihood functions

Let $X_1,X_2,\cdots,X_n \sim U(\theta – 1/2,~\theta + 1/2)$

$f(x_i,\theta) = 1 ~~~~~~~\theta – 1/2 < x_i < \theta + 1/2$

$L(\theta|X) = \prod_{i=1}^n 1 ~~~~$ whenever $~~\theta – 1/2 < x_i < \theta + 1/2$

$i.e.~ Y_1 = Min X_i > \theta – 1/2$

$Y_n = Max X_i < \theta + 1/2$

or $Y_n – 1/2 < \theta < Y_1 + 1/2 ~~~~~~~~~~~~~~~~~~~~~~(1)$

Any $\theta$ satisfies $(1)$ could be an MLE

$\hat\theta_{ML}  = \frac{Y_1+1/2+Y_n-1/2}{2}$

$= \frac{Y_n + Y_1}{2}$

This is average of maximum and minimum of X

This idea can be used to the following situations also

$X_1,X_2,\cdots,X_n \sim u(0,\theta)$ to obtain $\hat\theta_{ML} = Max X_n$  

Let $X_1,\cdots,X_n \sim u(\mu – \sqrt 3 \sigma , \mu + \sqrt 3 \sigma)$  

$f(x_i,\theta) = \frac{1}{2\sqrt 3 ~\sigma} ~~~~$if $~\mu – \sqrt 3 \sigma < x_i < \mu + \sqrt 3 \sigma$

$\Rightarrow L(\theta|X) = \prod_{i=1}^n \frac{1}{2\sqrt3 \sigma}$

$L(\theta|X) = (\frac{1}{2\sqrt3 \sigma})^n  ~~~~ x_i \in [ \mu – \sqrt3 \sigma ,\mu + \sqrt3 \sigma)$

Hence, minimum of X, $Y_1 > \mu – \sqrt3 \sigma$ and maximum of X $Y_n < \mu + \sqrt3 \sigma$

$\hat\mu_{ML} = \frac{Y_1+Y_n}{2}$

$\sqrt3 \sigma = Y_n – \mu$

$\sigma = Y_n – (\frac{Y_1 + Y_n}{2})$

$\sigma = \frac{Y_n – Y_1}{2 \sqrt3}$

Optimization for function of several variables

From Matrix Theory

Let $A = (a_{ij})_{n\times n}$. A $k \times k$ submatrix of A formed by deleting $n-k$ rows of $A$, and the same $n-k$ columns of $A$, is called $\textbf{principal sub matrix of A}$.

The determinant of a principal submatrix of A is called a $\textbf{principal minor}~\Delta_k$ of $A$.

The $\textbf{leading principal minor of A}$ of order $k$ is the minor of order $k$ obtained by deleting the last $n-k$ rows and columns. We denote leading principal minor of $A$ as $L\Delta_k$

For example, if $A = (a_{ij})_{2\times 2}$ then $L\Delta_1=a_{11}$ and $L\Delta_2= det(A)=a_{11}a_{22}-a_{21}a_{12}$

Let $A$ be a symmetric $n \times n$ matrix. Then we have:

1. $A$ is positive definite $\iff$ $L\Delta_k>0$ for all leading principal minors

2. $A$ is negative definite $\iff$ $(-1)^kL\Delta_k >0$ forall leading principal minors

Consider $$ A =\left(\begin{array}{cccc}   2 & -1 & 0 \\   -1 & 2 & -1 \\ 0 & -1 & 2 \end{array}\right) $$

$L\Delta_1=A_{11}=2$

$L\Delta_2= det\left(\begin{array}{cccc}   2 & -1 \\ -1 & 2\end{array}\right) = 3 $

$L\Delta_3=detA=4$

All $L\Delta_k$ are positive. $A$ is positive definite

Consider $$ A =\left(\begin{array}{cccc}   2 & -1 \\   -1 & -2 \end{array}\right) $$

$L\Delta_1=A_{11}=2>0$

$L\Delta_2=detA = -5<0$

$A$ is negative definite

Hessian Matrix

Let $f(\textbf{X})$ be a real valued function in $\mathbb{R^n}$ that is $f:\mathbb{R^n}\rightarrow\mathbb{R}$ $\textbf{X}=(x_1,x_2,\cdots,x_n)$

Gradient of $f$ is $\nabla f=\Big(\frac{\partial~f} {\partial~x_1},\frac{\partial~f} {\partial~x_2},\cdots,\frac{\partial~f} {\partial~x_n}\Big)$

Hessian matrix is $$H =\left(\begin{array}{cccc}   f_{11} & f_{12} & \cdots & f_{1n} \\   f_{21} & f_{22} & \cdots & f_{2n} \\ \vdots  & \vdots  & \ddots & \vdots  \\ f_{n1} & f_{n2} & \cdots & f_{nn} \end{array}\right) $$ where $f_{ij}=\frac{\partial^2~f} {\partial~x_i~\partial~x_j}~~\forall ~i, ~ j=1,2,\cdots,n$ and it is easy to observe that $f_{ij}=f_{ji}$ and $H$ is symmetric

Hessian Matrix and convexity

1. $H$ is Positive definite $\Rightarrow$ $f$ is convex

2. $H$ is Negative definite $\Rightarrow$ $f$ is concave

It is also to note that if $f$ is concave all stationary points $\{x_0/\nabla f(x_0)=0\}$ ($x_0$ is not on the boundary of the open set S) are global maximum points. Read for more details $\textbf{TOFI}$

$l(\theta~|~X) = 3\ln\theta + 2 \ln (1-\theta)$ so that $\nabla f = \frac{3-5\theta}{\theta(1-\theta)}$ and $H=-\Big[\frac{3}{\theta~^2}+\frac{2}{(1-\theta)~^2}\Big]$

It is easy to observe that at $\theta = 0.6, \nabla f = 0$; if $\theta < 0.6, \nabla f > 0$ and if $\theta > 0.6, \nabla f < 0$

Also it can be observed that H is always negative.

This example illustrates the behaviour of $\nabla f$ and Hessian matrix near a local maximum as well as the notion of concavity.

$f(x_1,x_2) = x^2_1-x^2_2-3x_1x_2$ can easily be verified as a concave function by observing the Hessian matrix as negative definite $L\Delta_1=H_{11}>0$ and $L\Delta_2=\text{det}H <0$

$\textbf{NOTE: Maximization For Two Parameters}$ $\theta_1,~\theta_2$

1. Solve $\frac{\partial ~l}{\partial~\theta_1} = 0$ and $\frac{\partial~l}{\partial~\theta_2}=0$ for $\theta_1,~\theta_2$

2. Compute Hessian matrix.

3. Check whether $H$ is negative definite.

Let $X_1,X_2,…,X_n \sim N(\mu,\eta) ~~~\eta=\sigma^2$

$L(\theta|X) = \prod f(x_i|\theta)$ where $\theta = \mu~\&~\eta$

$L(\theta|X) = \prod \Big[\frac{1}{\sqrt{2\pi\eta}}\Big] e^{-\frac{1}{2\eta}(X_i – \mu)^2}$

$L(\theta|X) =  \Big[\frac{1}{\sqrt{2\pi\eta}}\Big]^n e^{-\frac{1}{2\eta}\sum(X_i – \mu)^2}$

$l(\theta|X) = k – \frac{n}{2} \ln \eta – \frac{1}{2\eta} \sum(X_i – \mu)^2$ where $k$ is constant free from parameters $\mu,\eta$

$\frac{\partial~l}{\partial~\mu} = 0$

$\Rightarrow -~\frac{1}{2\eta} 2\sum(X_i – \mu)^2 (-1) = 0$

$\Rightarrow \sum X_i -n\mu = 0$

$\Rightarrow \hat\mu = \overline {X}$

$\frac{\partial~l}{\partial~\eta} = 0$

$\Rightarrow -\frac{n}{2\eta} + \frac{1}{2\eta^2}\sum(X_i – \mu)^2  = 0$

$\Rightarrow \frac{1}{2\eta}\Big[-n + \frac{1}{\eta}\sum(X_i – \mu)^2\Big] = 0$

$\Rightarrow \frac{1}{\eta} \sum(X_i – \mu)^2 = n$

$\Rightarrow \eta = \frac{\sum(X_i – \mu)^2}{n}$

With $\hat\mu_{ML} = \overline {X}$ we have

$\hat \eta_{ML} = \widehat {\sigma^2_{ML}} =\frac{\sum(X_i – \overline {X})^2}{n}$

Scroll to Top