Expectation & Moments of Bivariate Distributions

Random variable characterises a random phenomenon by listing the range and the corresponding probability distribution (that is, pmf in the discrete case or pdf in the continuous case). As we can find this in example say if $p$ is the price of a certain commodity and $S_i$ its total sales, then we may be interested in finding the expected receipts $(R = PS)$ for the commodity. This notes deals with expectation $E(X)$ expectation of a random variable $X$ or more generally expectation $E[g(x,y)]$ of a random variable (or more variables, as in the preceeding example $E(R) = E((PS))$.

If $X$ and $Y$ are random variables and $f_{XY}(x, y)$ is the joint probability distribution, then the expected value of $g_{XY}(x, y)$, an arbitrary function of $X$ and $Y$ is

$$ E[g_{XY}(x, y)] =\begin{cases} \displaystyle\sum_{x} \sum_{y} g_{XY}(x, y) f_{XY}(x, y) & X \text{ and } Y \text{ are discrete} \\[10pt] \displaystyle\int_{-\infty}^{\infty} \int_{-\infty}^{\infty} g_{XY}(x, y) f_{XY}(x, y) \, dx dy & X \text{ and } Y \text{ are continuous} \end{cases} $$

Let $X$ and $Y$ each take on either value $1$ or $-1$ with their joint pmf $\frac{1}{4}$. Find $E(X + Y)$.

The joint pmf of $X$ and $Y$ is as shown in the table.

$$\begin{array}{c|cc} X \backslash Y & -1 & 1 \\ \hline -1 & \frac{1}{4} & \frac{1}{4} \\ 1 & \frac{1}{4} & \frac{1}{4} \end{array} $$

$$ E(X + Y) = \sum_{x} \sum_{y} (X + Y) f_{XY}(x, y) $$

$$ = (-1)(-1)f(-1,-1) + (-1)(1)f(-1,1) + 1(-1)f(1,-1) + (1)(1)f(1,1) $$

$$ = \frac{1}{4}(1 – 1 – 1 + 1) = 0. $$


If the joint pdf of $X$ and $Y$ is given by

$$f_{XY}(x, y) = \begin{cases} \dfrac{1}{y} & 0 < x < y,\ 0 < y < 1 \\ 0 & \text{elsewhere} \end{cases} $$

find the expected value of $\dfrac{X}{Y}$.

$$ E(\dfrac{X}{Y}) = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} \left( \frac{X}{Y} \right) f_{XY}(x, y) \, dy dx $$

$$ = \int_{y=0}^{1} \int_{x=0}^{y} \frac{x}{y} \cdot \frac{1}{y} \, dx dy $$

$$ = \int_{y=0}^{1} \left( \frac{x^2}{2} \right)_0^{y} \cdot \frac{1}{y^2} \, dy = \frac{1}{2} \int_0^1 (y^2 – 0) \frac{1}{y^2} \, dy = \frac{1}{2} $$


Let us recall the moments for a random variable. Now, extend the definition of moments for the two random variable case. Let us define the product moments of two random variables.

 The $r^{th}$ and $s^{th}$ product moment about the means of the random variables $X$ and $Y$ denoted by $\mu_{rs}$ is the expected value of $(X – \mu_X)^r (Y – \mu_Y)^s$; that is

$$ \mu_{rs} = E\left[(X – \mu_X)^r (Y – \mu_Y)^s\right] $$

$$= \begin{cases} \displaystyle\sum_{x} \sum_{y} (x – \mu_X)^r (y – \mu_Y)^s f_{XY}(x, y) & \text{when } X \text{ and } Y \text{ are discrete rvs} \\[10pt] \displaystyle\int_{-\infty}^{\infty} \int_{-\infty}^{\infty} (x – \mu_X)^r (y – \mu_Y)^s f_{XY}(x, y) \, dx dy & \text{when } X \text{ and } Y \text{ are continuous rvs} \end{cases} $$

for $r = 0, 1, 2 \cdots n; \ s = 0, 1, 2 \cdots n$.

 The $r^{th}$ and $s^{th}$ product moment about the origin of the random variables $X$ and $Y$, denoted by $\mu’_{rs}$ is the expected value of $X^r Y^s$; that is

$$\mu’_{rs} = E[X^r Y^s]$$

$$=n\begin{cases} \displaystyle\sum_{x} \sum_{y} x^r y^s f_{XY}(x, y) & X \text{ and } Y \text{ are discrete rvs} \\[10pt] \displaystyle\int_{-\infty}^{\infty} \int_{-\infty}^{\infty} x^r y^s f_{XY}(x, y) dx dy & X \text{ and } Y \text{ are continuous rvs} \end{cases} $$

for $r = 0, 1, 2, \cdots$ and $s = 0, 1, 2, \cdots$.

In both cases, if $r = s = 0$ then

$$ \mu’_{00} = E(X^0 Y^0) = 1 \text{ and} $$

$$ \mu_{00} = E[(X – \mu_X)^0 (Y – \mu_Y)^0] = 1  $$

Further $\mu_{11}$ is of special importance in statistics because it indicates the relationship between the values of the two random variables $X$ and $Y$. We define this as covariance of $X$ and $Y$ and it is denoted by $\sigma_{XY}$, $\text{Cov}(X, Y)$ or $C(X, Y)$ that is $\text{Cov}(X, Y) = \mu_{11} = E[(X – \mu_X)(Y – \mu_Y)]$. Let us write $\mu_X$ as $E(X)$ and $\mu_Y$ as $E(Y)$ respective means of $X$ and $Y$, so that $\text{Cov}(X, Y) = E[(X – E(X))(Y – E(Y))]$. see property (property 9) for the computation of $\text{Cov}(X, Y)$.

Let $X, Y$ be any two random variables. $f_{XY}(x|y)$ is the value of the conditional pdf of $X$ given $Y = y$ at $x$. Let $u(X)$ be a function of $X$. Then, the conditional expectation of $u(X)$ given $Y = y$ denoted by $E[u(X)|y]$ is defined as

$$E[u(X)|Y = y] =\begin{cases}\displaystyle\sum_{x} u(X) f_{XY}(x|y) & \text{for discrete case} \\[10pt] \displaystyle\int_{-\infty}^{\infty} u(X) f_{XY}(x|y) dx & \text{for continuous case} \end{cases} $$

Such conditional expectation of $\phi(Y)$ (some function of $Y$) given $X = x$ can be defined in a similar way.

Two special cases of $u(X)$ (similarly $\phi(Y)$) are of more important that is $u(X) = X$ and $u(X) = X^2$.

The conditional mean of the random variable $X$ given $Y = y$ is defined as $E[X|Y = y]$ and

$$ E[X|Y = y] = \begin{cases} \displaystyle\sum_{x} x f_{XY}(x|y) & \text{for discrete case} \\[10pt] \displaystyle\int_{-\infty}^{\infty} x f_{XY}(x|y) dx & \text{for continuous case} \end{cases} $$

The conditional variance of $X$ given $Y = y$ is defined as

$$ V(X|Y = y) = E[(X – E(X|Y = y))^2 | Y = y] $$

Remember that we use the expression $V(X) = E(X^2) – E(X)^2$ to compute variance of a random variable $X$. In a similar way, the conditional variance can be calculated using an expression

$$ E[(X – E(X|Y = y))^2 | Y = y] = E[X^2|Y = y] – (E[X|Y = y])^2 $$

where $E[X^2|Y = y]$ is obtained from definition (1) with $u(X) = X^2$. All these measures can be defined for conditional expectations of $Y$ given $X = x$. One should not find it difficult to write the appropriate expressions.

Let $X$ and $Y$ have the joint pdf $f_{XY}(x, y) = x + y,\ 0 < x < 1,\ 0 < y < 1$; zero elsewhere. Find the conditional mean and variance of $Y$ given $X = x$.

Let us first find the conditional pdf of $Y$ given $X$.

$$ f_{XY}(y|x) = \frac{f_{XY}(x, y)}{f_X(x)} $$

$$ f_X(x) = \int_{-\infty}^{\infty} f_{XY}(x, y) \, dy $$

$$ = \int_0^1 (x + y) \, dy = \left( xy + \frac{y^2}{2} \right)_0^1 =\begin{cases} x + \frac{1}{2} & 0 < x < 1 \\ 0 & \text{elsewhere} \end{cases} $$

$$ \therefore f_{XY}(y|x) = \begin{cases} \dfrac{x+y}{x+\frac{1}{2}} & 0 < x < 1,\ 0 < y < 1 \\ 0 & \text{elsewhere} \end{cases} $$

$$ E[Y|X = x] = \int_{-\infty}^{\infty} y f_{XY}(y|x) \, dy = \int_0^1 y \frac{x+y}{x+\frac{1}{2}} \, dy = \frac{1}{x+\frac{1}{2}} \left( x \frac{y^2}{2} + \frac{y^3}{3} \right)_0^1 $$

$$ = \frac{2}{2x+1} \left( \frac{x}{2} + \frac{1}{3} \right) = \frac{3x+2}{6x+3}, \quad 0 < x < 1 $$

$$ E[Y^2|X = x] = \int_{-\infty}^{\infty} y^2 f_{XY}(y|x) \, dy = \int_0^1 y^2 \frac{x+y}{x+\frac{1}{2}} \, dy $$

$$ = \frac{1}{x+\frac{1}{2}} \left( x \frac{y^3}{3} + \frac{y^4}{4} \right)_0^1 = \frac{4x+3}{6(2x+1)} $$

$$ \therefore \quad V[Y|X = x] = E[Y^2|X = x] – E[Y|X = x]^2 $$

$$ = \frac{4x+3}{6(2x+1)} – \frac{(3x+2)^2}{9(2x+1)^2} $$

$$ = \frac{6x^2 + 6x + 1}{18(2x+1)^2}, \quad 0 < x < 1 $$

Let $f_{XY}(x, y)$ denote the joint pdf of two random variables $X$ and $Y$ and if $f_X(x)$ is the marginal pdf of $X$ then we have the conditional pdf of $Y$ given $X = x$ is

$$ f_{Y|X}(y|x) = \frac{f_{XY}(x, y)}{f_X(x)} $$

and the conditional mean of $Y$ given $X = x$ is given by

$$ E[Y|X = x] = \int_{-\infty}^{\infty} y f_{Y|X}(y|x) \, dy \quad \text{(or)} \quad \sum_{y} y f_{Y|X}(y|x) $$

accordingly $(X, Y)$ is continuous or discrete. This conditional mean of $Y$, given $X = x$ is a function of $x$ alone. Similarly, the conditional mean $X$, given $Y = y$ is a function of $y$ alone.

The conditional mean $E(Y|X = x)$ for a continuous distribution is called the regression function of $Y$ on $X$ and the graph of this function of $x$ is known as the regression curve of $Y$ on $X$ or sometimes it is called as the regression curve for the mean of $Y$. Similarly, we define $E(X|Y = y)$ as the regression curve of $X$ on $Y$ or regression curve for the mean of $X$.

If this curve is linear (straight line), it is called the line of regression between the variables. Otherwise regression is said to be curvilinear.

Further, the regression lines are obtained by the following two equations.

(a) Regression line of $Y$ on $X$ is

$$ Y = E(Y|X) = \overline{Y} + r \frac{\sigma_Y}{\sigma_X} (X – \overline{X}) \quad \text{and} $$

(b) Regression line of $X$ on $Y$ is

$$ X = E(X|Y) = \overline{X} + r \frac{\sigma_X}{\sigma_Y} (Y – \overline{Y}) \quad \text{where} $$

$$ \overline{X} = E(X); \ \overline{Y} = E(Y) \text{ and } r = \rho_{XY} $$

$b_{XY} = r \dfrac{\sigma_X}{\sigma_Y}$ and $b_{YX} = r \dfrac{\sigma_Y}{\sigma_X}$ are called the regression coefficients of $X$ on $Y$ and $Y$ on $X$ respectively.

1. $b_{XY}$ and $b_{YX}$ have always same sign.

2. Correlation coefficient is the geometric mean of regression coefficients that is $\rho_{XY} = \pm \sqrt{b_{XY} \cdot b_{YX}}$ where the sign of $\rho$ is the same as that of the regression coefficients.

3. Both coefficients cannot exceed one simultaneously.

4. Regression coefficients are independent of the change of origin but not of scale.

The two regression lines are not reversible. Also, the mean $(\overline {X}, \overline{Y})$ can be obtained as the point of intersection of the two regression lines.

If the conditional mean of $Y$ given $X = x$ is linear function of $x$ say $a + bx$, we say that the conditional mean of $Y$ is linear in $x$. In such cases, we write the conditional mean as

$$ E[Y|X = x] = E(Y) + \frac{\text{Cov}(X, Y)}{V(X)} (X – E(X)) $$

Similarly, the conditional mean of $X$ given $Y = y$ is linear in $y$ then we write

$$ E[X|Y = y] = E(X) + \frac{\text{Cov}(X, Y)}{V(Y)} (Y – E(Y)) $$

Scroll to Top