Introduction
Probability allows us to quantify the variability in the outcome of any experiment whose exact outcome cannot be predicted with certainty. A major problem that an engineer encounters is uncertainity. By using the theory of probability one can incorporate this uncertainity into analysis and thus make rational decisions. However, before we introduce probability there is a basic question to be answered. What do we mean when we say that the probability of an event is, say, 0.50, 0.02 or 0.85?
Historically, the oldest way of measuring uncertainities is the classical probability. It applies when all possible outcomes are equally likely, in which case we define the probability as follows.
If there are ‘$n$’ equally likely possibilities, of which one must occur and ‘$m$’ are regarded as favourable or as a “success”, then the probability of a success is given by $\dfrac{m}{n}$.
For example, if we wish to calculate the probability of getting a head when a fair coin is tossed, then there is one head ($= m$) among $n = 2$ outcomes (head, tail), so we get $P(\text{getting a head}) = \dfrac{m}{n} = \dfrac{1}{2}$.
However, the classical probability concept has a major disadvantage. That is, there are many situations in which the various probabilities cannot be regarded as equally likely. For example, if we are concerned with the question of whether a bus will come within 10 minutes, whether a batsman scores a century, or whether it will rain the next day.
Furthermore, in the above example of tossing a single fair coin, we calculated the probability of ‘success’ (getting a head) is $\dfrac{1}{2}$ or 50%. What we really mean by this number? Does it mean that every time we toss a coin there is a 50% chance of getting a head? If so, imagine that 10 of your friends toss a coin simultaneously. Accordingly, you should get 5 heads and 5 tails. But could be your observation from the experiment, it may be 8 heads, 2 tails or 10 heads, 0 tail or so on. then, what ‘probability’ really means?
Among the various probability concepts, most widely held is the frequency interpretation. It states that the probability of an event is the proportion of times the event would occur in a long run of repeated experiments. Mathematically, we can write this as
$$P(\text{an event}) = \lim_{n \to \infty} \dfrac{m}{n}$$
In the coin toss example, this concept explains the 50% as under the same conditions we will get head 50% of the time. But we cannot guarantee what will happen on any particular occassion – we may get a head or a tail. But if we keep records over a long period of time, we should find that the proportion of “successes” is very close to 50% or 0.5. In fact, we will study later (Chebyshev’s inequality) that we require 1,000,000 flips of a fair coin to get the probability is atleast 0.99 that the proportion of heads will fall between 0.495 and 0.505 i.e., 49.5% to 50.5%.
So, in accordance with the frequency interpretation of probability, we estimate the probability of an event by observing what fraction of the time similar events have occurred in the past.
Alternatively, probabilities can be expressed based on one’s belief with regard to the uncertainities that are involved, and they apply especially when there is no direct evidence available. Such probabilities are referred to as personal or subjective evaluations. Such probabilities are explained later (in example or Bayes’ Theorem).
After answering our first question, now it is time to frame two more questions, such as
- How to determine or to measure the numbers we call probabilities in practice?
- What are the mathematical rules which probabilities must obey?
Now, we define probabilities mathematically as the values of “additive” set functions. Here “additive” means the number which it assigns to the union of the two subsets which have no elements in common is the sum of the numbers assigned to the individual subsets.
In the preceeding discussion, we used the words, “experiment”, “event” frequently. Let us define them (and more) formally for our subsequent discussions.
An ‘experiment‘ is defined as a scientific procedure undertaken to make a discovery, test a hypothesis or demonstrate a known fact. Consequence of an experiment may be defined as an ‘outcome‘. The outcome may be the result of a direct measurement or count, or it may be an answer obtained after extensive measurements and calculations.
Example:
- We wish to measure the heights of students in a classroom. Then the outcome may be a range of numbers say 150–180 cms.
- On the otherhand, if our task is to choose an elective subject from a list, then outcome may be any one of the subjects from a list of say 10 subjects.
Definition
- An experiment whose outcomes cannot be predicted or obtained in prior is said to be a Random experiment.
- The set of all possible outcomes of a random experiment is said to be sample space.
- A subset of sample space is said to be an event.
Let us follow a conventional way of denoting Random Experiment a “$E$” sample space as ‘$S$’ and events say $A, B, C, \cdots$ or $A_1, A_2, A_3, \cdots$ (if we have countable events).
Generally, sample spaces are classified according to the number of elements that they contain. A sample space is either discrete or continuous. If a sample space has finitely many or countably infinite number of elements then it is said to be discrete. On the other hand, if the elements of a sample space constitutes a continuum (certain interval on the line of real numbers), then the sample space is said to be continuous.
In the above examples, if we let $S_1, S_2, \cdots S_{10}$ denote the 10 subjects, then the sample space is $S = \{S_1, S_2, \cdots, S_{10}\}$.
Also, the sample space of measurement of heights is a real number in $[150, 180]$.
Further, a discrete sample space is classified as finite or countably infinite.
For example, if the experiment is to select a President and a Vice President from among 25 members, the sample space has 600 elements (In Example 1.5, we discuss this later). So it is finite discrete space.
Consider a sample space results for the experiment of tossing a coin until a head is observed.
i.e., $S = \{H, TH, TTH, TTTH, \cdots\}$
which means $1^{st}$ time get head, or $2^{nd}$ time we get head or $3^{rd}$ time and so on.
So, if the set of all outcomes can be put into a one-to-one correspondence with the natural numbers is said to be countably infinite.
Now we define the term probability. Let $S$ be a sample space of a random experiment. Let $P(A)$ be the probability measure associated with event $A$. Then
$$P(A) = \dfrac{n(A)}{n(S)}$$
This $P(A)$ is a value of an additive set function that satisfies the following three conditions.
AXIOMS OF PROBABILITY
$\textbf{Axiom 1}$: $0 \le P(A) \le 1$ for each event $A$ in $S$.
$\textbf{Axiom 2}$: $P(S) = 1$.
$\textbf{Axiom 3}$: If $A$ and $B$ are two events which are disjoint then $P(A \cup B) = P(A) + P(B)$.
Also we review the following laws.
- Commutative
- (i) $A \cup B = B \cup A$
- (ii) $A \cap B = B \cap A$
- Associative
- (i) $A \cup (B \cup C) = (A \cup B) \cup C$
- (ii) $A \cap (B \cap C) = (A \cap B) \cap C$
- Distributive
- (i) $A \cup (B \cap C) = (A \cup B) \cap (A \cup C)$
- (ii) $A \cap (B \cup C) = (A \cap B) \cup (A \cap C)$
- De Morgan’s law
- (i) $(A \cup B)’ = A’ \cap B’$
- (ii) $(A \cap B)’ = A’ \cup B’$
In the set theoretical approach of calculating probability, it is necessary to know the following items.
$E$: Random experiment.
$S$: Sample space.
$A$: Defined event.
$n(S)$: Number of elements in $S$.
$n(A)$: number of elements in $A$.
The following examples illustrate this approach.
Example 1
Find the probability of getting exactly 2 heads if 3 fair coins are tossed simultaneously.
Let $E$: Tossing 3 coins simultaneously.
$S$: $\{HHH, HHT, HTH, THH, TTT, TTH, THT, TTH\}$ $\therefore n(S) = 8$
$A$: Getting exactly 2 heads $\{HHT, HTH, THH\}$ $\therefore n(A) = 3$
$$\therefore \quad p(A) = \dfrac{n(A)}{n(S)} = \dfrac{3}{8}$$
Example 2
Find the probability of selecting a multiple of 5 from the integer set 1 to 50.
Let $E$: Selecting on integer from 1 to 50
$S$: $\{1, 2, 3 \cdots 50\}$ $\qquad \therefore n(S) = 50$
$A$: Getting a multiple of $5 = \{5, 10, 15, 20, 25, 30, 35, 40, 45, 50\}$ $\qquad \therefore n(A) = 10$
$$\therefore \quad P(A) = \dfrac{n(A)}{n(S)} = \dfrac{10}{50} = \dfrac{1}{5}$$
From the above two examples, we can observe that probabilities can be computed by such exhaustive way of counting the possibilities. But this method is having a major disadvantage. If the above 2 examples are generalised or extended to larger values, then the total enumeration may be impractical.
i.e., If we do the experiment for 10 coins (instead of 3 as in Examples 1) we need to write 1024 (How?) outcomes in the set ‘$S$’.
Also, to find the probability of any event, the number of outcomes is important rather than the actual pattern of outcomes.
This idea leads to the introduction of some counting principles in the next section.
Counting Principle 1
Product rule
If an event occurs $m$ times, and another event occurs $n$ times then the number of ways both event occur is $mn$ ways.
i.e., If a procedure can be broken into 2 events. If there are $m$ ways to do the first event and $n$ ways to the second event after the first event has been done, then there are ‘$m \times n$’ ways to do the procedure.
Note:
If a certain thing can be done in $n_1$, ways and then a second thing can be done in $n_2$ ways and then a third thing in $n_3$ ways and so on the total number of ways of doing all the things is $n_1 \times n_2 \times n_3 \times \cdots$.
Example 3
If a test consists of 10 true-false questions, in how many ways can a student mark the test paper with one answer to each question?
Since each question can be answered in two ways (either ‘T’ or ‘F’), there are altogether.
$$2 \times 2 \times 2 \times \cdots 2 = 2^{10} = 1024 \ \text{Possibilities}$$
(10 times)
The sum rule
If a task can be done in $m$ ways and a second task in $n$ ways and if these tasks cannot be done at the same time, then there are $m + n$ ways to do either task.
Note:
We can extend the sum rule to more than two tasks. Suppose that the tasks (or events) $A_1, A_2, \cdots, A_m$ can be done in $n_1, n_2, \cdots, n_m$ ways and no two of these tasks can be done at the same time. Then the number of ways to do one of these tasks is $n_1 + n_2 + \cdots + n_m$.
Example 4
Consider a bag which consists of 4 red, 3 black, 3 white balls and two experiments $E_1$, and $E_2$.
Let $E_1$: Selection of a ball.
$S_1$: $\{R_1, R_2, R_3, R_4, B_1, B_2, B_3, W_1, W_2, W_3\}$ $n(S_1) = 10$
$A$: Selecting a white ball. $\{W_1, W_2, W_3\}$ $n(A_1) = 3$
$$\therefore \quad P(A_1) = \dfrac{3}{10}$$
i.e., Probability of selecting 1 white ball from 10 balls $= \dfrac{3}{10}$
$E_2$: Selection of 3 balls from the collection.
$S_2$: $\{R_1R_2R_3, R_1R_3R_4, R_1B_1B_2, R_1B_2B_3, R_1R_2W_1, R_1B_2W_3 \text{ and so on} \cdots\}$
$$n(S_2) = 120 \ (\text{How?})$$
$A_2$: Event that all the 3 are red balls.
$$= \{R_1R_2R_3, R_1R_2R_4 \cdots\}$$
$$n(A_2) = 4 \ (\text{How?}) \quad \text{and} \quad P(A_2) = \dfrac{4}{120} = \dfrac{1}{30}$$
From this the preceeding example, we can observe that selection (or arranging) of more than one object from a collection of ‘$n$’ distinct objects requires more than natural counting, and complete enumeration may be impractical sometimes.
Counting Principle 2
Permutation and Combination
We assume an arrangement of ‘$n$’ objects along a straightline or in a row of ‘$r$’ position where $n \ge r$.
Assume, the first position is filled by selecting an object from the whole set of $n$ objects, the second selection is made from the $n – 1$ objects which remain after the first selection has been made and the $r^{th}$ selection is made from the $n – (r-1) = n – r + 1$ objects which remain after the first $r – 1$ selections have been made.
Therefore by the product role, the total number of permutations of $r$ objects selected from a set of $n$ distinct objects is
$$nP_r = n(n-1)(n-2)\cdots n – r + 1 \quad \text{for } r = 1, 2, \cdots, n$$
Theorem 1
The number of permutations of $r$ objects selected from a set of $n$ distinct objects is
$$nP_r = \dfrac{n!}{(n-r)!}$$
$\textbf{Corollary}$
Hence arrangement of ‘$r$’ objects in ‘$r$’ places $= rP_r = \dfrac{r!}{(r-r)!} = r!$
Theorem 2
The number of ways in which $r$ objects can be selected from a set of $n$ distinct objects is
$$nC_r = \dfrac{n!}{r!(n-r)!}$$
Note:
1. From the preceeding discussions, we observe that the main difference between permutation and combination is the “order”.
i.e., ordered selection is permutation and mere selection (Irrespective of the position) is combination.
2. The formulas in Theorems 1 and 2 can be calculated using a calculator which directly yields the factorials and/or ratios of factorials.
Example 5
A car rental agency has 10 compact cars and 20 intermediate size cars. Four cars are randomly selected for a safety check. Find the following probabilities:
- Exactly two cars of each kind are selected.
- All four cars are of the same kind.
- The majority of the selected cars are compact.
$E$: Total number of ways of selecting $4$ cars from $30$ cars$(10+20)$:
$$n(S) = 30C_4$$
1. let $A1$: Event that two cars of each kind
Number of ways of selecting 2 compact cars from 10 and 2 intermediate cars from 20:
$$= {}^{10}C_2 \times {}^{20}C_2$$
Therefore,
$$P(A1) = \dfrac{{}^{10}C_2 \times {}^{20}C_2}{{}^{30}C_4}$$
2. let $A2$: Event that all four cars are of the same kind
Therefore, all the 4 cars must be EITHER compact cars OR all the 4 cars are intermediate size cars.
i.e., 4 compact AND no intermediate size
OR
No compact AND 4 intermediate size
$$\therefore \quad \text{Number of ways} = (10C4 \times 20C0) + (10C0 \times 20C4)$$
$$\therefore \quad P(A) = \dfrac{10C4 + 20C4}{30C4}$$
3. Majority are compact cars
to determine the probability that majority of the selection is compact cars then the possible ways are:
(i) 3 compact AND 1 intermediate size car.
OR
(ii) 4 compact AND 0 intermediate size car.
Therefore, the number of ways is
$$(10C_3 \times 20C_1) + (10C_4 \times 20C_0)$$
[Using product rule and sum rule].
Therefore required probability
$$= \dfrac{(10C3 \times 20C1) + (10C4)}{30C4} = \dfrac{2}{21}$$
Some elementary theorems
Theorem 3
$$p(\phi) = 0$$
Theorem 4
$$p(A’) = 1 – p(A) \quad \text{where } A’ \text{ is the compliment of } A.$$
Theorem 5
For any two events $A$ and $B$.
- $P(A) = P(A \cap B) + P(A \cap B’)$
- $P(B) = P(A \cap B) + P(A’ \cap B)$
Theorem 6
Addition rule for probability
If $A$ and $B$ are any two events in $S$ then $P(A \cup B) = P(A) + P(B) – P(A \cap B)$.
Note:
- When $A$ and $B$ are mutually exclusive so that $P(A \cap B) = 0$, Theorem 6 reduces to Axiom 3.
- For 3 events $A_1, A_2, A_3$, Theorem 6 can be expressed as
$$P(A_1 \cup A_2 \cup A_3) = P(A_1) + P(A_2) + P(A_3) – [P(A_1 \cap A_2) + P(A_2 \cap A_3) + P(A_3 \cap A_1)] + P(A_1 \cap A_2 \cap A_3)$$ - Theorem 6 can be generalised for $n$ events $A_1, A_2, \cdots, A_n$ as follows.
$$P\left(\bigcup_{i=1}^{n} A_i\right) = \sum_{i=1}^{n} P(A_i) – \sum_{i<j} P(A_i \cap A_j) + \sum_{i<j<k} P(A_i \cap A_j \cap A_k) – \cdots (-1)^{n-1} P(A_1 \cap A_2 \cap \cdots \cap A_n)$$
Remark
Above Theorems or similar results do not directly help us to calculate probabilities, but after identifying the basic events, connect them with the correct logic then this theorems or similar results may be applied appropriately. The subsequent examples illustrate the preceeding argument. Apart from this methods any other approach to the examples if any, is appreciable.
Example 6
Given that $P(A) = 0.45$, $P(B) = 0.63$, $P(A \cap B) = 0.24$.
Find
- $P(A \cup B)$,
- $P(A’ \cap B)$,
- $P(A \cap B’)$ and
- $P(A’ \cup B’)$
i. $P(A \cup B) = P(A) + P(B) – P(A \cap B)$ (By Theorem 6)
$$= 0.45 + 0.63 – 0.24 = 0.84$$
ii. $P(A’ \cap B) = P(B) – P(A \cap B)$ (By Theorem 5)
$$= 0.63 – 0.24 = 0.39$$
iii. $P(A \cap B’) = P(A) – P(A \cap B)$ (By Theorem 5)
$$= 0.45 – 0.24 = 0.21$$
iv. $P(A’ \cup B’) = P[(A \cap B)’]$ (By De Morgan’s Law)
$$= 1 – P(A \cap B) \text{ (By Theorem 4)}$$
$$= 1 – 0.24 = 0.76.$$
Example 7
The probability that a person buys a car is 0.16, the probability that he will buy a bike is 0.24 and the probability that he will buy both is 0.11.
(a) What is the probability that he will buy atleast one of the two items?
(b) What is the probability that he will buy only one of two items?
Let $A$ and $B$ be the events that the person will buy a car and he will buy a bike respectively.
$$\therefore \quad \text{Given} \quad P(A) = 0.16;\ P(B) = 0.24$$
$$P(A \cap B) = 0.11$$
(a) $P(\text{he will buy atleast one}) = P(A \cup B)$
$$= P(A) + P(B) – P(A \cap B) \ (\text{Using Theorem 5})$$
$$= 0.16 + 0.24 – 0.11 = 0.29$$
(b) $P(\text{he will buy only one}) = P(\text{ ‘car but not bike’ or ‘bike but not car’})$
$$= P[(A \cap B’) \cup (A’ \cap B)] = P(A \cap B’) + P(A’ \cap B)$$
$$= [P(A) – P(A \cap B)] + [P(B) – P(A \cap B)]$$
$$= (0.16 – 0.11) + (0.24 – 0.11) = 0.18$$
We shall formally define two kind of events which may help us to solve the problem by classifying the basic events and applying the results or theorems, if any appropriately.
Mutually exclusive events
Two events $A$ and $B$ are said to be mutually exclusive if the occurance of $A$ excludes the occurance of $B$ and vice-versa.
i.e., If $A$ happens then $B$ will not happen and vice-versa.
Independent events
Two events $A$ and $B$ are said to be independent events if the occurance or non-occurance of $A$ will not affect the occurance or non-occurance of $B$ and vice-versa.
Example 8
Discuss the winning chance of two horses if they run in (a) two different races and (b) in the same race.
Let $A_1$ be the event that the first horse wins the race and $A_2$ be the event that the second horse wins the race.
(a) If they run in two different races, then whether the first horse wins or loses, it will not affect the winning or losing possibility of the second horse and vice-versa.
i.e., $A_1$ may or may not happen, which does not affect the possibilities of $A_2$.
Therefore $A_1$ and $A_2$ are independent.
(b) If they run in the same race, if the first horse wins the race then the second horse cannot win and vice-versa. $\therefore A_1$ excludes the possibility of $A_2$ and vice versa.
Therefore $A_1$ and $A_2$ are mutually exclusive events.
Let $E$: Selecting an integer from 1 to 50.
$S$: $\{1, 2, 3, 4, \cdots, 50\} \qquad n(S) = 50$
$A$: Selecting a multiple of $5 = \{5, 10, \cdots, 50\}$ so that $n(A) = 10$
$B$: Selecting a multiple of $7 = \{7, 14, \cdots, 40\}$ so that $n(B) = 7$
$$\therefore \quad p(A) = \dfrac{10}{50}, \quad p(B) = \dfrac{7}{50} \tag{1}$$
Now we are interested in selecting a multiple of 5 among the multiples of 7. This requires the following two steps (i) first select the multiples of 7 and (ii) search the multiple of 5 in that set.
Therefore Step (i) is the set ‘$B$’ and Step (ii) yields only one element (What is it?)
i.e., The favourable element of the required event is only one among the elements of ‘$B$’ and $n(B) = 7$
$$\therefore \quad \text{Required probability} = \dfrac{1}{7} \tag{2}$$
Similarly if we try to calculate the probability of selecting a multiple of 7 among the multiples of 5, then it is
$$= \dfrac{1}{10} \tag{3}$$
The probabilities in (1) are calculated based on the original sample space ‘$S$’, where as the probability in (2) is calculated based on ‘$B$’ and that of in (3) is calculated based on ‘$A$’. Such probabilities (in (2) and (3)) are defined as “Probabilities on Reduced Sample Space”. This idea leads the following definition.
CONDITIONAL PROBABILITY
If $A$ and $B$ are any events in $S$ and $P(B) \ne 0$, the conditional probability of $A$ given $B$ is
$$P(A|B) = \dfrac{P(A \cap B)}{P(B)}$$
Note:
- When we use the symbol $P(A|B)$, we mean the probability of $A$ given some space $B$, (i.e.,) the event $B$ has already happened.
- In a similar way we can define $P(B|A)$ as $P(B|A) = \dfrac{P(A \cap B)}{P(A)}$ which calculates the probability of $B$ when $A$ has already happened and $P(A) \ne 0$.
- Indeed every probability is a conditional probability. But the choice of the original sample space $S$ is evident and the choice of $S$ is clearly understood, we use the simplified notation $P(A)$ instead of $P(A|S)$.
- In this context, the probabilities calculated in (2) and (3) in the above example are $P(A|B)$ and $P(B|A)$ respectively.
The following theorem is an immediate consequence of the definition of conditional probability.
Theorem 7
Multiplication rule of probability
If $A$ and $B$ are events in ‘$S$’ then
$$P(A \cap B) = P(A) \cdot P(B|A) \quad \text{if } P(A) \ne 0$$
$$P(B) \cdot P(A|B) \quad \text{if } P(B) \ne 0$$
Note:
If ‘$A$’ and ‘$B$’ are independent, then $P(A|B) = P(A)$ and $P(B|A) = P(B)$
$$\therefore \quad P(A \cap B) = P(A) \cdot P(B)$$
The following two theorems, are immediate consequence of this result.
Theorem 8
If $A$ and $B$ are independent events then $A$ and $B’$ are independent.
Note:
In a similar way, using $P(B) = P(A \cap B) + P(A’ \cap B)$ we can prove that if $A$ and $B$ are independent then $A’$ and $B$ are independent.
Theorem 9
If $A$ and $B$ are independent then $A’$ and $B’$ are independent.
Note:
- The multiplicative law can be extended to ‘$n$’ events. Let us prove for 3 events $A_1, A_2$ and $A_3$.
$$P(A_1 \cap A_2 \cap A_3) = P(A_1) \cdot P(A_2|A_1) \cdot P(A_3|A_1 \cap A_2).$$
Let $A_1 \cap A_2 = A$
$$\therefore \quad P(A_1 \cap A_2 \cap A_3) = P(A \cap A_3)$$
$$= P(A) \cdot P(A_3|A) \ (\text{Theorem 7})$$
$$= P(A_1 \cap A_2) \cdot P(A_3|A_1 \cap A_2)$$
$$= P(A_1) \cdot P(A_2|A_1) \cdot P(A_3|A_1 \cap A_2) \ (\text{Theorem 7})$$
Similarly we state the Theorem for ‘$n$’ events
$$P(A_1 \cap A_2 \cap \cdots \cap A_n) = P(A_1) \cdot P(A_2|A_1) \cdot P(A_3|A_1 \cap A_2) \cdots P(A_n|A_1 \cap \cdots \cap A_n)$$
2. The multiplication law of ‘$n$’ independent events $A_1, A_2, \cdots, A_n$ can also be extended as
$$P(A_1 \cap A_2 \cap \cdots \cap A_n) = P(A_1) \cdot P(A_2) \cdots P(A_n)$$
The above discussion can be illustrated by the following example.
Example 9
Consider drawing four cards from an ordinary 52-card pack.
Let $A_1$: Drawing an ace on the first draw
$A_2$: Drawing an ace on the second draw
$A_3$: Drawing an ace on the third draw
$A_4$: Drawing an ace on the fourth draw
Case (i): Cards are drawn assuming each is replaced after the draw.
$$\therefore \quad \text{Required probability} = P(A_1 \cap A_2 \cap A_3 \cap A_4)$$
$$= P(A_1) \cdot P(A_2) \cdot P(A_3) \cdot P(A_4) \ (\text{assuming independent events})$$
$$= \dfrac{4}{52} \cdot \dfrac{4}{52} \cdot \dfrac{4}{52} \cdot \dfrac{4}{52}$$
$$= \left(\dfrac{4}{52}\right)^4 = 3.5 \times 10^{-5}$$
Case (ii): Cards are not replaced after each draw.
$$P(A_1 \cap A_2 \cap A_3 \cap A_4) = P(A_1) \cdot P(A_2|A_1) \cdot P(A_3|A_1 \cap A_2) \cdot P(A_4|A_1 \cap A_2 \cap A_3) \tag{1}$$
(1) represents that the result of $2^{nd}$ draw is ‘influenced’ by $1^{st}$ draw and $3^{rd}$ draw is ‘influenced’ by both $1^{st}$ draw and $2^{nd}$ draw and so on.
$$\therefore \quad \text{From (1), Required probability} = \dfrac{4}{52} \times \dfrac{3}{51} \times \dfrac{3}{50} \times \dfrac{1}{49}$$
$$= 3.69 \times 10^{-6}$$
Comparing the answers of Case (i) of Case (ii) of the above Example, We can observe that there is a better chance of drawing four aces when cards are replaced than when kept. This is an intuitively satisfying result since replacing the ace drawn raises chances for an ace on the succeeding draw.
Multiple Events
When more than two events are involved independence by pairs is not sufficient to establish the events as statistically independent.
In the case of 3 events $A_1, A_2, A_3$ they are said to be mutually independent if they satisfy
$$P(A_1 \cap A_2) = P(A_1) \times P(A_2)$$
$$P(A_2 \cap A_3) = P(A_2) \times P(A_3)$$
$$P(A_3 \cap A_1) = P(A_3) \times P(A_1)$$
$$P(A_1 \cap A_2 \cap A_3) = P(A_1) \times P(A_2) \times P(A_3)$$
Similarly we may write 11 conditions for four events $A_1, A_2, A_3$ and $A_4$ to be mutually independent events. (From 4 events, select 2 events in 6 ways, select 3 events in 4 ways and select 4 events in 1 way) Proceeding the argument we require $2^n – n – 1$ conditions for $n$ events $A_1, A_2$ and $A_n$ to be statistically independent.
Following examples may illustrate the way in which the preceeding discussions can be applied in solving the problems.
Example 10
A problem in statistics is given to three students whose chances of solving it are $\dfrac{1}{2}$, $\dfrac{1}{3}$ and $\dfrac{1}{4}$. What is the probability that the problem will be solved?
Let $A$, $B$ and $C$ be three events that the problem is solved by the 3 students.
Given $P(A) = \dfrac{1}{2}$; $P(B) = \dfrac{1}{3}$; $P(C) = \dfrac{1}{4}$
$P(\text{Problem will be solved}) = P(\text{Atleast one of them must solve})$
$$= P(A \cup B \cup C)$$
$$= P(A) + P(B) + P(C) – [P(A \cap B) + P(B \cap C) + P(C \cap A)] + P(A \cap B \cap C)$$
$$= P(A) + P(B) + P(C) – P(A)\cdot P(B) – P(B)\cdot P(C) – P(C)\cdot P(A) + P(A)\cdot P(B)\cdot P(C)$$
(We assume the 3 students try independently)
$$= \dfrac{1}{2} + \dfrac{1}{3} + \dfrac{1}{4} – \dfrac{1}{2}\cdot\dfrac{1}{3} – \dfrac{1}{3}\cdot\dfrac{1}{4} – \dfrac{1}{4}\cdot\dfrac{1}{2} + \dfrac{1}{2}\cdot\dfrac{1}{3}\cdot\dfrac{1}{4} = \dfrac{3}{4}$$
$\textbf{Aliter}$: $P(\text{Problem will be solved}) = P(\text{Atleast one of them will solve})$
$$= 1 – P(\text{None of them solve})$$
$$= 1 – P(A’ \cap B’ \cap C’)$$
$$= 1 – P(A’) \cdot P(B’) \cdot P(C’) \ (\text{Theorem 9})$$
$$= 1 – [(1 – P(A))(1 – P(B))(1 – P(C))] \ (\text{Theorem 4})$$
$$= 1 – \left[\dfrac{1}{2}\cdot\dfrac{2}{3}\cdot\dfrac{3}{4}\right] = \dfrac{3}{4}$$
Example 11
A town has two doctors $X$ and $Y$ operating independently. If the probability that doctor $X$ is available is 0.9 and that for $Y$ is 0.8. What is the probability that atleast one doctor is available when needed?
Let $A$ and $B$ be the events of availability of doctor $X$ and $Y$ respectively.
Given $P(A) = 0.9$, $P(B) = 0.8$. Also given that $A$ and $B$ are independent.
$P(\text{Atleast one Dr. is available}) = P(\text{Either Dr. ‘}X\text{‘ or }Y\text{ or Both})$
$$= P(A \cup B)$$
$$= P(A) + P(B) – P(A \cap B) \ (\text{Theorem 1.6})$$
$$= P(A) + P(B) – P(A) \cdot P(B) \ (\because A \text{ and } B \text{ are independent})$$
$$= (0.9) + (0.8) – (0.9) \cdot (0.8) = 0.98.$$
Note:
If the probability of event $A$ is $p$, then the ‘odds’ that it will happen are given by the ratio of $p$ to $1 – p$. Odds are usually given as a ratio of two positive integers having no common factor, and if an event is more likely not to occur than to occur, then we usually give the odds that it will not occur.
If the odds for the occurrence of an event $A$ are $a$ to $b$ ($a, b$ are positive integers) then $p = \dfrac{a}{a+b}$ and hence $1 – p = \dfrac{b}{a+b}$. This idea is illustrated in the examples.
Example 12
In an Engineering college there are 4 branches namely EEE, ECE, CSE and IT. If 2% of the students of EEE, 3% of the students of ECE, 4% of the students of CSE and 5% of the students of IT have a chance to fail in a paper, what is the probability that a randomly selected student of that college will fail in that paper?
Event $A$: A student fails in that paper.
$E_1$: student of EEE
$E_2$: student of ECE
$E_3$: student of CSE
$E_4$: student of IT
(observe that $E_1, E_2, E_3$ and $E_4$ are mutually exclusive).
Define $A$ as follows:
(A failed student of EEE) or (failed student of ECE) or (failed student of CSE) or (failed student of IT)
$$\therefore \quad A = (A \cap E_1) \cup (A \cap E_2) \cup (A \cap E_3) \cup (A \cap E_4) \tag{1}$$
$$\therefore \quad P(A) = P(A \cap E_1) + P(A \cap E_2) + P(A \cap E_3) + P(A \cap E_4)$$
$$= P(E_1) \cdot P(A|E_1) + P(E_2) \cdot P(A|E_2) + P(E_3) \cdot P(A|E_3) + P(E_4) \cdot P(A|E_4)$$
$$= (0.25)(0.02) + (0.25)(0.03) + (0.25)(0.04) + (0.25)(0.05)$$
(Assuming equal chance for $E_1, E_2, E_3, E_4$)
$$= 0.035.$$
In the above example if we are interested to know the probability of randomly selected student who has failed in that paper is a student of “EEE or ECE or CSE or IT” then we define such events as $P(E_i|A)$ in Example 12 $(i = 1, 2, 3, 4)$.
Now we use the definition of conditional probability for the above events
$$P(E_i|A) = \dfrac{P(E_i \cap A)}{P(A)} = \dfrac{P(E_i) \cdot P(A|E_i)}{P(A)}$$
$P(A)$ will be replaced by the equation of Total Probability [as in Examples (1.26)].
The proceeding discussion leads a theorem called Bayes’ theorem. Before providing the actual Theorem and the proof of Bayes’ Theorem we can generalise the concepts of Total Probability.
Suppose we are given ‘$n$’ mutually exclusive events $E_1, E_2, \cdots, E_n$ whose union equals ‘$S$’ the sample space. These events satisfy
(i) $E_i \cap E_j = \phi$ when $i \ne j$
(ii) $\displaystyle\bigcup_{i=1}^{n} E_i = S.$
Then, (generalising the ideas in 1.26)
$$A = (A \cap E_1) \cup (A \cap E_2) \cup \cdots \cup (A \cap E_n)$$
$$P(A) = P(A \cap E_1) + P(A \cap E_2) + \cdots + P(A \cap E_n)$$
$$P(A) = P(E_1)P(A|E_1) + P(E_2) \cdot P(A|E_2) + P(E_3)P(A|E_3) + \cdots + P(A_n)P(A|E_n)$$
which is known as the Total Probability of event $A$. Now using these ideas we formally state and prove the Bayes’ Theorem. While solving such problems we should observe whether the given problem has the hypothesis of Bayes’ Theorem such as ‘$n$’ mutually exclusive events and a “common event $A$” and the required probabilities
$$P(E_1), P(E_2), \cdots, P(E_n) \quad \text{and} \quad P(A|E_1), P(A|E_2), \cdots, P(A|E_n)$$
Many of us follow the probabilistic approach perhaps intuitively at times. (We apply our experience for such estimation.) Bayes’ theorem is to provide a foundation to the methods of sampling and experimentation. However the significance of the results depends on the data from which the probabilities are estimated.
BAYES’ THEOREM
Let $E_1, E_2, \cdots E_n$ be $n$ mutually exclusive events. Let $A$ be a subset of union of all $E_i$’s. The probabilities namely $P(E_1), P(E_2), \cdots P(E_n)$; $P(A|E_1), P(A|E_2), \cdots, P(A|E_n)$ are known. Then,
$$P(E_i|A) = \dfrac{P(E_i)\,P(A|E_i)}{\displaystyle\sum_{j=1}^{n} P(E_j)P(A|E_j)} \quad (\text{for each } i)$$
Note:
- In Bayes’ theorem equation (3) is said to be rule of Elimination or the rule of Total Probability which we have discussed already.
- Bayes’ theorem provides a formula for finding the probability that the “Effect” $A$ was “Caused” by the event $E_i$. The probabilities $P(E_i)$ are called the prior or ‘a apriori’ probabilities of the causes $E_i$. The probabilities $P(A|E_i)$ are sometimes called transition probabilities and $P(E_i|A)$ (which we calculate using Bayes’ formula) are called Posterior probabilities since they apply after the experiment’s performance when some event $A$ is obtained.
- If we extend the idea of the event ‘A’ in Bayes’ theorem to more such events, then we can calculate the more posterior probabilities using Bayes’ theorem.
Following Examples illustrates this idea.
Example 13
Three bags are given each containing red and white balls as indicated.
Bag 1: 6 red and 4 white; Bag 2: 2 red and 6 white Bag 3: 1 red and 8 white.
- A bag is chosen at random a ball is drawn from this urn. The ball is red. Find the probability that the bag chosen was bag 1.
- A bag is chosen at random and two balls are drawn without replacement from this bag. If both balls are red, find the probability that bag 1 was chosen. Under these conditions, what is the probability that bag 3 was chosen.
Let $E_1, E_2, E_3$ be 3 events that represent the selection of the bags 1, 2 and 3 respectively.
Since a bag is selected at random.
$$P(E_1) = P(E_2) = P(E_3) \quad \text{and} \quad P(E_1) + P(E_2) + P(E_3) = 1$$
$$\therefore \quad P(E_1) = \dfrac{1}{3};\ P(E_2) = \dfrac{1}{3};\ P(E_3) = \dfrac{1}{3}$$
(1) Let $A$ be the event of selecting red ball.
$$P(A|E_1) = P(\text{red ball from bag 1})$$
$$= \dfrac{6}{10} = \dfrac{3}{5}$$
Similarly $$P(A|E_2) = \dfrac{2}{8} = \dfrac{1}{4}$$
$$P(A|E_3) = \dfrac{1}{9}$$
Using Total Probability
$$P(A) = P(E_1) \cdot P(A|E_1) + P(E_2) \cdot P(A|E_2) + P(E_3) \cdot P(A|E_3)$$
$$= \dfrac{1}{3}\cdot\dfrac{3}{5} + \dfrac{1}{3}\cdot\dfrac{1}{4} + \dfrac{1}{3}\cdot\dfrac{1}{9} = \dfrac{173}{540}$$
Required probability $= P(E_1|A)$
$$= \dfrac{P(E_1 \cap A)}{P(A)} \quad \text{(Definition of conditional probability)}$$
$$= \dfrac{P(E_1) \cdot P(A|E_1)}{P(A)}$$
$$= \dfrac{1/3 \cdot 3/5}{173/540} = \dfrac{108}{173}$$
(2) Let $A$ be the event of selecting 2 red balls. Without replacement
$$\therefore \quad P(A|E_1) = \dfrac{6C_2}{10C_2} = \dfrac{15}{45} = \dfrac{1}{3}$$
$$P(A|E_2) = \dfrac{2C_2}{8C_2} = \dfrac{1}{28} \tag{1}$$
$$P(A|E_3) = 0 \ (\because \text{There is only one red ball in bag 3})$$
$$\therefore \quad P(A) = P(E_1) \cdot P(A|E_1) + P(E_2) \cdot P(A|E_2) + P(E_3) \cdot P(A|E_3) \quad (\text{Total probability})$$
$$= \dfrac{1}{3}\cdot\dfrac{1}{3} + \dfrac{1}{3}\cdot\dfrac{1}{28} = \dfrac{31}{252}$$
Required probability $= P(E_1|A) = \dfrac{P(E_1) \cdot P(A|E_1)}{P(A)}$
$$= \dfrac{1/9}{31/252} = \dfrac{252}{279} = \dfrac{28}{31}$$
Similarly, $P(E_3|A) = 0$
$$= \dfrac{6}{10} \times \dfrac{5}{9} \ (\text{since we are not replacing})$$
$$= \dfrac{6C_2}{10C_2}$$
$\therefore$ Computationally, the combinatorics rules can be applied when the samples are without replacement from a lot.
Note:
In the above example the probabilities in should be computed as $P(A|E_1) = P(2 \text{ red balls without replacement in bag 1})$.
Example 14
Engineers incharge of maintaining nuclear fleet must continually check for corrosion inside the pipes that are part of the cooling systems. The inside condition of the pipes cannot be observed directly but a non-destructive test can give an indication of possible corrosion. This test is not infallible. The test has probability 0.7 of detecting corrosion when it is present but it also has probability 0.2 of falsely indicating internal corrosion. If the probability that any section of pipe has internal corrosion is 0.1.
(a) Determine the probability that a section of pipe has internal corrosion given that the test indicates its presence.
(b) Determine the probability that a section of pipe has corrosion given that the test is negative.
Let $E_1$ be the event that the section of pipe has internal corrosion.
$$\therefore \quad P(E_1) = 0.1$$
Let $E_2 = E_1′ =$ the pipe has no internal corrosion
$$P(E_2) = 1 – 0.1$$
$$P(E_2) = 0.9$$
Let $A$ be the event that the test indicates a corrosion.
$$\therefore \quad P(A|E_1) = 0.7$$
$$P(A|E_2) = 0.2$$
(a) $$P(E_1|A) = \dfrac{P(E_1) \cdot P(A|E_1)}{p(E_1) \cdot p(A|E_1) + P(E_2)(A|E_2)} \quad \text{Bayes formula}$$
$$= \dfrac{(0.1)(0.7)}{(0.1)(0.7) + (0.9)(0.2)} = \dfrac{0.07}{0.25} = 0.28$$
(b) $$P(E_1|A’) = \dfrac{P(E_1 \cap A’)}{P(A’)} \quad \text{(Definition of conditional probability)}$$
$$= \dfrac{P(E_1) – P(E_1 \cap A)}{P(A’)} \quad \text{(By Theorem 5)}$$
$$= \dfrac{P(E_1) – P(E_1) \cdot P(A|E_1)}{1 – P(A)}$$
$$= \dfrac{(0.1) – (0.1)(0.7)}{1 – 0.25} = \dfrac{0.03}{0.75} = 0.04$$