Probability Distributions in R

The d, p, q, r Framework


Introduction

Every probability distribution in R is accessed through four functions sharing the same prefix letter, followed by the distribution name (e.g. norm, binom). The four prefixes always mean the same thing, regardless of which distribution you are working with.


d — Density or Probability Mass

For a continuous distribution, d returns the value of the probability density function (PDF) at a given point x. This is the height of the distribution curve at x. It is not a probability. Probability for a continuous variable only arises as area under the curve over an interval:

$$P(a \leq X \leq b) = \int_a^b f(x)\,dx$$

The density value can be any non-negative number, including values greater than 1 — perfectly valid for a density.

For a discrete distribution, d returns the probability mass function (PMF) — the actual probability of observing exactly the value x:

$$P(X = x) = f(x)$$

Here the output is a probability, bounded between 0 and 1.

Key optional argument:

  • log = FALSE (default): returns the density or mass on its natural scale.
  • log = TRUE: returns the natural log of the density or mass, i.e. log f(x).
dnorm(0, mean=0, sd=1)             # ≈  0.3989  — density on natural scale
dnorm(0, mean=0, sd=1, log=TRUE)   # ≈ -0.9189  — log-density, i.e. log(0.3989)

Why use log = TRUE? Densities of many observations multiplied together (likelihoods) can underflow to zero in floating-point arithmetic. Working in log scale turns products into sums, which is numerically stable. This is essential in maximum likelihood estimation and Bayesian computation.


p — Cumulative Distribution Function (CDF)

p returns the cumulative probability for a given value q. By default this is the lower tail:

$$P(X \leq q)$$

This is always a true probability between 0 and 1. It is the workhorse for computing probabilities over ranges:

$$P(a \leq X \leq b) = p(b) – p(a)$$

Key optional arguments:

lower.tail

  • lower.tail = TRUE (default): returns P(X ≤ q) — probability to the left.
  • lower.tail = FALSE: returns P(X > q) — probability to the right.
pnorm(1.96)                          # P(X ≤ 1.96) ≈ 0.9750  — left tail
pnorm(1.96, lower.tail = FALSE)      # P(X > 1.96) ≈ 0.0250  — right tail

These two always sum to 1 (for continuous distributions). For discrete distributions they overlap by exactly P(X = q), so:

pbinom(4, 10, 0.3) + pbinom(4, 10, 0.3, lower.tail=FALSE)  # = 1 + dbinom(4,10,0.3)

Why use lower.tail = FALSE? Beyond convenience, it is more numerically precise for very small tail probabilities. When p is very close to 1, computing 1 - pnorm(x) loses precision because of floating-point cancellation. Passing lower.tail = FALSE directly avoids this subtraction and gives a more accurate small probability.

# These are mathematically equal but the second is more numerically accurate:
1 - pnorm(8)                         # may suffer precision loss
pnorm(8, lower.tail = FALSE)         # ≈ 6.22e-16  — more reliable

log.p

  • log.p = FALSE (default): probabilities on natural scale.
  • log.p = TRUE: probability is given or returned as its natural log.
pnorm(1.96, log.p = TRUE)            # ≈ -0.02535  — i.e. log(0.9750)
pnorm(8,    log.p = TRUE)            # ≈ -6.22e-16 — extreme tail, stable in log scale

log.p can be combined with lower.tail:

pnorm(8, lower.tail=FALSE, log.p=TRUE)   # log of the upper-tail probability

q — Quantile Function (Inverse CDF)

q is the inverse of p. Given a cumulative probability p, it returns the value x such that:

$$P(X \leq x) = p$$

This is the quantile function or inverse CDF. Essential for critical values, confidence interval thresholds, and percentile cutoffs. For discrete distributions, q snaps to the smallest integer satisfying the cumulative condition.

Since p and q are inverses:

qnorm(pnorm(1.96))    # = 1.96
pnorm(qnorm(0.975))   # = 0.975

Key optional arguments — mirror those of p:

  • lower.tail = TRUE (default): p is treated as P(X ≤ x), so returns the lower quantile.
  • lower.tail = FALSE: p is treated as an upper-tail probability P(X > x).
  • log.p = FALSE (default): p supplied on natural scale.
  • log.p = TRUE: p supplied as log(p) — useful when the probability came from a p call with log.p = TRUE.
qnorm(0.025)                             # ≈ -1.96  — 2.5th percentile (lower tail)
qnorm(0.025, lower.tail = FALSE)         # ≈  1.96  — upper 2.5% cutoff

# Passing a log-probability directly:
qnorm(log(0.025), log.p = TRUE)          # ≈ -1.96  — same result, log-scale input

r — Random Sampling

r draws n independent random samples from the distribution. The first argument is always n. There are no lower.tail or log.p arguments here — those concepts don’t apply to sampling.

rnorm(5, mean=0, sd=1)    # e.g. c(-0.3, 1.2, 0.8, -1.1, 0.4)

Universal Argument Summary

ArgumentApplies toDefaultEffect
logd onlyFALSEReturn log of density/mass
lower.tailp, qTRUETRUE = left tail P(X≤x); FALSE = right tail P(X>x)
log.pp, qFALSEProbabilities on log scale (input for q, output for p)

These three arguments behave identically across all distributions — binom, norm, pois, gamma, beta, and every other distribution in R.


Continuous vs Discrete: The Key Distinction

ContinuousDiscrete
d returnsPDF height — not a probabilityPMF value — a probability
d can exceed 1?YesNo
P(X = x) exactlyAlways 0Equal to d
Probability over rangeArea under curve via pDifference of p values

Examples by Distribution


Binomial — binom (Discrete)

Models the number of successes in size independent trials, each with success probability prob.

dbinom(x, size, prob, log = FALSE)
pbinom(q, size, prob, lower.tail = TRUE, log.p = FALSE)
qbinom(p, size, prob, lower.tail = TRUE, log.p = FALSE)
rbinom(n, size, prob)

X ~ Binomial(size = 10, prob = 0.3):

dbinom(4, size=10, prob=0.3)                      # P(X=4) ≈ 0.2001
pbinom(4, size=10, prob=0.3)                      # P(X ≤ 4) ≈ 0.8497
pbinom(4, size=10, prob=0.3, lower.tail=FALSE)    # P(X > 4) ≈ 0.1503
pbinom(6, size=10, prob=0.3) - pbinom(3, size=10, prob=0.3)  # P(4 ≤ X ≤ 6)
qbinom(0.90, size=10, prob=0.3)                   # 90th percentile = 5
rbinom(5, size=10, prob=0.3)

Normal — norm (Continuous)

Parameterised by mean (μ) and sd (σ — standard deviation, not variance).

dnorm(x, mean, sd, log = FALSE)
pnorm(q, mean, sd, lower.tail = TRUE, log.p = FALSE)
qnorm(p, mean, sd, lower.tail = TRUE, log.p = FALSE)
rnorm(n, mean, sd)

X ~ Normal(mean = 100, sd = 15):

dnorm(100, mean=100, sd=15)                        # PDF height ≈ 0.02659 — not a probability
dnorm(100, mean=100, sd=15, log=TRUE)              # ≈ -3.627 — log-density

pnorm(115, mean=100, sd=15)                        # P(X ≤ 115) ≈ 0.8413
pnorm(85,  mean=100, sd=15, lower.tail=FALSE)      # P(X > 85)  ≈ 0.8413  — same by symmetry
pnorm(115, mean=100, sd=15) - pnorm(85, mean=100, sd=15)  # P(85 ≤ X ≤ 115) ≈ 0.6827

qnorm(0.975, mean=100, sd=15)                      # 97.5th percentile ≈ 129.4
qnorm(0.025, mean=100, sd=15, lower.tail=FALSE)    # same: upper 2.5% cutoff ≈ 129.4

rnorm(5, mean=100, sd=15)

Note: R takes sd, not var. Pass sd = 15 for variance 225.


Poisson — pois (Discrete)

Models event counts in a fixed interval at average rate lambda (λ).

dpois(x, lambda, log = FALSE)
ppois(q, lambda, lower.tail = TRUE, log.p = FALSE)
qpois(p, lambda, lower.tail = TRUE, log.p = FALSE)
rpois(n, lambda)

X ~ Poisson(lambda = 4):

dpois(4, lambda=4)                           # P(X=4) ≈ 0.1954
ppois(6, lambda=4)                           # P(X ≤ 6) ≈ 0.8893
ppois(6, lambda=4, lower.tail=FALSE)         # P(X > 6) ≈ 0.1107
ppois(6, lambda=4) - ppois(2, lambda=4)      # P(3 ≤ X ≤ 6)
qpois(0.90, lambda=4)                        # 90th percentile = 7
rpois(5, lambda=4)

Gamma — gamma (Continuous)

Right-skewed distribution on (0, ∞). Use either rate or scale, never both.

⚠️ rate = 1/scale. R defaults to rate = 1. Textbooks often use scale (θ) — be careful.

dgamma(x, shape, rate=1, log = FALSE)          # or scale= instead of rate=
pgamma(q, shape, rate=1, lower.tail=TRUE, log.p=FALSE)
qgamma(p, shape, rate=1, lower.tail=TRUE, log.p=FALSE)
rgamma(n, shape, rate=1)

X ~ Gamma(shape = 3, rate = 0.5) ≡ Gamma(shape = 3, scale = 2):

dgamma(4, shape=3, rate=0.5)                   # PDF height ≈ 0.0916
pgamma(4, shape=3, rate=0.5)                   # P(X ≤ 4) ≈ 0.3233
pgamma(4, shape=3, rate=0.5, lower.tail=FALSE) # P(X > 4) ≈ 0.6767
qgamma(0.95, shape=3, rate=0.5)               # 95th percentile ≈ 12.24
rgamma(5, shape=3, rate=0.5)

dgamma(4, shape=3, scale=2)                    # identical — scale = 1/rate

Beta — beta (Continuous)

Defined on [0, 1]. Only shape parameters — no location or scale.

dbeta(x, shape1, shape2, log = FALSE)
pbeta(q, shape1, shape2, lower.tail=TRUE, log.p=FALSE)
qbeta(p, shape1, shape2, lower.tail=TRUE, log.p=FALSE)
rbeta(n, shape1, shape2)

X ~ Beta(shape1 = 2, shape2 = 5):

dbeta(0.3, shape1=2, shape2=5)                  # PDF height ≈ 1.852 — can exceed 1!
pbeta(0.3, shape1=2, shape2=5)                  # P(X ≤ 0.3) ≈ 0.4202
pbeta(0.3, shape1=2, shape2=5, lower.tail=FALSE) # P(X > 0.3) ≈ 0.5798
pbeta(0.5, shape1=2, shape2=5) - pbeta(0.2, shape1=2, shape2=5)  # P(0.2 ≤ X ≤ 0.5)
qbeta(0.50, shape1=2, shape2=5)                 # median ≈ 0.2793
rbeta(5, shape1=2, shape2=5)

dbeta(0.3, 2, 5) ≈ 1.85 is a density height — not a probability, and not required to be ≤ 1.


Summary

DistributionTypeKey Argumentsd returns
binomDiscretesize, proba probability
normContinuousmean, sdPDF height (not a probability)
poisDiscretelambdaa probability
gammaContinuousshape, rate or scalePDF height (not a probability)
betaContinuousshape1, shape2PDF height (can exceed 1)

p and q are always inverses of each other:

pnorm(qnorm(0.95))                               = 0.95
qnorm(pnorm(1.645))                              = 1.645
ppois(qpois(0.80, lambda=3), lambda=3)           ≥ 0.80

For discrete distributions, q returns the smallest integer where the cumulative probability first meets or exceeds p.

ppois(qpois(0.80, lambda=3), lambda=3) gives the cumulative probability at that integer, which is ≥ 0.80 but not necessarily equal to it.

Key Takeaways

R provides four functions for every distribution

Here’s the bulleted summary:
d — density height (continuous) or exact probability (discrete); use log=TRUE for numerical stability

p — cumulative probability P(X ≤ q); the only way to get probability over a range for continuous distributions

q — inverse CDF: given a probability, returns the corresponding value; exact for continuous, approximate for discrete

r — random samples; no tail or log arguments

lower.tail — switches p and q between left tail (default) and right tail; also more numerically precise for extreme probabilities

log.p — probabilities on log scale; output for p, input for q

Continuous d — a height, not a probability; can exceed 1 (e.g. Beta)

Discrete d — a probability; p and q round-trip is approximate due to step-function CDF

    Scroll to Top