The d, p, q, r Framework
Introduction
Every probability distribution in R is accessed through four functions sharing the same prefix letter, followed by the distribution name (e.g. norm, binom). The four prefixes always mean the same thing, regardless of which distribution you are working with.
d — Density or Probability Mass
For a continuous distribution, d returns the value of the probability density function (PDF) at a given point x. This is the height of the distribution curve at x. It is not a probability. Probability for a continuous variable only arises as area under the curve over an interval:
$$P(a \leq X \leq b) = \int_a^b f(x)\,dx$$
The density value can be any non-negative number, including values greater than 1 — perfectly valid for a density.
For a discrete distribution, d returns the probability mass function (PMF) — the actual probability of observing exactly the value x:
$$P(X = x) = f(x)$$
Here the output is a probability, bounded between 0 and 1.
Key optional argument:
log = FALSE(default): returns the density or mass on its natural scale.log = TRUE: returns the natural log of the density or mass, i.e. log f(x).
dnorm(0, mean=0, sd=1) # ≈ 0.3989 — density on natural scale
dnorm(0, mean=0, sd=1, log=TRUE) # ≈ -0.9189 — log-density, i.e. log(0.3989)
Why use
log = TRUE? Densities of many observations multiplied together (likelihoods) can underflow to zero in floating-point arithmetic. Working in log scale turns products into sums, which is numerically stable. This is essential in maximum likelihood estimation and Bayesian computation.
p — Cumulative Distribution Function (CDF)
p returns the cumulative probability for a given value q. By default this is the lower tail:
$$P(X \leq q)$$
This is always a true probability between 0 and 1. It is the workhorse for computing probabilities over ranges:
$$P(a \leq X \leq b) = p(b) – p(a)$$
Key optional arguments:
lower.tail
lower.tail = TRUE(default): returns P(X ≤ q) — probability to the left.lower.tail = FALSE: returns P(X > q) — probability to the right.
pnorm(1.96) # P(X ≤ 1.96) ≈ 0.9750 — left tail
pnorm(1.96, lower.tail = FALSE) # P(X > 1.96) ≈ 0.0250 — right tail
These two always sum to 1 (for continuous distributions). For discrete distributions they overlap by exactly P(X = q), so:
pbinom(4, 10, 0.3) + pbinom(4, 10, 0.3, lower.tail=FALSE) # = 1 + dbinom(4,10,0.3)
Why use
lower.tail = FALSE? Beyond convenience, it is more numerically precise for very small tail probabilities. When p is very close to 1, computing1 - pnorm(x)loses precision because of floating-point cancellation. Passinglower.tail = FALSEdirectly avoids this subtraction and gives a more accurate small probability.
# These are mathematically equal but the second is more numerically accurate:
1 - pnorm(8) # may suffer precision loss
pnorm(8, lower.tail = FALSE) # ≈ 6.22e-16 — more reliable
log.p
log.p = FALSE(default): probabilities on natural scale.log.p = TRUE: probability is given or returned as its natural log.
pnorm(1.96, log.p = TRUE) # ≈ -0.02535 — i.e. log(0.9750)
pnorm(8, log.p = TRUE) # ≈ -6.22e-16 — extreme tail, stable in log scale
log.pcan be combined withlower.tail:
pnorm(8, lower.tail=FALSE, log.p=TRUE) # log of the upper-tail probability
q — Quantile Function (Inverse CDF)
q is the inverse of p. Given a cumulative probability p, it returns the value x such that:
$$P(X \leq x) = p$$
This is the quantile function or inverse CDF. Essential for critical values, confidence interval thresholds, and percentile cutoffs. For discrete distributions, q snaps to the smallest integer satisfying the cumulative condition.
Since p and q are inverses:
qnorm(pnorm(1.96)) # = 1.96
pnorm(qnorm(0.975)) # = 0.975
Key optional arguments — mirror those of p:
lower.tail = TRUE(default): p is treated as P(X ≤ x), so returns the lower quantile.lower.tail = FALSE: p is treated as an upper-tail probability P(X > x).log.p = FALSE(default): p supplied on natural scale.log.p = TRUE: p supplied as log(p) — useful when the probability came from apcall withlog.p = TRUE.
qnorm(0.025) # ≈ -1.96 — 2.5th percentile (lower tail)
qnorm(0.025, lower.tail = FALSE) # ≈ 1.96 — upper 2.5% cutoff
# Passing a log-probability directly:
qnorm(log(0.025), log.p = TRUE) # ≈ -1.96 — same result, log-scale input
r — Random Sampling
r draws n independent random samples from the distribution. The first argument is always n. There are no lower.tail or log.p arguments here — those concepts don’t apply to sampling.
rnorm(5, mean=0, sd=1) # e.g. c(-0.3, 1.2, 0.8, -1.1, 0.4)
Universal Argument Summary
| Argument | Applies to | Default | Effect |
|---|---|---|---|
log | d only | FALSE | Return log of density/mass |
lower.tail | p, q | TRUE | TRUE = left tail P(X≤x); FALSE = right tail P(X>x) |
log.p | p, q | FALSE | Probabilities on log scale (input for q, output for p) |
These three arguments behave identically across all distributions —
binom,norm,pois,gamma,beta, and every other distribution in R.
Continuous vs Discrete: The Key Distinction
| Continuous | Discrete | |
|---|---|---|
d returns | PDF height — not a probability | PMF value — a probability |
d can exceed 1? | Yes | No |
| P(X = x) exactly | Always 0 | Equal to d |
| Probability over range | Area under curve via p | Difference of p values |
Examples by Distribution
Binomial — binom (Discrete)
Models the number of successes in size independent trials, each with success probability prob.
dbinom(x, size, prob, log = FALSE)
pbinom(q, size, prob, lower.tail = TRUE, log.p = FALSE)
qbinom(p, size, prob, lower.tail = TRUE, log.p = FALSE)
rbinom(n, size, prob)
X ~ Binomial(size = 10, prob = 0.3):
dbinom(4, size=10, prob=0.3) # P(X=4) ≈ 0.2001
pbinom(4, size=10, prob=0.3) # P(X ≤ 4) ≈ 0.8497
pbinom(4, size=10, prob=0.3, lower.tail=FALSE) # P(X > 4) ≈ 0.1503
pbinom(6, size=10, prob=0.3) - pbinom(3, size=10, prob=0.3) # P(4 ≤ X ≤ 6)
qbinom(0.90, size=10, prob=0.3) # 90th percentile = 5
rbinom(5, size=10, prob=0.3)
Normal — norm (Continuous)
Parameterised by mean (μ) and sd (σ — standard deviation, not variance).
dnorm(x, mean, sd, log = FALSE)
pnorm(q, mean, sd, lower.tail = TRUE, log.p = FALSE)
qnorm(p, mean, sd, lower.tail = TRUE, log.p = FALSE)
rnorm(n, mean, sd)
X ~ Normal(mean = 100, sd = 15):
dnorm(100, mean=100, sd=15) # PDF height ≈ 0.02659 — not a probability
dnorm(100, mean=100, sd=15, log=TRUE) # ≈ -3.627 — log-density
pnorm(115, mean=100, sd=15) # P(X ≤ 115) ≈ 0.8413
pnorm(85, mean=100, sd=15, lower.tail=FALSE) # P(X > 85) ≈ 0.8413 — same by symmetry
pnorm(115, mean=100, sd=15) - pnorm(85, mean=100, sd=15) # P(85 ≤ X ≤ 115) ≈ 0.6827
qnorm(0.975, mean=100, sd=15) # 97.5th percentile ≈ 129.4
qnorm(0.025, mean=100, sd=15, lower.tail=FALSE) # same: upper 2.5% cutoff ≈ 129.4
rnorm(5, mean=100, sd=15)
Note: R takes
sd, notvar. Passsd = 15for variance 225.
Poisson — pois (Discrete)
Models event counts in a fixed interval at average rate lambda (λ).
dpois(x, lambda, log = FALSE)
ppois(q, lambda, lower.tail = TRUE, log.p = FALSE)
qpois(p, lambda, lower.tail = TRUE, log.p = FALSE)
rpois(n, lambda)
X ~ Poisson(lambda = 4):
dpois(4, lambda=4) # P(X=4) ≈ 0.1954
ppois(6, lambda=4) # P(X ≤ 6) ≈ 0.8893
ppois(6, lambda=4, lower.tail=FALSE) # P(X > 6) ≈ 0.1107
ppois(6, lambda=4) - ppois(2, lambda=4) # P(3 ≤ X ≤ 6)
qpois(0.90, lambda=4) # 90th percentile = 7
rpois(5, lambda=4)
Gamma — gamma (Continuous)
Right-skewed distribution on (0, ∞). Use either rate or scale, never both.
⚠️
rate = 1/scale. R defaults torate = 1. Textbooks often use scale (θ) — be careful.
dgamma(x, shape, rate=1, log = FALSE) # or scale= instead of rate=
pgamma(q, shape, rate=1, lower.tail=TRUE, log.p=FALSE)
qgamma(p, shape, rate=1, lower.tail=TRUE, log.p=FALSE)
rgamma(n, shape, rate=1)
X ~ Gamma(shape = 3, rate = 0.5) ≡ Gamma(shape = 3, scale = 2):
dgamma(4, shape=3, rate=0.5) # PDF height ≈ 0.0916
pgamma(4, shape=3, rate=0.5) # P(X ≤ 4) ≈ 0.3233
pgamma(4, shape=3, rate=0.5, lower.tail=FALSE) # P(X > 4) ≈ 0.6767
qgamma(0.95, shape=3, rate=0.5) # 95th percentile ≈ 12.24
rgamma(5, shape=3, rate=0.5)
dgamma(4, shape=3, scale=2) # identical — scale = 1/rate
Beta — beta (Continuous)
Defined on [0, 1]. Only shape parameters — no location or scale.
dbeta(x, shape1, shape2, log = FALSE)
pbeta(q, shape1, shape2, lower.tail=TRUE, log.p=FALSE)
qbeta(p, shape1, shape2, lower.tail=TRUE, log.p=FALSE)
rbeta(n, shape1, shape2)
X ~ Beta(shape1 = 2, shape2 = 5):
dbeta(0.3, shape1=2, shape2=5) # PDF height ≈ 1.852 — can exceed 1!
pbeta(0.3, shape1=2, shape2=5) # P(X ≤ 0.3) ≈ 0.4202
pbeta(0.3, shape1=2, shape2=5, lower.tail=FALSE) # P(X > 0.3) ≈ 0.5798
pbeta(0.5, shape1=2, shape2=5) - pbeta(0.2, shape1=2, shape2=5) # P(0.2 ≤ X ≤ 0.5)
qbeta(0.50, shape1=2, shape2=5) # median ≈ 0.2793
rbeta(5, shape1=2, shape2=5)
dbeta(0.3, 2, 5) ≈ 1.85is a density height — not a probability, and not required to be ≤ 1.
Summary
| Distribution | Type | Key Arguments | d returns |
|---|---|---|---|
binom | Discrete | size, prob | a probability |
norm | Continuous | mean, sd | PDF height (not a probability) |
pois | Discrete | lambda | a probability |
gamma | Continuous | shape, rate or scale | PDF height (not a probability) |
beta | Continuous | shape1, shape2 | PDF height (can exceed 1) |
p and q are always inverses of each other:
pnorm(qnorm(0.95)) = 0.95
qnorm(pnorm(1.645)) = 1.645
ppois(qpois(0.80, lambda=3), lambda=3) ≥ 0.80
For discrete distributions, q returns the smallest integer where the cumulative probability first meets or exceeds p.
ppois(qpois(0.80, lambda=3), lambda=3) gives the cumulative probability at that integer, which is ≥ 0.80 but not necessarily equal to it.
Key Takeaways
R provides four functions for every distribution
Here’s the bulleted summary: d — density height (continuous) or exact probability (discrete); use log=TRUE for numerical stability
p — cumulative probability P(X ≤ q); the only way to get probability over a range for continuous distributions
q — inverse CDF: given a probability, returns the corresponding value; exact for continuous, approximate for discrete
r — random samples; no tail or log arguments
lower.tail — switches p and q between left tail (default) and right tail; also more numerically precise for extreme probabilities
log.p — probabilities on log scale; output for p, input for q
Continuous d — a height, not a probability; can exceed 1 (e.g. Beta)
Discrete d — a probability; p and q round-trip is approximate due to step-function CDF