Formula sheet

The key formulas of the course, grouped by topic. Search by name, symbol or a word from the note.

42 formulas

Events

Addition rule

P(A \cup B) = P(A) + P(B) - P(A \cap B)

Subtract the intersection so it is not counted twice.

Complement

P(A') = 1 - P(A)

The complement is usually easier — "at least one" is almost always 1 minus "none".

De Morgan's laws

(A \cup B)' = A' \cap B' \qquad (A \cap B)' = A' \cup B'

Disjoint events

P(A \cap B) = 0 \;\Rightarrow\; P(A \cup B) = P(A) + P(B)

Disjoint ≠ independent. Two disjoint events with positive probability are in fact dependent.

Counting

Classical probability

P(A) = \frac{|A|}{|\Omega|}

Only when all outcomes are equally likely.

Combinations — order does not matter

\binom{n}{k} = \frac{n!}{k!\,(n-k)!}

Choosing k out of n when order does not matter.

Permutations — order matters

P(n,k) = \frac{n!}{(n-k)!}

Order matters. k! times more than combinations.

Conditional & Bayes

Conditional probability

P(A \mid B) = \frac{P(A \cap B)}{P(B)}

The denominator is the event that has already happened — it is the new sample space.

Law of total probability

P(B) = \sum_i P(B \mid A_i)\,P(A_i)

When the Aᵢ split the space into a complete, disjoint partition.

Bayes' formula

P(A \mid B) = \frac{P(B \mid A)\,P(A)}{P(B)}

Reverses the direction of the conditioning.

Independence

P(A \cap B) = P(A)P(B) \iff P(A \mid B) = P(A)

This is a definition, not a property you assume.

Discrete RVs

Expectation — discrete

E(X) = \sum_x x\,P(X = x)

A weighted average by the probabilities.

Variance — definition and shortcut

Var(X) = E\left[(X-\mu)^2\right] = E(X^2) - \left[E(X)\right]^2

The second form is almost always easier to compute.

Standard deviation

SD(X) = \sqrt{Var(X)}

In the units of X itself, which is why it is the one you compare against.

Linear transformation

E(aX + b) = aE(X) + b \qquad Var(aX + b) = a^2\,Var(X)

In Var the coefficient is squared and b disappears — a shift does not change the spread.

Discrete distributions

Binomial

P(X=k) = \binom{n}{k} p^k (1-p)^{n-k} \qquad E=np,\; Var=np(1-p)

n independent trials, constant probability p.

Bernoulli

E(X) = p \qquad Var(X) = p(1-p)

A single trial: 0 or 1.

Geometric

P(X=k) = (1-p)^{k-1}p \qquad E=\frac{1}{p},\; Var=\frac{1-p}{p^2}

The number of trials up to and including the first success.

Poisson

P(X=k) = \frac{e^{-\lambda}\lambda^k}{k!} \qquad E = Var = \lambda

Expectation and variance are equal — the classic trap is using λ instead of √λ for the standard deviation.

Hypergeometric

P(X=k) = \frac{\binom{K}{k}\binom{N-K}{n-k}}{\binom{N}{n}}

Sampling without replacement — which is why it is not binomial.

Joint distributions

Covariance

Cov(X,Y) = E(XY) - E(X)E(Y)

Zero if independent — but not the other way round.

Correlation coefficient

\rho = \frac{Cov(X,Y)}{SD(X)\,SD(Y)}

Always between −1 and 1, and unitless.

Variance of a sum — the general case

Var(X+Y) = Var(X) + Var(Y) + 2\,Cov(X,Y)

The last term vanishes only under independence.

Sums & indicators

Expectation of a sum

E\left(\sum_{i=1}^{n} X_i\right) = n\mu

Always holds — even without independence. This is the strong property of expectation.

Variance of a sum

Var\left(\sum_{i=1}^{n} X_i\right) = n\sigma^2

Only when the variables are independent.

Standard deviation of a sum

SD\left(\sum_{i=1}^{n} X_i\right) = \sigma\sqrt{n}

The square root of nσ² — not n·σ.

Indicators

X = \sum_i I_i \;\Rightarrow\; E(X) = \sum_i P(A_i)

The way to compute the expectation of "how many out of" without touching the distribution.

Continuous RVs

Density and CDF

F(x) = \int_{-\infty}^{x} f(t)\,dt \qquad f(x) = F'(x)

Probability is area — so P(X=x)=0 for any single x.

Expectation and variance — continuous

E(X) = \int x f(x)\,dx \qquad Var(X) = \int x^2 f(x)\,dx - \left[E(X)\right]^2

Probability of an interval

P(a < X < b) = F(b) - F(a)

For a continuous variable there is no difference between < and ≤.

Continuous distributions

Continuous uniform

f(x) = \frac{1}{b-a} \qquad E = \frac{a+b}{2},\; Var = \frac{(b-a)^2}{12}

The 12 in the variance's denominator is the thing that gets forgotten.

Exponential

f(x) = \lambda e^{-\lambda x} \qquad E = \frac{1}{\lambda},\; Var = \frac{1}{\lambda^2}

Memoryless: P(X>s+t \mid X>s) = P(X>t).

Standardization

Z = \frac{X - \mu}{\sigma}

If X is normal, the result is distributed N(0,1).

The normal table

P(Z < z) = \Phi(z) \qquad \Phi(-z) = 1 - \Phi(z)

For a negative value use the complement — the table only lists positive values.

CLT

Expectation of the sample mean

E(\bar{X}) = \mu

Averaging does not move the expectation.

Variance of the sample mean

Var(\bar{X}) = \frac{\sigma^2}{n}

Divide by n, not by √n. That is the common mistake here.

Standard error of the sample mean

SD(\bar{X}) = \frac{\sigma}{\sqrt{n}}

Take the square root of the variance of the mean: σ/√n, not σ/n.

Standardizing the sample mean

Z = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}}

The denominator is the standard deviation of what is in the numerator — of the mean, not of a single variable.

Distribution of the sample mean

\bar{X} \sim N\left(\mu,\; \frac{\sigma^2}{n}\right)

Inside the brackets is the variance, not the standard deviation.

Distribution of the sum

\sum_{i=1}^{n} X_i \sim N\left(n\mu,\; n\sigma^2\right)

Markov & Chebyshev

Markov's inequality

P(X \ge a) \le \frac{E(X)}{a}

Only for a non-negative variable, and a very loose bound.

Chebyshev's inequality

P\left(|X - \mu| \ge k\sigma\right) \le \frac{1}{k^2}

You do not need to know the distribution — which is why the bound is crude.