contentintech
Learn/data science/Statistics & Probability
Intermediate~20 min read

Statistics & Probability

The statistical foundations for data science: descriptive stats, distributions, the CLT, Bayes' theorem, hypothesis testing, confidence intervals, and A/B testing.

statisticsprobabilityhypothesis-testingdistributions

Why Statistics for Data Science?

Statistics is the language of uncertainty. Every model estimate, every A/B test result, and every metric comes with noise. Understanding distributions, inference, and hypothesis testing lets you separate real signal from randomness.

Descriptive Statistics

Descriptive stats summarize a dataset. Central tendency (mean, median, mode) and spread (variance, standard deviation, range) are the core measures.

The mean is mean = sum(x_i) / n. Population variance is var = sum((x_i - mean)^2) / n; the sample version divides by n - 1 (Bessel's correction). Standard deviation is std = sqrt(var).

import numpy as np

x = np.array([4, 8, 15, 16, 23, 42])
np.mean(x)             # 18.0
np.median(x)           # 15.5
np.var(x, ddof=1)      # sample variance (n-1)
np.std(x, ddof=1)      # sample std dev
np.percentile(x, [25, 50, 75])   # quartiles

Mean vs Median

The median resists outliers; the mean does not. For skewed data (income, response times), report the median. Always set ddof=1 in NumPy for sample statistics.

Probability Basics

Probability measures the chance of an event, from 0 (impossible) to 1 (certain). Key rules:

Complement:  P(not A) = 1 - P(A)
Union:       P(A or B) = P(A) + P(B) - P(A and B)
Conditional: P(A | B)  = P(A and B) / P(B)
Independent: P(A and B) = P(A) * P(B)   if A, B independent

A random variable maps outcomes to numbers. Its distribution is described by a PMF (discrete) or PDF (continuous), and its expected value is the long-run average.

Common Distributions

DistributionTypeMeanVarianceModels
Bernoulli(p)Discretepp(1-p)Single yes/no trial
Binomial(n,p)Discretenpnp(1-p)Successes in n trials
Poisson(λ)DiscreteλλCounts per interval
Normal(μ,σ²)Continuousμσ²Bell-curve quantities
Exponential(λ)Continuous1/λ1/λ²Time between events

Key formulas: Binomial PMF P(X=k) = C(n,k) p^k (1-p)^(n-k); Poisson PMF P(X=k) = e^(-lambda) lambda^k / k!; Normal PDF f(x) = (1/(sigma sqrt(2 pi))) exp(-(x-mu)^2 / (2 sigma^2)).

from scipy import stats

stats.binom.pmf(k=3, n=10, p=0.5)      # P(X=3)
stats.poisson.pmf(k=2, mu=4)           # lambda=4
stats.norm.cdf(1.96, loc=0, scale=1)   # P(Z <= 1.96) approx 0.975
stats.norm.ppf(0.975)                  # inverse CDF -> 1.96
stats.expon.mean(scale=1/2)            # lambda=2 -> mean 0.5

The Central Limit Theorem

The CLT states that the distribution of the sample mean approaches a Normal distribution as sample size grows, regardless of the population's shape. If the population has mean mu and std sigma, the sample mean has mean mu and standard error SE = sigma / sqrt(n).

import numpy as np
pop = np.random.exponential(scale=2, size=100000)  # skewed
means = [np.random.choice(pop, 50).mean() for _ in range(5000)]
# Histogram of `means` is approximately Normal even though
# the population is exponential.

Why It Matters

The CLT is what makes confidence intervals and t-tests valid on means even when the raw data is not Normal, as long as n is reasonably large (a common rule of thumb is 30 or more).

Bayes' Theorem

Bayes' theorem updates a belief given new evidence: P(A|B) = P(B|A) P(A) / P(B), where P(A) is the prior, P(B|A) the likelihood, and P(A|B) the posterior.

Worked Example: Medical Test

A disease affects 1% of people. A test is 99% sensitive (detects the disease when present) and has a 5% false-positive rate. If you test positive, what is the probability you actually have the disease?

P(D)      = 0.01          # prior: has disease
P(pos|D)  = 0.99          # sensitivity
P(pos|~D) = 0.05          # false positive rate

P(pos) = 0.99*0.01 + 0.05*0.99 = 0.0099 + 0.0495 = 0.0594
P(D|pos) = (0.99 * 0.01) / 0.0594 = 0.1667

Despite a positive result, there is only a 16.7% chance of disease. The low prior (base rate) dominates. This base-rate fallacy trips up intuition constantly.

Hypothesis Testing

We test a null hypothesis (H0, "no effect") against an alternative (H1). The p-value is the probability of observing data at least as extreme as ours if H0 were true. If p < alpha (commonly 0.05), we reject H0.

DecisionH0 TrueH0 False
Reject H0Type I error (α)Correct (power)
Fail to rejectCorrectType II error (β)

A Type I error is a false positive (crying wolf); a Type II error is a false negative (missing a real effect). Power = 1 - beta.

t-test & Chi-Square

from scipy import stats

# Two-sample t-test: are two group means different?
t, p = stats.ttest_ind(group_a, group_b, equal_var=False)
if p < 0.05:
    print("Reject H0: means differ")

# One-sample t-test vs a known value
stats.ttest_1samp(sample, popmean=100)

# Chi-square test of independence (categorical)
chi2, p, dof, expected = stats.chi2_contingency(contingency_table)

Use a t-test to compare means, and a chi-square test to check whether two categorical variables are associated.

Confidence Intervals

A confidence interval (CI) gives a plausible range for a parameter. A 95% CI for the mean is mean +/- z * (sigma / sqrt(n)), where z = 1.96 for 95%. Use the t-distribution when sigma is unknown and n is small.

import numpy as np
from scipy import stats

data = np.array([...])
n = len(data)
mean, se = data.mean(), stats.sem(data)
ci = stats.t.interval(0.95, df=n-1, loc=mean, scale=se)
print(ci)   # (lower, upper)

Correct Interpretation

A 95% CI means: if we repeated the experiment many times, 95% of such intervals would contain the true parameter. It does NOT mean there is a 95% probability the true value lies in this one specific interval.

Correlation vs Causation

Correlation (e.g. Pearson's r in [-1, 1]) measures linear association. It does not imply causation: a hidden confounder can drive both variables. Establishing causation requires a randomized experiment or careful causal inference.

r, p = stats.pearsonr(x, y)      # linear correlation
rho, p = stats.spearmanr(x, y)   # rank correlation (monotonic)

A/B Testing

A/B testing is a randomized controlled experiment: split users into control (A) and treatment (B), then test whether the metric difference is statistically significant. For conversion rates (proportions), compare with a two-proportion z-test or a chi-square test.

from statsmodels.stats.proportion import proportions_ztest

# A: 120/2000 converted, B: 150/2000 converted
count = [120, 150]
nobs  = [2000, 2000]
z, p = proportions_ztest(count, nobs)
print(f"z={z:.2f}, p={p:.4f}")   # reject H0 if p < 0.05

Before launching, run a power analysis to compute the sample size needed to detect the minimum effect you care about. Peeking early and stopping when significant inflates the Type I error rate.

Practice Exercises

  1. Given a skewed dataset, compute the mean, median, sample variance, and IQR. Explain why the mean and median differ.
  2. Simulate 5000 sample means from a Poisson population and plot the histogram. Confirm it looks Normal (the CLT in action).
  3. Rework the medical-test Bayes example with a disease prevalence of 10%. How does the posterior P(D|pos) change?
  4. Use stats.ttest_ind on two samples and state your null hypothesis, p-value, and conclusion at alpha = 0.05.
  5. Compute a 95% confidence interval for a sample mean and write a one-sentence correct interpretation.
  6. Run a two-proportion z-test on A/B conversion data (110/1500 vs 140/1500). Is the lift significant?
  7. Explain, with an example, one situation where two variables are strongly correlated but not causally related.

Section navigation