Descriptive Stats (NumPy)
np.mean(x) # average
np.median(x) # 50th percentile
np.var(x, ddof=1) # sample variance (n-1)
np.std(x, ddof=1) # sample std dev
np.percentile(x,[25,50,75]) # quartiles
stats.mode(x) # most frequent
stats.iqr(x) # interquartile range
stats.zscore(x) # standardize
Formulas
mean = sum(x_i) / n
var = sum((x_i - mean)^2) / (n - 1)
std = sqrt(var)
SE = std / sqrt(n) # standard error
z = (x - mean) / std
Probability Rules
P(not A) = 1 - P(A)
P(A or B) = P(A) + P(B) - P(A and B)
P(A | B) = P(A and B) / P(B)
P(A and B) = P(A) * P(B) # if independent
Bayes: P(A|B) = P(B|A) P(A) / P(B)
E[X] = sum(x * P(x)) # expected value
Distributions
| Dist | Mean | Var |
|---|---|---|
| Bernoulli(p) | p | p(1-p) |
| Binomial(n,p) | np | np(1-p) |
| Poisson(λ) | λ | λ |
| Normal(μ,σ²) | μ | σ² |
| Exponential(λ) | 1/λ | 1/λ² |
Binomial PMF: C(n,k) p^k (1-p)^(n-k)
Poisson PMF: e^(-lambda) lambda^k / k!
Normal PDF: (1/(s*sqrt(2pi))) exp(-(x-m)^2/(2 s^2))
scipy.stats API
from scipy import stats
stats.binom.pmf(k, n, p)
stats.poisson.pmf(k, mu)
stats.norm.pdf(x, loc, scale)
stats.norm.cdf(1.96) # 0.975
stats.norm.ppf(0.975) # 1.96 (inverse CDF)
stats.expon.cdf(x, scale=1/lmbda)
dist.rvs(size=1000) # sample random draws
dist.mean(), dist.std()
Hypothesis Tests
stats.ttest_1samp(x, popmean) # vs known mean
stats.ttest_ind(a, b, equal_var=False) # two groups
stats.ttest_rel(before, after) # paired
stats.chi2_contingency(table) # categorical
stats.f_oneway(g1, g2, g3) # ANOVA
stats.shapiro(x) # normality
stats.mannwhitneyu(a, b) # non-parametric
# reject H0 if p < alpha (0.05)
| Error | Meaning |
|---|---|
| Type I (α) | False positive: reject true H0 |
| Type II (β) | False negative: miss real effect |
| Power | 1 - β |
Confidence Intervals
# 95% CI for mean
CI = mean +/- z * (sigma / sqrt(n)) # z=1.96 for 95%
from scipy import stats
se = stats.sem(data)
stats.t.interval(0.95, df=n-1, loc=data.mean(), scale=se)
stats.norm.interval(0.95, loc=mean, scale=se)
Interpretation
95% of repeated CIs contain the true parameter, not "95% probability this interval does." Use t for small n / unknown sigma.
Correlation & A/B
stats.pearsonr(x, y) # linear, r in [-1,1]
stats.spearmanr(x, y) # monotonic rank
np.corrcoef(x, y) # correlation matrix
# A/B two-proportion z-test
from statsmodels.stats.proportion import proportions_ztest
z, p = proportions_ztest([120,150], [2000,2000])
# sample size / power
from statsmodels.stats.power import TTestIndPower
TTestIndPower().solve_power(effect_size=0.2,
alpha=0.05, power=0.8)