The Two Workhorses
NumPy gives you fast, typed n-dimensional arrays and the vectorized math that powers almost every data library. pandas builds labeled tables on top of NumPy so you can slice, group, and join real-world data by name. Together they are the foundation of the Python data stack.
import numpy as np
import pandas as pd
Creating ndarrays
An ndarray is a grid of values of the same type, described by its shape and dtype.
a = np.array([1, 2, 3]) # from a list
m = np.array([[1, 2], [3, 4]]) # 2D
np.zeros((2, 3)) # all zeros
np.ones((2, 3)) # all ones
np.full((2, 2), 7) # constant
np.arange(0, 10, 2) # [0 2 4 6 8]
np.linspace(0, 1, 5) # 5 evenly spaced
np.eye(3) # identity matrix
np.random.default_rng(0).random(3) # modern RNG
m.shape # (2, 2)
m.ndim # 2
m.size # 4
m.dtype # dtype('int64')
Dtypes
Every array has one dtype. Choosing the right one controls memory and precision. Mixing types promotes to the most general (int + float becomes float).
np.array([1, 2], dtype=np.float64)
np.array([1, 2], dtype=np.int32)
arr.astype(np.float32) # convert dtype
np.array([1, 2.0]) # -> float64 (promotion)
np.array([1, 2], dtype=bool) # -> [ True True]
Vectorization vs Loops
The core idea of NumPy: apply an operation to a whole array at once. This is both faster and clearer than a Python loop.
a = np.arange(1_000_000)
# slow: pure-Python loop
result = [x * 2 + 1 for x in a]
# fast: vectorized, runs in C
result = a * 2 + 1
# elementwise ufuncs
np.sqrt(a); np.exp(a); np.log1p(a)
np.where(a > 5, "hi", "lo") # vectorized ternary
| Feature | Python list | NumPy array |
|---|---|---|
| Elementwise math | Loop / comprehension | a * 2 |
| Speed | Slow | 10-100x faster |
| Memory | Boxed objects | Compact buffer |
| Multi-dimensional | Nested lists | Native n-D |
Broadcasting
Broadcasting lets NumPy combine arrays of different shapes by stretching the smaller one along dimensions of size 1, without copying data.
a = np.array([[1, 2, 3],
[4, 5, 6]]) # shape (2, 3)
a + 10 # scalar broadcast to every element
a + np.array([100, 200, 300]) # row vector -> each row
col = np.array([[10], [20]]) # shape (2, 1)
a + col # column broadcast to each column
Broadcasting rule
Compare shapes from the right. Dimensions are compatible when they are equal or one of them is 1. Otherwise NumPy raises a shape error.
Indexing & Slicing
Slices return views (not copies), so writing to a slice changes the original. Boolean and fancy indexing return copies.
a = np.arange(10)
a[2:5] # [2 3 4]
a[::-1] # reversed
a[a > 5] # boolean mask -> [6 7 8 9]
a[[0, 2, 4]] # fancy indexing
m = np.arange(12).reshape(3, 4)
m[1, 2] # single element
m[:, 1] # whole column
m[0:2, 1:3] # sub-block
m[m % 2 == 0] # all even values
Aggregation
Reductions collapse an array to summary values. The axis argument controls the direction: axis=0 down columns, axis=1 across rows.
m = np.array([[1, 2, 3],
[4, 5, 6]])
m.sum() # 21 (all)
m.sum(axis=0) # [5 7 9] (per column)
m.sum(axis=1) # [6 15] (per row)
m.mean(); m.std(); m.min(); m.max()
m.argmax(); m.cumsum()
pandas Series & DataFrame
A Series is a labeled 1D array; a DataFrame is a labeled 2D table (a dict of Series sharing an index).
s = pd.Series([10, 20, 30], index=["a", "b", "c"])
s["b"] # 20
df = pd.DataFrame({
"name": ["Ada", "Linus", "Grace"],
"team": ["A", "B", "A"],
"score": [91, 88, 95],
})
df.head(); df.info(); df.describe()
df.shape; df.columns; df.dtypes
df["score"] # a column (Series)
df[["name", "score"]] # multiple columns
Selection: loc vs iloc
Use .loc for label-based selection and .iloc for integer-position selection. Mixing them up is the most common pandas mistake.
| Aspect | .loc | .iloc |
|---|---|---|
| Selects by | Labels | Integer positions |
| Slice end | Inclusive | Exclusive |
| Boolean mask | Yes | No |
df.loc[0, "name"] # by label
df.loc[df["score"] > 90] # boolean filter
df.loc[0:2, ["name", "score"]] # label slice (inclusive)
df.iloc[0] # first row
df.iloc[0:2, 0:2] # first 2 rows/cols (exclusive)
df.iloc[-1] # last row
Avoid chained assignment
Write df.loc[mask, "col"] = value instead of df[mask]["col"] = value. The chained form may write to a temporary copy and silently do nothing.
GroupBy
The split-apply-combine pattern: group rows by a key, apply an aggregation, and combine the results into a new table.
df.groupby("team")["score"].mean()
# multiple aggregations at once
df.groupby("team").agg(
avg_score=("score", "mean"),
n=("score", "size"),
)
# group by more than one column
df.groupby(["team", "name"]).sum()
Merge & Join
Combine tables on shared keys, just like a SQL join. The how argument controls which rows are kept.
teams = pd.DataFrame({"team": ["A", "B"], "lead": ["Kai", "Mo"]})
pd.merge(df, teams, on="team", how="left")
pd.merge(df, teams, on="team", how="inner")
# stack tables (union of rows)
pd.concat([df1, df2], ignore_index=True)
Missing Data
Real data has gaps. pandas represents them as NaN. Detect, drop, or fill them explicitly.
df.isna().sum() # count nulls per column
df.dropna() # drop rows with any NaN
df.dropna(subset=["score"]) # only if key column null
df["score"].fillna(df["score"].mean())
df.ffill(); df.bfill() # forward / backward fill
Reshaping: Pivot & Melt
Switch between wide and long layouts. pivot_table makes data wider (spreads values into columns); melt makes it longer (collapses columns into rows).
# long -> wide
wide = df.pivot_table(
index="team", columns="name",
values="score", aggfunc="mean",
)
# wide -> long
long = wide.reset_index().melt(
id_vars="team",
var_name="name", value_name="score",
)
Tidy data
Most analysis and plotting is easiest in long/tidy form: one observation per row, one variable per column. Reshape to wide only for final presentation.
Practice Exercises
- Create a 4x4 array of the numbers 0-15, then extract the 2x2 block in its bottom-right corner.
- Given an array of temperatures, use a boolean mask to replace every value below 0 with 0.
- Use broadcasting to subtract each column's mean from a 2D array (mean-centering).
- Load a CSV into a DataFrame and report the mean score per team using
groupby. - Merge a sales DataFrame with a products lookup on
product_idusing a left join. - Fill missing values in a numeric column with its median, then pivot the table to a wide layout.