contentintech
Learn/data science/NumPy & Pandas
Beginner~20 min read

NumPy & Pandas

The essential NumPy and pandas toolkit: ndarrays, broadcasting, vectorization, Series/DataFrame, loc/iloc, groupby, merges, missing data, and reshaping.

numpypandasdataframevectorization

The Two Workhorses

NumPy gives you fast, typed n-dimensional arrays and the vectorized math that powers almost every data library. pandas builds labeled tables on top of NumPy so you can slice, group, and join real-world data by name. Together they are the foundation of the Python data stack.

python
import numpy as np
import pandas as pd

Creating ndarrays

An ndarray is a grid of values of the same type, described by its shape and dtype.

python
a = np.array([1, 2, 3])              # from a list
m = np.array([[1, 2], [3, 4]])       # 2D

np.zeros((2, 3))                     # all zeros
np.ones((2, 3))                      # all ones
np.full((2, 2), 7)                   # constant
np.arange(0, 10, 2)                  # [0 2 4 6 8]
np.linspace(0, 1, 5)                 # 5 evenly spaced
np.eye(3)                            # identity matrix
np.random.default_rng(0).random(3)   # modern RNG

m.shape      # (2, 2)
m.ndim       # 2
m.size       # 4
m.dtype      # dtype('int64')

Dtypes

Every array has one dtype. Choosing the right one controls memory and precision. Mixing types promotes to the most general (int + float becomes float).

python
np.array([1, 2], dtype=np.float64)
np.array([1, 2], dtype=np.int32)
arr.astype(np.float32)               # convert dtype

np.array([1, 2.0])                   # -> float64 (promotion)
np.array([1, 2], dtype=bool)         # -> [ True  True]

Vectorization vs Loops

The core idea of NumPy: apply an operation to a whole array at once. This is both faster and clearer than a Python loop.

python
a = np.arange(1_000_000)

# slow: pure-Python loop
result = [x * 2 + 1 for x in a]

# fast: vectorized, runs in C
result = a * 2 + 1

# elementwise ufuncs
np.sqrt(a); np.exp(a); np.log1p(a)
np.where(a > 5, "hi", "lo")          # vectorized ternary
FeaturePython listNumPy array
Elementwise mathLoop / comprehensiona * 2
SpeedSlow10-100x faster
MemoryBoxed objectsCompact buffer
Multi-dimensionalNested listsNative n-D

Broadcasting

Broadcasting lets NumPy combine arrays of different shapes by stretching the smaller one along dimensions of size 1, without copying data.

python
a = np.array([[1, 2, 3],
              [4, 5, 6]])           # shape (2, 3)

a + 10                              # scalar broadcast to every element
a + np.array([100, 200, 300])       # row vector -> each row

col = np.array([[10], [20]])        # shape (2, 1)
a + col                             # column broadcast to each column

Broadcasting rule

Compare shapes from the right. Dimensions are compatible when they are equal or one of them is 1. Otherwise NumPy raises a shape error.

Indexing & Slicing

Slices return views (not copies), so writing to a slice changes the original. Boolean and fancy indexing return copies.

python
a = np.arange(10)
a[2:5]                # [2 3 4]
a[::-1]               # reversed
a[a > 5]              # boolean mask -> [6 7 8 9]
a[[0, 2, 4]]          # fancy indexing

m = np.arange(12).reshape(3, 4)
m[1, 2]               # single element
m[:, 1]               # whole column
m[0:2, 1:3]           # sub-block
m[m % 2 == 0]         # all even values

Aggregation

Reductions collapse an array to summary values. The axis argument controls the direction: axis=0 down columns, axis=1 across rows.

python
m = np.array([[1, 2, 3],
              [4, 5, 6]])

m.sum()               # 21  (all)
m.sum(axis=0)         # [5 7 9]  (per column)
m.sum(axis=1)         # [6 15]   (per row)
m.mean(); m.std(); m.min(); m.max()
m.argmax(); m.cumsum()

pandas Series & DataFrame

A Series is a labeled 1D array; a DataFrame is a labeled 2D table (a dict of Series sharing an index).

python
s = pd.Series([10, 20, 30], index=["a", "b", "c"])
s["b"]                # 20

df = pd.DataFrame({
    "name": ["Ada", "Linus", "Grace"],
    "team": ["A", "B", "A"],
    "score": [91, 88, 95],
})

df.head(); df.info(); df.describe()
df.shape; df.columns; df.dtypes
df["score"]           # a column (Series)
df[["name", "score"]] # multiple columns

Selection: loc vs iloc

Use .loc for label-based selection and .iloc for integer-position selection. Mixing them up is the most common pandas mistake.

Aspect.loc.iloc
Selects byLabelsInteger positions
Slice endInclusiveExclusive
Boolean maskYesNo
python
df.loc[0, "name"]              # by label
df.loc[df["score"] > 90]       # boolean filter
df.loc[0:2, ["name", "score"]] # label slice (inclusive)

df.iloc[0]                     # first row
df.iloc[0:2, 0:2]              # first 2 rows/cols (exclusive)
df.iloc[-1]                    # last row

Avoid chained assignment

Write df.loc[mask, "col"] = value instead of df[mask]["col"] = value. The chained form may write to a temporary copy and silently do nothing.

GroupBy

The split-apply-combine pattern: group rows by a key, apply an aggregation, and combine the results into a new table.

python
df.groupby("team")["score"].mean()

# multiple aggregations at once
df.groupby("team").agg(
    avg_score=("score", "mean"),
    n=("score", "size"),
)

# group by more than one column
df.groupby(["team", "name"]).sum()

Merge & Join

Combine tables on shared keys, just like a SQL join. The how argument controls which rows are kept.

python
teams = pd.DataFrame({"team": ["A", "B"], "lead": ["Kai", "Mo"]})

pd.merge(df, teams, on="team", how="left")
pd.merge(df, teams, on="team", how="inner")

# stack tables (union of rows)
pd.concat([df1, df2], ignore_index=True)

Missing Data

Real data has gaps. pandas represents them as NaN. Detect, drop, or fill them explicitly.

python
df.isna().sum()               # count nulls per column
df.dropna()                   # drop rows with any NaN
df.dropna(subset=["score"])   # only if key column null
df["score"].fillna(df["score"].mean())
df.ffill(); df.bfill()        # forward / backward fill

Reshaping: Pivot & Melt

Switch between wide and long layouts. pivot_table makes data wider (spreads values into columns); melt makes it longer (collapses columns into rows).

python
# long -> wide
wide = df.pivot_table(
    index="team", columns="name",
    values="score", aggfunc="mean",
)

# wide -> long
long = wide.reset_index().melt(
    id_vars="team",
    var_name="name", value_name="score",
)

Tidy data

Most analysis and plotting is easiest in long/tidy form: one observation per row, one variable per column. Reshape to wide only for final presentation.

Practice Exercises

  1. Create a 4x4 array of the numbers 0-15, then extract the 2x2 block in its bottom-right corner.
  2. Given an array of temperatures, use a boolean mask to replace every value below 0 with 0.
  3. Use broadcasting to subtract each column's mean from a 2D array (mean-centering).
  4. Load a CSV into a DataFrame and report the mean score per team using groupby.
  5. Merge a sales DataFrame with a products lookup on product_id using a left join.
  6. Fill missing values in a numeric column with its median, then pivot the table to a wide layout.

Section navigation