Why Python for Data?
Python is the lingua franca of data science because it pairs a readable, beginner-friendly syntax with a mature ecosystem of libraries like NumPy, pandas, and scikit-learn. Before reaching for those libraries, you need a firm grip on the language fundamentals. This guide covers the core Python skills you will use every single day when wrangling data.
Modern Python
Examples target Python 3.12+. Use f-strings for formatting, type hints for clarity, and pathlib over raw string paths where possible.
Core Data Types
Python's built-in types map naturally onto the shapes of data you will encounter. Knowing which one to reach for keeps code fast and correct.
| Type | Example | Mutable? | Typical use |
|---|---|---|---|
int / float | 42, 3.14 | No | Counts, measurements |
str | "revenue" | No | Labels, text |
list | [1, 2, 3] | Yes | Ordered sequences |
tuple | (x, y) | No | Fixed records, keys |
dict | {"a": 1} | Yes | Key-value lookups |
set | Yes | Uniqueness, membership |
prices = [19.99, 5.49, 12.00] # list: ordered, mutable
record = ("SKU-42", 19.99, 3) # tuple: fixed record
product = {"sku": "SKU-42", "qty": 3} # dict: labeled fields
seen = {"SKU-42", "SKU-99"} # set: unique membership
total = sum(prices)
print(f"Total: ${total:.2f}") # f-string -> Total: $37.48
List & Dict Comprehensions
Comprehensions are the idiomatic way to transform and filter collections. They are shorter and usually faster than an equivalent for loop with append.
nums = [1, 2, 3, 4, 5, 6]
squares = [n * n for n in nums] # [1, 4, 9, 16, 25, 36]
evens = [n for n in nums if n % 2 == 0] # [2, 4, 6]
# dict comprehension: sku -> price
skus = ["A", "B", "C"]
prices = [9.99, 4.50, 12.00]
catalog = {sku: price for sku, price in zip(skus, prices)}
# set comprehension: unique first letters
initials = {word[0] for word in ["apple", "avocado", "banana"]} # {'a', 'b'}
Keep it readable
If a comprehension needs more than one condition or a nested loop, prefer an explicit loop. Clever one-liners that nobody can read are a false economy.
Functions
Functions package reusable logic. Use type hints to document intent and default arguments for optional parameters. Avoid mutable default arguments.
def normalize(values: list[float], target_max: float = 1.0) -> list[float]:
"""Scale values so the largest equals target_max."""
hi = max(values)
return [v / hi * target_max for v in values]
print(normalize([2, 4, 8])) # [0.25, 0.5, 1.0]
# *args / **kwargs for flexible signatures
def summarize(*columns: str, sep: str = ", ") -> str:
return sep.join(columns)
print(summarize("name", "age", sep=" | ")) # name | age
Gotcha: mutable defaults
Never write def f(x, cache=[]). The list is created once and shared across calls. Use cache=None then if cache is None: cache = [].
File I/O
Always open files with a with block so they close automatically, even if an error occurs. Modern code prefers pathlib.Path for filesystem paths.
from pathlib import Path
path = Path("data") / "notes.txt"
# write
path.write_text("line one\nline two\n", encoding="utf-8")
# read all lines, stripping the newline
with path.open(encoding="utf-8") as f:
for line in f:
print(line.rstrip())
# quick one-liner read
text = path.read_text(encoding="utf-8")
Working with CSV
CSV is the most common flat data format. The standard library csv module handles quoting and delimiters correctly, so avoid splitting on commas manually.
import csv
from pathlib import Path
rows = [
{"name": "Ada", "score": 91},
{"name": "Linus", "score": 88},
]
# write with a header
with Path("scores.csv").open("w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["name", "score"])
writer.writeheader()
writer.writerows(rows)
# read back into dicts
with Path("scores.csv").open(newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
print(row["name"], int(row["score"]))
Working with JSON
JSON is the standard for API responses and config. The json module converts between Python objects and JSON text.
import json
from pathlib import Path
config = {"model": "v2", "threshold": 0.75, "tags": ["a", "b"]}
# object -> JSON string / file
Path("config.json").write_text(json.dumps(config, indent=2))
# JSON file -> object
loaded = json.loads(Path("config.json").read_text())
print(loaded["threshold"]) # 0.75
# parse an API-style string
payload = '{"ok": true, "count": 3}'
data = json.loads(payload)
print(data["count"]) # 3
Generators & Lazy Evaluation
Generators produce values one at a time instead of building a full list in memory. They are essential when streaming large files or infinite sequences.
def read_large(path):
"""Yield rows lazily; never loads the whole file."""
with open(path, encoding="utf-8") as f:
for line in f:
yield line.rstrip()
# generator expression: like a comprehension but lazy
total = sum(len(row) for row in read_large("big.csv"))
# generators are single-use and evaluated on demand
gen = (n * n for n in range(3))
print(next(gen)) # 0
print(list(gen)) # [1, 4]
List vs generator
Use a list when you need to index, reuse, or know the length. Use a generator when the data is large or you only need to iterate once.
Why NumPy? A Motivation
Python lists are flexible but slow for numerical math because each element is a full Python object. NumPy stores numbers in a compact typed buffer and runs operations in optimized C, which is often 10-100x faster.
| Aspect | Python list | NumPy array |
|---|---|---|
| Element type | Mixed, boxed objects | Single fixed dtype |
| Math on all elements | Manual loop | Vectorized |
| Speed | Slow for big data | Fast (C-backed) |
| Memory | High overhead | Compact |
# list: element-by-element
prices = [10, 20, 30]
with_tax = [p * 1.08 for p in prices]
# numpy: one vectorized operation
import numpy as np
arr = np.array([10, 20, 30])
with_tax = arr * 1.08 # array([10.8, 21.6, 32.4])
A First Glimpse of pandas
pandas builds on NumPy to give you labeled tables (DataFrame). It reads CSV/JSON in one line and makes filtering and aggregation trivial. You will go deep on this in the NumPy & Pandas guide.
import pandas as pd
df = pd.read_csv("scores.csv")
print(df.head())
# filter and aggregate in one expression
top = df[df["score"] > 90]
print(df["score"].mean())
Virtual Environments
A virtual environment isolates a project's packages so different projects do not clash. Create one per project and never install libraries globally.
# create and activate (macOS / Linux)
python -m venv .venv
source .venv/bin/activate
# Windows
.venv\Scripts\activate
# install and freeze dependencies
pip install numpy pandas
pip freeze > requirements.txt
# recreate later
pip install -r requirements.txt
Modern tooling
Tools like uv and poetry manage environments and lockfiles automatically. uv venv and uv pip install are dramatically faster than plain pip.
Practice Exercises
- Given a list of temperatures in Celsius, use a comprehension to build a new list of Fahrenheit values, keeping only those above freezing.
- Write a function
word_count(text)that returns a dict mapping each word to how many times it appears. - Read a CSV of products, and write a new CSV containing only rows where
price > 100. - Load a JSON config file, add a new key, and write it back with 2-space indentation.
- Write a generator that yields the running total of a sequence of numbers.
- Create a virtual environment, install pandas, and freeze the dependencies to
requirements.txt.