contentintech
Learn/data science/Python for Data
Beginner~20 min read

Python for Data

Core Python skills for data work: types, comprehensions, functions, file/CSV/JSON I/O, generators, virtual environments, and a first glimpse of pandas.

pythondata-typesfile-iogenerators

Why Python for Data?

Python is the lingua franca of data science because it pairs a readable, beginner-friendly syntax with a mature ecosystem of libraries like NumPy, pandas, and scikit-learn. Before reaching for those libraries, you need a firm grip on the language fundamentals. This guide covers the core Python skills you will use every single day when wrangling data.

Modern Python

Examples target Python 3.12+. Use f-strings for formatting, type hints for clarity, and pathlib over raw string paths where possible.

Core Data Types

Python's built-in types map naturally onto the shapes of data you will encounter. Knowing which one to reach for keeps code fast and correct.

TypeExampleMutable?Typical use
int / float42, 3.14NoCounts, measurements
str"revenue"NoLabels, text
list[1, 2, 3]YesOrdered sequences
tuple(x, y)NoFixed records, keys
dict{"a": 1}YesKey-value lookups
setYesUniqueness, membership
python
prices = [19.99, 5.49, 12.00]        # list: ordered, mutable
record = ("SKU-42", 19.99, 3)         # tuple: fixed record
product = {"sku": "SKU-42", "qty": 3} # dict: labeled fields
seen = {"SKU-42", "SKU-99"}           # set: unique membership

total = sum(prices)
print(f"Total: ${total:.2f}")         # f-string -> Total: $37.48

List & Dict Comprehensions

Comprehensions are the idiomatic way to transform and filter collections. They are shorter and usually faster than an equivalent for loop with append.

python
nums = [1, 2, 3, 4, 5, 6]

squares = [n * n for n in nums]            # [1, 4, 9, 16, 25, 36]
evens   = [n for n in nums if n % 2 == 0]  # [2, 4, 6]

# dict comprehension: sku -> price
skus = ["A", "B", "C"]
prices = [9.99, 4.50, 12.00]
catalog = {sku: price for sku, price in zip(skus, prices)}

# set comprehension: unique first letters
initials = {word[0] for word in ["apple", "avocado", "banana"]}  # {'a', 'b'}

Keep it readable

If a comprehension needs more than one condition or a nested loop, prefer an explicit loop. Clever one-liners that nobody can read are a false economy.

Functions

Functions package reusable logic. Use type hints to document intent and default arguments for optional parameters. Avoid mutable default arguments.

python
def normalize(values: list[float], target_max: float = 1.0) -> list[float]:
    """Scale values so the largest equals target_max."""
    hi = max(values)
    return [v / hi * target_max for v in values]

print(normalize([2, 4, 8]))          # [0.25, 0.5, 1.0]

# *args / **kwargs for flexible signatures
def summarize(*columns: str, sep: str = ", ") -> str:
    return sep.join(columns)

print(summarize("name", "age", sep=" | "))  # name | age

Gotcha: mutable defaults

Never write def f(x, cache=[]). The list is created once and shared across calls. Use cache=None then if cache is None: cache = [].

File I/O

Always open files with a with block so they close automatically, even if an error occurs. Modern code prefers pathlib.Path for filesystem paths.

python
from pathlib import Path

path = Path("data") / "notes.txt"

# write
path.write_text("line one\nline two\n", encoding="utf-8")

# read all lines, stripping the newline
with path.open(encoding="utf-8") as f:
    for line in f:
        print(line.rstrip())

# quick one-liner read
text = path.read_text(encoding="utf-8")

Working with CSV

CSV is the most common flat data format. The standard library csv module handles quoting and delimiters correctly, so avoid splitting on commas manually.

python
import csv
from pathlib import Path

rows = [
    {"name": "Ada", "score": 91},
    {"name": "Linus", "score": 88},
]

# write with a header
with Path("scores.csv").open("w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "score"])
    writer.writeheader()
    writer.writerows(rows)

# read back into dicts
with Path("scores.csv").open(newline="", encoding="utf-8") as f:
    for row in csv.DictReader(f):
        print(row["name"], int(row["score"]))

Working with JSON

JSON is the standard for API responses and config. The json module converts between Python objects and JSON text.

python
import json
from pathlib import Path

config = {"model": "v2", "threshold": 0.75, "tags": ["a", "b"]}

# object -> JSON string / file
Path("config.json").write_text(json.dumps(config, indent=2))

# JSON file -> object
loaded = json.loads(Path("config.json").read_text())
print(loaded["threshold"])   # 0.75

# parse an API-style string
payload = '{"ok": true, "count": 3}'
data = json.loads(payload)
print(data["count"])         # 3

Generators & Lazy Evaluation

Generators produce values one at a time instead of building a full list in memory. They are essential when streaming large files or infinite sequences.

python
def read_large(path):
    """Yield rows lazily; never loads the whole file."""
    with open(path, encoding="utf-8") as f:
        for line in f:
            yield line.rstrip()

# generator expression: like a comprehension but lazy
total = sum(len(row) for row in read_large("big.csv"))

# generators are single-use and evaluated on demand
gen = (n * n for n in range(3))
print(next(gen))   # 0
print(list(gen))   # [1, 4]

List vs generator

Use a list when you need to index, reuse, or know the length. Use a generator when the data is large or you only need to iterate once.

Why NumPy? A Motivation

Python lists are flexible but slow for numerical math because each element is a full Python object. NumPy stores numbers in a compact typed buffer and runs operations in optimized C, which is often 10-100x faster.

AspectPython listNumPy array
Element typeMixed, boxed objectsSingle fixed dtype
Math on all elementsManual loopVectorized
SpeedSlow for big dataFast (C-backed)
MemoryHigh overheadCompact
python
# list: element-by-element
prices = [10, 20, 30]
with_tax = [p * 1.08 for p in prices]

# numpy: one vectorized operation
import numpy as np
arr = np.array([10, 20, 30])
with_tax = arr * 1.08          # array([10.8, 21.6, 32.4])

A First Glimpse of pandas

pandas builds on NumPy to give you labeled tables (DataFrame). It reads CSV/JSON in one line and makes filtering and aggregation trivial. You will go deep on this in the NumPy & Pandas guide.

python
import pandas as pd

df = pd.read_csv("scores.csv")
print(df.head())

# filter and aggregate in one expression
top = df[df["score"] > 90]
print(df["score"].mean())

Virtual Environments

A virtual environment isolates a project's packages so different projects do not clash. Create one per project and never install libraries globally.

bash
# create and activate (macOS / Linux)
python -m venv .venv
source .venv/bin/activate

# Windows
.venv\Scripts\activate

# install and freeze dependencies
pip install numpy pandas
pip freeze > requirements.txt

# recreate later
pip install -r requirements.txt

Modern tooling

Tools like uv and poetry manage environments and lockfiles automatically. uv venv and uv pip install are dramatically faster than plain pip.

Practice Exercises

  1. Given a list of temperatures in Celsius, use a comprehension to build a new list of Fahrenheit values, keeping only those above freezing.
  2. Write a function word_count(text) that returns a dict mapping each word to how many times it appears.
  3. Read a CSV of products, and write a new CSV containing only rows where price > 100.
  4. Load a JSON config file, add a new key, and write it back with 2-space indentation.
  5. Write a generator that yields the running total of a sequence of numbers.
  6. Create a virtual environment, install pandas, and freeze the dependencies to requirements.txt.

Section navigation