What "Serverless" Really Means
Serverless does not mean there are no servers — it means you never provision or manage them. The cloud provider runs your code on demand, scales it automatically from zero to thousands of concurrent executions, and bills you only for the compute you actually consume, usually per millisecond.
The defining traits of a serverless platform are:
- No server management — no patching, capacity planning, or OS upkeep.
- Automatic scaling — horizontal scale is transparent and instant.
- Scale to zero — with no traffic you pay nothing for compute.
- Event-driven — code runs in response to a trigger (HTTP, queue, file upload, cron).
- Stateless — each invocation is independent; persistent state lives elsewhere.
FaaS Across the Big Three
Function-as-a-Service (FaaS) is the core serverless building block. The three major clouds offer conceptually identical products with different names and limits.
| Capability | AWS | Azure | GCP |
|---|---|---|---|
| FaaS | Lambda | Azure Functions | Cloud Run functions |
| Container serverless | App Runner / Fargate | Container Apps | Cloud Run |
| API gateway | API Gateway | API Management | API Gateway |
| Message queue | SQS | Queue Storage | Pub/Sub |
| Orchestration | Step Functions | Durable Functions | Workflows |
Writing and Deploying a Function
AWS Lambda (Node.js)
A Lambda handler receives an event and a context, and returns a response object.
// index.mjs
export const handler = async (event) => {
const name = event.queryStringParameters?.name ?? "world";
return {
statusCode: 200,
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ message: `Hello, ${name}!` }),
};
};
# Package and deploy with the AWS CLI
zip function.zip index.mjs
aws lambda create-function \
--function-name hello \
--runtime nodejs22.x \
--handler index.handler \
--zip-file fileb://function.zip \
--role arn:aws:iam::123456789012:role/lambda-exec
# Update code later and invoke
aws lambda update-function-code --function-name hello --zip-file fileb://function.zip
aws lambda invoke --function-name hello --payload '{}' out.json
GCP Cloud Run functions (Python)
# main.py
import functions_framework
@functions_framework.http
def hello(request):
name = request.args.get("name", "world")
return {"message": f"Hello, {name}!"}
gcloud functions deploy hello \
--gen2 --runtime=python312 --region=us-central1 \
--source=. --entry-point=hello \
--trigger-http --allow-unauthenticated
Azure Functions (C#, CLI)
func init MyFuncApp --dotnet
func new --name Hello --template "HTTP trigger"
func azure functionapp publish my-func-app
Triggers & Event Sources
Functions are wired to event sources. The trigger determines the shape of the event payload and the invocation model (synchronous vs asynchronous vs stream/poll).
| Trigger | Example use | Model |
|---|---|---|
| HTTP | REST API / webhook | Synchronous |
| Queue / topic | Order processing, fan-out | Async / poll |
| Object storage | Thumbnail on upload | Async |
| Schedule (cron) | Nightly report | Async |
| Stream | Kinesis / DynamoDB changes | Poll (batched) |
Cold Starts
When a function scales from zero, the platform must download your code, start a runtime, and initialize it before the first request runs. This latency is the cold start. Warm invocations reuse the container and skip it.
Reducing Cold Starts
Keep deployment packages small, initialize expensive clients (DB pools, SDKs) outside the handler so they persist across warm invocations, prefer lighter runtimes, and use provisioned concurrency (Lambda) or min instances (Cloud Run) for latency-sensitive paths.
// GOOD: client created once, reused on warm invocations
import { DynamoDBClient } from "@aws-sdk/client-dynamodb";
const db = new DynamoDBClient({}); // init OUTSIDE the handler
export const handler = async (event) => {
// db is already warm here
};
// Keep a Lambda warm with provisioned concurrency
// aws lambda put-provisioned-concurrency-config \
// --function-name hello --qualifier prod \
// --provisioned-concurrent-executions 5
Concurrency, Scaling & Limits
Serverless scales by running one concurrent execution per in-flight request (for classic FaaS). Ten simultaneous requests means ten instances. Cloud Run differs: a single container can serve many concurrent requests, which lowers cost for I/O-bound workloads.
Watch the platform limits — they shape architecture decisions:
- Timeout — Lambda maxes at 15 minutes; long jobs need Step Functions/Workflows.
- Memory / CPU — CPU scales with memory allocation; tune both together.
- Payload size — request/response bodies are capped (e.g. 6 MB sync for Lambda).
- Account concurrency — a soft cap that can throttle bursts; request increases early.
Common Patterns & Pitfalls
Design for the serverless model, not against it:
- Idempotency — async events can be delivered more than once; make handlers safe to retry.
- Dead-letter queues — route failed events to a DLQ instead of losing them.
- No local state — write to a database, cache, or object store between invocations.
- Fan-out with queues — decouple producers and consumers via SQS/Pub/Sub.
Cost Watch
Serverless is cheap at low and spiky traffic but can be more expensive than a VM at sustained high volume. Model your steady-state requests-per-second before committing an always-hot workload to FaaS.
Practice Exercises
- Deploy an HTTP-triggered function on any cloud that returns JSON, then invoke it from
curl. - Add an object-storage trigger so uploading a file logs its name and size.
- Move an expensive client initialization outside the handler and measure the cold vs warm latency difference.
- Wire a queue (SQS or Pub/Sub) to a function and make the handler idempotent so duplicate deliveries are safe.
- Configure provisioned concurrency (or min instances) and confirm cold starts disappear under load.
- Build a two-step Step Functions/Workflows pipeline that calls two functions in sequence and handles a failure with a retry.