Search Topics

Search across all FastAPI topics

GitHub

Uvicorn & Gunicorn

Under the Hood

Your API crashed at 3 AM. No auto-restart. No backup process. Four hours of downtime before anyone noticed.

You deploy with "uvicorn main:app" in production. It works fine until your single process crashes after a memory leak. No auto-restart. No load distribution. Your entire API is down and nobody knows.

terminal
# Production server (single uvicorn process):
$ uvicorn main:app --host 0.0.0.0 --port 8000

# 3:47 AM — Process crashes due to memory leak
MemoryError: Unable to allocate 512 MiB

# No process manager → No auto-restart
# API is DOWN. No one is alerted.
# Users see: ERR_CONNECTION_REFUSED
# Duration of outage: 4 hours (until someone checks manually)

Question

One process. One point of failure. When it dies, everything dies. How do production apps avoid this? They use a process manager to keep multiple copies of the app running — and restart any that crash.

Two Tools, Two Jobs

Uvicorn and Gunicorn do completely different things. Understanding which does what is the key to a production-ready deployment.

Gunicorn

The manager. Spawns workers, monitors health, restarts crashes.

Uvicorn Worker 1

Runs your app. Handles requests.

Uvicorn Worker 2

Another copy. Same app.

Uvicorn Worker 3

Yet another. Load distributed.

uvicorn vs gunicorn

What you just learned

Uvicorn is the ASGI server — it runs your app, handles HTTP/WebSocket, manages the event loop

Gunicorn is the process manager — it spawns workers, monitors health, restarts crashes

Together they give you multi-core utilization and fault tolerance

Uvicorn: The Engine

Uvicorn is the process that actually runs your FastAPI app. It listens for connections, parses HTTP, and feeds requests to your async handlers. Think of it as the engine of a car — it does the actual work.

terminal
# Run your FastAPI app with Uvicorn
uvicorn main:app --reload --host 0.0.0.0 --port 8000

# What this means:
# main   → the file "main.py"
# app    → the FastAPI() instance in that file
# --reload  → restart on code changes (dev only!)
# --host    → listen on all interfaces
# --port    → listen on port 8000

Watch out

A single Uvicorn process can handle thousands of concurrent connections. But it's still one process. If it crashes, your API is dead. If you have 4 CPU cores, you're only using one of them.

Gunicorn: The Fleet Manager

Gunicorn doesn't run your app directly. It spawns multiple worker processes, each running their own copy of your app. It's the fleet manager — it doesn't drive any cars, but it manages the drivers.

terminal
# Gunicorn with Uvicorn workers
gunicorn main:app -k uvicorn.workers.UvicornWorker -w 4

# What this means:
# main:app → your FastAPI application
# -k uvicorn.workers.UvicornWorker → each worker runs Uvicorn
# -w 4     → spawn 4 worker processes

# Each worker is a separate process with its own event loop
# 4 workers on 4 cores = full CPU utilization
# Worker crashes? Gunicorn restarts it automatically.

production deployment

What you just learned

Uvicorn alone = great for dev, risky for production (single point of failure)

Gunicorn + Uvicorn workers = multi-core utilization with automatic crash recovery

Gunicorn handles the hard stuff: spawning, health checks, graceful restarts

When to Use Which

The right choice depends on where you're deploying. Here's the cheat sheet:

Development

Uvicorn with --reload

terminal
uvicorn main:app --reload

Auto-restarts on code changes. Single process, easy to debug.

Production (single core)

Uvicorn alone

terminal
uvicorn main:app --host 0.0.0.0 --port 8000

Single process, no reload overhead. Good for containers (one process per container).

Production (multi-core)

Gunicorn + Uvicorn workers

terminal
gunicorn main:app -k uvicorn.workers.UvicornWorker -w 4

Multiple worker processes utilize all CPU cores. Gunicorn handles process management.

Docker

Uvicorn directly

Dockerfile
# Dockerfile
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
# Scale by running multiple containers instead of multiple workers

One process per container. Scale horizontally with Kubernetes or Docker Compose replicas.

Real-world

The Docker approach is increasingly popular. Instead of one server running 4 Gunicorn workers, you run 4 containers each with one Uvicorn process. Kubernetes handles the "process management" that Gunicorn would. If a container crashes, Kubernetes restarts it. Same concept, different level.

Go Deeper: The Full Production Setup

Here's what a production Gunicorn + Uvicorn deployment looks like with all the important options spelled out:

terminal
# Full production command with all options
gunicorn main:app \
    -k uvicorn.workers.UvicornWorker \  # Use Uvicorn as worker
    -w 4 \                              # 4 worker processes
    --bind 0.0.0.0:8000 \              # Listen on all interfaces
    --timeout 120 \                     # Worker timeout (seconds)
    --graceful-timeout 30 \             # Graceful shutdown time
    --access-logfile - \                # Log to stdout
    --error-logfile -                    # Errors to stdout

# Architecture:
# ┌─────────────────────────────────┐
# │  Gunicorn (Master Process)      │
# │  - Manages worker lifecycle     │
# │  - Handles signals (SIGTERM)    │
# │  - Restarts crashed workers     │
# ├─────────┬──────────┬────────────┤
# │ Worker 1│ Worker 2 │ Worker 3   │  (Uvicorn)
# │ (loop)  │ (loop)   │ (loop)     │
# │ ~1000   │ ~1000    │ ~1000      │  concurrent
# │ conns   │ conns    │ conns      │  connections
# └─────────┴──────────┴────────────┘

deployment patterns

What you just learned

Development: uvicorn with --reload. Production: gunicorn + uvicorn workers.

Docker changes the game: one process per container, scale with orchestration

--timeout, --graceful-timeout, and logging flags are critical for production

Try It: Process Architecture

Toggle between Uvicorn solo and Gunicorn + Uvicorn to see how requests flow through each architecture. Hit "Simulate Traffic" to watch requests get distributed across workers in real time.

Think about it...

You have a 4-core server. Should you run 4 Uvicorn workers or 8?

Hint: Think about what your workers spend most of their time doing.

Key Points

Uvicorn = ASGI Server

Runs your app, handles HTTP/WebSocket, manages the event loop

Gunicorn = Process Manager

Spawns worker processes, monitors health, handles restarts

Dev = Uvicorn Only

Use --reload for auto-restart during development

Prod = Gunicorn + Uvicorn

Or scale with Docker containers running single Uvicorn processes

These are the patterns that trip up developers most often. Switch between Wrong and Fixed to compare the code side by side.

1
Using --reload in production
The dev reload flag adds overhead and is unstable for prod
Don't do this
terminal
# WRONG: --reload watches files and restarts on changes
# This adds overhead and can cause downtime
uvicorn main:app --reload --host 0.0.0.0 --port 8000
The --reload flag watches your file system for changes and restarts the server. In production, this wastes resources and can cause unexpected restarts. Only use --reload during development.
2
Not setting worker count based on CPU cores
Using arbitrary worker counts instead of a formula
Don't do this
terminal
# Arbitrary number — might be too many or too few
gunicorn main:app -k uvicorn.workers.UvicornWorker -w 20
Too many workers waste memory and cause context-switching overhead. Too few waste CPU cores. The formula (2 x cores) + 1 is a good starting point. For async FastAPI apps, you can often use fewer workers since each handles many concurrent requests.