Search Topics

Search across all FastAPI topics

GitHub

Lifespan Events

Core Concept

Lifespan events let you run code when your app starts up and shuts down — perfect for connecting to databases, loading ML models, or warming caches.

Your ML model takes 30 seconds to load. The first user to hit your API after deployment waits 30 seconds for a response. Every subsequent request is fast. The cold start is killing your user experience.

terminal
# First request after deploy:
GET /predict {"text": "classify this"}

# Server log:
INFO: Loading model from ./model.pkl...
INFO: Model loaded successfully (took 31.2s)
INFO: 200 /predict — 31,247ms  ← First user waited 31 seconds!

# Second request:
INFO: 200 /predict — 42ms  ← Everyone after is fine

Question

What if you could load the model before the first request arrives? What if your app didn't start accepting traffic until the model, database pool, and cache were all ready? That's what lifespan events give you.

The Lifespan Timeline

Your app has three phases. Lifespan events let you run code in the startup and shutdown phases — the parts that happen before any user touches your API and after the last request is handled.

Startup

Load models, connect DBs, warm caches

Ready

App starts accepting requests

Running

Handle requests normally

Shutting Down

Close connections, cleanup

The key: your app doesn't accept any traffic during startup. Users never see that 31-second model load. By the time requests arrive, everything is ready.

The Problem

What you just learned

Loading resources on first request causes cold start delays

Lifespan events run setup code BEFORE the app accepts traffic

Shutdown events handle cleanup AFTER all requests are done

Old Approach (Deprecated)

You might see @app.on_event() in older tutorials. It still works, but it's deprecated. If you're starting fresh, skip ahead to the modern approach below.

main.py
from fastapi import FastAPI

app = FastAPI()

# Don't use in new code — deprecated
@app.on_event("startup")
async def startup():
    print("Starting up...")

@app.on_event("shutdown")
async def shutdown():
    print("Shutting down...")

Watch out

The problem with on_event? Startup and shutdown are separate functions that can't share variables. If you create a database pool in startup, how do you close it in shutdown? You need a global variable. The modern approach solves this elegantly.

Modern Lifespan Function

The recommended approach uses an asynccontextmanager. Code before yield runs on startup. Code after yield runs on shutdown. And they share the same scope — no globals needed.

main.py
from contextlib import asynccontextmanager
from fastapi import FastAPI

@asynccontextmanager
async def lifespan(app: FastAPI):
    # --- STARTUP ---
    # Everything here runs before the first request
    print("Starting up...")

    yield  # App runs here — handles requests

    # --- SHUTDOWN ---
    # Everything here runs after the last request
    print("Shutting down...")

app = FastAPI(lifespan=lifespan)

Lifespan Basics

What you just learned

asynccontextmanager is the modern, recommended way to handle startup/shutdown

Code before yield = startup, code after yield = shutdown

The same function scope means startup and shutdown can share variables

Pass the lifespan function to FastAPI(lifespan=lifespan)

Real-World Use Cases

The most common use cases are database connection pools, ML model loading, and HTTP clients. Here's what a real startup sequence looks like:

main.py
from contextlib import asynccontextmanager
from fastapi import FastAPI
import httpx

@asynccontextmanager
async def lifespan(app: FastAPI):
    # 1. Database connection pool
    app.state.db_pool = await create_db_pool(
        "postgresql://localhost/mydb",
        min_size=5,
        max_size=20,
    )

    # 2. HTTP client for external APIs
    app.state.http_client = httpx.AsyncClient()

    print("DB pool and HTTP client ready")
    yield  # App handles requests here

    # Clean up in reverse order
    await app.state.http_client.aclose()
    await app.state.db_pool.close()
    print("Connections closed")

app = FastAPI(lifespan=lifespan)
ml_app.py
from contextlib import asynccontextmanager
from fastapi import FastAPI

# ML model loading — the exact fix for our cold start problem
@asynccontextmanager
async def lifespan(app: FastAPI):
    # Load once on startup — before any request arrives
    app.state.model = load_ml_model("model.pkl")
    print(f"Model loaded: {app.state.model.version}")

    yield  # Model is ready for all requests

    del app.state.model  # Cleanup if needed

app = FastAPI(lifespan=lifespan)

@app.post("/predict")
async def predict(data: InputData):
    # No cold start — model is already loaded!
    result = app.state.model.predict(data.features)
    return {"prediction": result}

Sharing State via app.state

Resources created in the lifespan function are stored on app.state. In your endpoints, you access them through the Request object. Here's how:

main.py
from fastapi import FastAPI, Request

@app.get("/items")
async def list_items(request: Request):
    # Access the DB pool created in lifespan
    async with request.app.state.db_pool.acquire() as conn:
        items = await conn.fetch("SELECT * FROM items")
    return items

@app.get("/health")
async def health(request: Request):
    # Check if resources are available
    pool = request.app.state.db_pool
    return {
        "status": "healthy",
        "db_pool_size": pool.get_size(),
        "db_pool_free": pool.get_idle_size(),
    }

Insight

The pattern is always request.app.state.your_resource. The request.app gives you the FastAPI instance, and .state is a simple namespace where you can attach anything during lifespan.

State Management

What you just learned

Store startup resources on app.state (e.g., app.state.db_pool)

Access them in endpoints via request.app.state

This avoids global variables — state is tied to the app instance

The health endpoint pattern is great for monitoring

Startup Exception Kills the App

Your lifespan function tries to connect to Redis on startup. But Redis isn't running. What happens to your app?

Broken code
main.py
@asynccontextmanager
async def lifespan(app: FastAPI):
    # This throws ConnectionRefusedError if Redis is down
    app.state.redis = await aioredis.from_url(
        "redis://localhost:6379"
    )
    print("Redis connected")

    yield

    await app.state.redis.close()

See It: Full Lifespan Cycle

Watch a real-world app initialize Redis, OpenAI, Gemini, and a cache on startup — serve requests using them — then clean everything up on shutdown.

Go Deeper: Full Production Example

Here's a production-grade lifespan that initializes multiple services — the exact pattern from the visualization above.

main.py
from contextlib import asynccontextmanager
from fastapi import FastAPI, Request
from openai import AsyncOpenAI
import google.generativeai as genai
import aioredis

@asynccontextmanager
async def lifespan(app: FastAPI):
    # ── STARTUP (before yield) ──
    # 1. Redis connection pool
    app.state.redis = await aioredis.from_url(
        "redis://localhost:6379", max_connections=20
    )

    # 2. OpenAI async client
    app.state.openai = AsyncOpenAI(api_key=settings.openai_key)

    # 3. Gemini model (loaded once, reused across requests)
    genai.configure(api_key=settings.gemini_key)
    app.state.gemini = genai.GenerativeModel("gemini-pro")

    # 4. Warm cache with popular data
    app.state.cache = {}
    async with app.state.redis.client() as conn:
        for key in ["config", "models", "limits"]:
            app.state.cache[key] = await conn.get(f"app:{key}")

    print("All services ready")
    yield  # ── APP RUNS HERE ──

    # ── SHUTDOWN (after yield) ──
    app.state.cache.clear()
    await app.state.openai.close()
    await app.state.redis.close()
    print("All services cleaned up")

app = FastAPI(lifespan=lifespan)

@app.post("/chat")
async def chat(prompt: str, request: Request):
    response = await request.app.state.openai.chat.completions.create(
        model="gpt-4", messages=[{"role": "user", "content": prompt}]
    )
    # Cache the response in Redis
    await request.app.state.redis.setex(
        f"chat:{prompt[:50]}", 3600, response.choices[0].message.content
    )
    return {"reply": response.choices[0].message.content}

Think about it...

If your lifespan startup function raises an exception, does the app still start accepting requests?

Hint: Think about what happens if code before yield raises — does yield ever execute?

Key Points

asynccontextmanager

The modern, recommended way to handle startup and shutdown

on_event Deprecated

@app.on_event() still works but should not be used in new code

app.state

Share resources between lifespan and endpoints via request.app.state

Cleanup Guarantee

Code after yield always runs, even if the app crashes