Search Topics
Search across all FastAPI topics
Lifespan Events
Core ConceptLifespan events let you run code when your app starts up and shuts down — perfect for connecting to databases, loading ML models, or warming caches.
Your ML model takes 30 seconds to load. The first user to hit your API after deployment waits 30 seconds for a response. Every subsequent request is fast. The cold start is killing your user experience.
# First request after deploy:
GET /predict {"text": "classify this"}
# Server log:
INFO: Loading model from ./model.pkl...
INFO: Model loaded successfully (took 31.2s)
INFO: 200 /predict — 31,247ms ← First user waited 31 seconds!
# Second request:
INFO: 200 /predict — 42ms ← Everyone after is fineQuestion
What if you could load the model before the first request arrives? What if your app didn't start accepting traffic until the model, database pool, and cache were all ready? That's what lifespan events give you.
The Lifespan Timeline
Your app has three phases. Lifespan events let you run code in the startup and shutdown phases — the parts that happen before any user touches your API and after the last request is handled.
Startup
Load models, connect DBs, warm caches
Ready
App starts accepting requests
Running
Handle requests normally
Shutting Down
Close connections, cleanup
The key: your app doesn't accept any traffic during startup. Users never see that 31-second model load. By the time requests arrive, everything is ready.
The Problem
What you just learned
Loading resources on first request causes cold start delays
Lifespan events run setup code BEFORE the app accepts traffic
Shutdown events handle cleanup AFTER all requests are done
Old Approach (Deprecated)
You might see @app.on_event() in older tutorials. It still works, but it's deprecated. If you're starting fresh, skip ahead to the modern approach below.
from fastapi import FastAPI
app = FastAPI()
# Don't use in new code — deprecated
@app.on_event("startup")
async def startup():
print("Starting up...")
@app.on_event("shutdown")
async def shutdown():
print("Shutting down...")Watch out
The problem with on_event? Startup and shutdown are separate functions that can't share variables. If you create a database pool in startup, how do you close it in shutdown? You need a global variable. The modern approach solves this elegantly.
Modern Lifespan Function
The recommended approach uses an asynccontextmanager. Code before yield runs on startup. Code after yield runs on shutdown. And they share the same scope — no globals needed.
from contextlib import asynccontextmanager
from fastapi import FastAPI
@asynccontextmanager
async def lifespan(app: FastAPI):
# --- STARTUP ---
# Everything here runs before the first request
print("Starting up...")
yield # App runs here — handles requests
# --- SHUTDOWN ---
# Everything here runs after the last request
print("Shutting down...")
app = FastAPI(lifespan=lifespan)Lifespan Basics
What you just learned
asynccontextmanager is the modern, recommended way to handle startup/shutdown
Code before yield = startup, code after yield = shutdown
The same function scope means startup and shutdown can share variables
Pass the lifespan function to FastAPI(lifespan=lifespan)
Real-World Use Cases
The most common use cases are database connection pools, ML model loading, and HTTP clients. Here's what a real startup sequence looks like:
from contextlib import asynccontextmanager
from fastapi import FastAPI
import httpx
@asynccontextmanager
async def lifespan(app: FastAPI):
# 1. Database connection pool
app.state.db_pool = await create_db_pool(
"postgresql://localhost/mydb",
min_size=5,
max_size=20,
)
# 2. HTTP client for external APIs
app.state.http_client = httpx.AsyncClient()
print("DB pool and HTTP client ready")
yield # App handles requests here
# Clean up in reverse order
await app.state.http_client.aclose()
await app.state.db_pool.close()
print("Connections closed")
app = FastAPI(lifespan=lifespan)from contextlib import asynccontextmanager
from fastapi import FastAPI
# ML model loading — the exact fix for our cold start problem
@asynccontextmanager
async def lifespan(app: FastAPI):
# Load once on startup — before any request arrives
app.state.model = load_ml_model("model.pkl")
print(f"Model loaded: {app.state.model.version}")
yield # Model is ready for all requests
del app.state.model # Cleanup if needed
app = FastAPI(lifespan=lifespan)
@app.post("/predict")
async def predict(data: InputData):
# No cold start — model is already loaded!
result = app.state.model.predict(data.features)
return {"prediction": result}Sharing State via app.state
Resources created in the lifespan function are stored on app.state. In your endpoints, you access them through the Request object. Here's how:
from fastapi import FastAPI, Request
@app.get("/items")
async def list_items(request: Request):
# Access the DB pool created in lifespan
async with request.app.state.db_pool.acquire() as conn:
items = await conn.fetch("SELECT * FROM items")
return items
@app.get("/health")
async def health(request: Request):
# Check if resources are available
pool = request.app.state.db_pool
return {
"status": "healthy",
"db_pool_size": pool.get_size(),
"db_pool_free": pool.get_idle_size(),
}Insight
The pattern is always request.app.state.your_resource. The request.app gives you the FastAPI instance, and .state is a simple namespace where you can attach anything during lifespan.
State Management
What you just learned
Store startup resources on app.state (e.g., app.state.db_pool)
Access them in endpoints via request.app.state
This avoids global variables — state is tied to the app instance
The health endpoint pattern is great for monitoring
Startup Exception Kills the App
Your lifespan function tries to connect to Redis on startup. But Redis isn't running. What happens to your app?
@asynccontextmanager
async def lifespan(app: FastAPI):
# This throws ConnectionRefusedError if Redis is down
app.state.redis = await aioredis.from_url(
"redis://localhost:6379"
)
print("Redis connected")
yield
await app.state.redis.close()See It: Full Lifespan Cycle
Watch a real-world app initialize Redis, OpenAI, Gemini, and a cache on startup — serve requests using them — then clean everything up on shutdown.
Go Deeper: Full Production Example
Here's a production-grade lifespan that initializes multiple services — the exact pattern from the visualization above.
from contextlib import asynccontextmanager
from fastapi import FastAPI, Request
from openai import AsyncOpenAI
import google.generativeai as genai
import aioredis
@asynccontextmanager
async def lifespan(app: FastAPI):
# ── STARTUP (before yield) ──
# 1. Redis connection pool
app.state.redis = await aioredis.from_url(
"redis://localhost:6379", max_connections=20
)
# 2. OpenAI async client
app.state.openai = AsyncOpenAI(api_key=settings.openai_key)
# 3. Gemini model (loaded once, reused across requests)
genai.configure(api_key=settings.gemini_key)
app.state.gemini = genai.GenerativeModel("gemini-pro")
# 4. Warm cache with popular data
app.state.cache = {}
async with app.state.redis.client() as conn:
for key in ["config", "models", "limits"]:
app.state.cache[key] = await conn.get(f"app:{key}")
print("All services ready")
yield # ── APP RUNS HERE ──
# ── SHUTDOWN (after yield) ──
app.state.cache.clear()
await app.state.openai.close()
await app.state.redis.close()
print("All services cleaned up")
app = FastAPI(lifespan=lifespan)
@app.post("/chat")
async def chat(prompt: str, request: Request):
response = await request.app.state.openai.chat.completions.create(
model="gpt-4", messages=[{"role": "user", "content": prompt}]
)
# Cache the response in Redis
await request.app.state.redis.setex(
f"chat:{prompt[:50]}", 3600, response.choices[0].message.content
)
return {"reply": response.choices[0].message.content}Think about it...
If your lifespan startup function raises an exception, does the app still start accepting requests?
Hint: Think about what happens if code before yield raises — does yield ever execute?
Key Points
asynccontextmanager
The modern, recommended way to handle startup and shutdown
on_event Deprecated
@app.on_event() still works but should not be used in new code
app.state
Share resources between lifespan and endpoints via request.app.state
Cleanup Guarantee
Code after yield always runs, even if the app crashes