Search Topics
Search across all FastAPI topics
Rate Limiting
SecurityRate limiting protects your API from abuse by capping how many requests a client can make in a given time window. Without it, a single client can overwhelm your server.
Your API goes viral on Hacker News. 10,000 requests per second. Your database connection pool is exhausted in seconds. Every legitimate user gets 503 Service Unavailable. One bot can bring down your entire service.
# 2:15 PM — Hacker News front page
# Traffic spikes from 50 req/s to 10,000 req/s
sqlalchemy.exc.TimeoutError: QueuePool limit of 5 overflow 10
reached, connection timed out, timeout 30.00
# All endpoints return:
HTTP 503 Service Unavailable
{"detail": "Service temporarily unavailable"}
# 100% of real users affected. Revenue loss: $2,400/hour.Question
How do you keep your API alive when 10,000 requests per second hit it? You can't just scale infinitely — that's expensive and slow to react. What you need is a bouncer at the door: rate limiting.
It caps how many requests each client can make, so one runaway bot can't ruin the experience for everyone else.
How Rate Limiting Works
Every request gets checked against a counter. If you're under the limit, the request goes through. If you're over, you get a 429 and have to wait.
Request arrives
From IP 192.168.1.1
Check counter
This IP: 4 of 5 allowed
Under limit?
4 < 5 → YES
Process request
Increment counter to 5
Next request
Same IP, same minute
Check counter
This IP: 5 of 5 allowed
Under limit?
5 < 5 → NO
429 Too Many Requests
Retry-After: 60
The Concept
What you just learned
Rate limiting counts requests per client within a time window
When the limit is hit, the server returns 429 Too Many Requests
The Retry-After header tells clients when they can try again
Different endpoints can have different limits based on sensitivity
Setting Up slowapi
slowapi is the go-to rate limiting library for FastAPI. It wraps the limits library and plugs into Starlette middleware. Three lines to set up, one decorator per route.
# pip install slowapi
from fastapi import FastAPI, Request
from slowapi import Limiter, _rate_limit_exceeded_handler
from slowapi.util import get_remote_address
from slowapi.errors import RateLimitExceeded
# Create limiter keyed by client IP
limiter = Limiter(key_func=get_remote_address)
app = FastAPI()
app.state.limiter = limiter
# This handler returns a proper 429 response
app.add_exception_handler(RateLimitExceeded, _rate_limit_exceeded_handler)Insight
get_remote_address extracts the client's IP from the request. That's fine for most cases, but if your app is behind a load balancer or reverse proxy, you'll need to read the X-Forwarded-For header instead. Otherwise, ALL your users look like they're coming from the same IP (the proxy's).
Per-Route Limits
Not all endpoints are equal. Your login endpoint should be locked down tight (brute-force prevention), while read-heavy endpoints can be more generous. Here's how you set different limits per route.
from fastapi import FastAPI, Request
from slowapi import Limiter
from slowapi.util import get_remote_address
limiter = Limiter(key_func=get_remote_address)
app = FastAPI()
# Strict limit on login — prevent brute force
@app.post("/login")
@limiter.limit("5/minute")
async def login(request: Request):
return {"message": "Login endpoint"}
# Moderate limit on writes
@app.post("/items")
@limiter.limit("30/minute")
async def create_item(request: Request):
return {"message": "Item created"}
# Generous limit on reads
@app.get("/items")
@limiter.limit("100/minute")
async def list_items(request: Request):
return {"items": []}Per-Route Configuration
What you just learned
Use @limiter.limit() decorator to set per-route limits
Login/auth endpoints should be the most restrictive (5-10/minute)
Write endpoints need moderate limits (20-50/minute)
Read endpoints can be generous (100+/minute)
Global Rate Limiting
Don't want to decorate every single route? Set a default limit for all routes, then override specific ones. And yes — you can exempt endpoints like health checks entirely.
from slowapi import Limiter
from slowapi.util import get_remote_address
limiter = Limiter(
key_func=get_remote_address,
default_limits=["60/minute"], # All routes: 60 req/min
)
app = FastAPI()
app.state.limiter = limiter
# Uses the global default: 60/minute
@app.get("/items")
async def list_items(request: Request):
return {"items": []}
# Override: stricter limit for sensitive endpoint
@app.post("/login")
@limiter.limit("5/minute")
async def login(request: Request):
return {"message": "Login"}
# Override: no limit for health checks
@app.get("/health")
@limiter.exempt
async def health():
return {"status": "ok"}Custom Rate Limit Responses
The default 429 response is bare-bones. Your frontend team will appreciate a structured error with a clear retry_after field and a human-readable message.
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
from slowapi import Limiter
from slowapi.errors import RateLimitExceeded
from slowapi.util import get_remote_address
limiter = Limiter(key_func=get_remote_address)
app = FastAPI()
app.state.limiter = limiter
@app.exception_handler(RateLimitExceeded)
async def rate_limit_handler(request: Request, exc: RateLimitExceeded):
return JSONResponse(
status_code=429,
content={
"error": "rate_limit_exceeded",
"message": f"Too many requests. Limit: {exc.detail}",
"retry_after": "60 seconds",
},
headers={"Retry-After": "60"}, # Clients can read this header
)Production Setup
What you just learned
default_limits sets a baseline for all routes
@limiter.exempt removes limits from health check and status endpoints
Custom 429 handlers give clients actionable information
The Retry-After header is an HTTP standard — well-behaved clients respect it
Try It: Rate Limit Simulator
Click "Send Request" repeatedly and watch what happens when you exceed the limit. The gauge fills up, and once you hit the cap, requests get rejected with 429. Frustrating, right? Now imagine that's what your users see when a bot hammers your API without limits.
Go Deeper: The NAT Problem
IP-based rate limiting has a blind spot. What happens when hundreds of users share one IP?
Corporate NAT Blocks Legitimate Users
A corporate office with 500 employees shares one public IP address. Your rate limit of 100 requests/minute per IP means ALL 500 employees share one bucket.
# Your rate limit config:
limiter = Limiter(key_func=get_remote_address)
@app.get("/api/data")
@limiter.limit("100/minute") # Per IP
async def get_data(request: Request):
return {"data": "..."}Think about it...
You set a rate limit of 100 requests/minute per IP. A corporate office has 500 employees behind one NAT IP. What happens?
Hint: Think about what the server sees — one IP address for all 500 employees.
Key Points
slowapi
The go-to rate limiting library for FastAPI applications
Per-Route Limits
Apply stricter limits to sensitive endpoints like login
429 Status Code
Returned when a client exceeds the rate limit with Retry-After header
Key Function
Rate limits keyed by IP address, user ID, or API key