Released · improving
Rate limiting guide · 1/6
A rate limiter caps how many requests are accepted within a period of time. You set a rule such as "100 requests per minute per user" or "5 login attempts per minute per IP", and requests beyond it are rejected or delayed. This chapter explains why we limit requests, which algorithms exist, and the vocabulary used throughout the rest of the guide.
Server resources are finite. When one client floods a service in a short time, everyone else gets slower responses, and in the worst case the whole service goes down. A rate limiter protects against:
| Term | Meaning |
|---|---|
| Limit | Maximum number of requests allowed in one window |
| Window | The time span over which requests are counted (1 second, 1 minute) |
| Key | The unit that gets its own counter (user ID, API key, IP address) |
| Burst | A group of requests arriving close together |
| Rate | The long-run average speed allowed (for example 10 per second) |
| Capacity | The most tokens a token bucket can hold, which is the largest burst it admits |
A fixed window counter takes only a few lines. The following chapters fix its weaknesses one at a time.
import time
from collections import defaultdict
LIMIT, WINDOW = 5, 60 # 5 requests per minute
counters = defaultdict(int) # (key, window number) -> count
def allow(key: str) -> bool:
window_id = int(time.time() // WINDOW)
counters[(key, window_id)] += 1
return counters[(key, window_id)] <= LIMITThis sketch never deletes old counters. Real code cleans up past windows or uses a store with expiry, such as Redis.
The same algorithm behaves very differently depending on what you count against.
user:42:POST /search give fine-grained control.Behind a reverse proxy, be careful with X-Forwarded-For. Clients can set it to anything, so only trust entries appended by proxies you operate.
def client_ip(remote_addr: str, forwarded_for: str | None,
trusted_proxies: set[str]) -> str:
# Only trust the header when the direct peer is our own proxy
if forwarded_for and remote_addr in trusted_proxies:
hops = [h.strip() for h in forwarded_for.split(",")]
# Walk from the right; the first untrusted hop is the real client
for hop in reversed(hops):
if hop not in trusted_proxies:
return hop
return remote_addrOver HTTP, a request that exceeds the limit normally gets 429 Too Many Requests, a status code defined in RFC 6585. Adding a Retry-After header, either a number of seconds or an HTTP date, tells well-behaved clients exactly how long to back off. Many APIs also report the remaining quota and reset time in headers, such as GitHub's x-ratelimit-remaining.
HTTP/1.1 429 Too Many Requests
Retry-After: 12
Content-Type: application/json
{"error": "rate_limited", "message": "Too many requests. Retry in 12 seconds."}Retry-After so clients know when to try again.
0 comments
Sign in · Sign in to leave a comment.
Be the first to comment.