Diagramming System Design: Rate Limiters
Rate limiting controls how many requests users can make in a set time period. It protects systems from overload, abuse, and excessive costs, and ensures fair usage.
Key Design Factors
- Request volume
- Latency tolerance
- Client type
- Infrastructure
Decisions to Make
- Limit type: per user/IP
- Response: e.g. HTTP 429
- Rate limiter location: CDN, app, etc.
- Storage: in-memory, centralized, or distributed
Popular Algorithms
Fixed Window
- Simple but can be bypassed with bursts
Sliding Window Log
- Accurate but resource-heavy
Sliding Window Counter
- Balanced and efficient
Token Bucket
- Allows bursts, easy to implement
Leaky Bucket
- Queues requests, smooths traffic
Scaling requires a distributed design and shared or cached state storage.
All production-grade services rely on rate limiters. For instance, a user may post at most five articles per minute, or an API responds at most 500 times an hour.
Benefits of Rate Limiting
- Prevents excessive use of resources: High volume requests may bring down servers; a rate limiter prevents cascading failures.
- Prevents excessive use of external resources: Rate limits can prevent cost overruns when using paid external services.
- Ensures fair access for all users: No single user can monopolize resources.
Considerations When Designing a Rate Limiter
- Anticipated number of callers
- Average and peak requests per client
- Average request size
- Likelihood of traffic surges
- Maximum acceptable latency
Responding to Rate-Limited Requests
- Common responses include blocking excess requests with HTTP status codes 429 (Too Many Requests) or 503 (Service Unavailable).
- Another approach is to accept excess requests but not process them, returning HTTP status code 200 (OK).
Locations for Implementing Rate Limiters
- Content Delivery Network (CDN): Some may be configured to rate-limit requests.
- Reverse Proxy: Like Nginx, which offers rate-limiting options.
- API Gateway: Controls access to groups of endpoints.
- Application: Implement custom middleware for fine-grained control.
Tracking Rate Limiting State
- In-memory storage fits low-scale services.
- Centralized storage introduces latency and potential bottlenecks.
Common Rate-Limiting Algorithms
Fixed Window Counter
- Divides time into fixed windows with counters. Exceeding limits means denying excess requests until the window resets.
Sliding Window Log
- Logs timestamps for requests to manage counts accurately.
Token Bucket Algorithm
- Uses a bucket to track tokens for requests. Tokens are added at a predefined rate.
Conclusion
Rate limiters are essential for protecting against overloads, optimizing performance, and ensuring fair access to resources. Design decisions include setting request limits, deciding on tracking methods, and selecting the right algorithm based on application needs.