Diagramming System Design: Rate Limiters

Rate limiting controls how many requests users can make in a set time period. It protects systems from overload, abuse, and excessive costs, and ensures fair usage.

Key Design Factors

  • Request volume
  • Latency tolerance
  • Client type
  • Infrastructure

Decisions to Make

  • Limit type: per user/IP
  • Response: e.g. HTTP 429
  • Rate limiter location: CDN, app, etc.
  • Storage: in-memory, centralized, or distributed

Popular Algorithms

Fixed Window

  • Simple but can be bypassed with bursts

Sliding Window Log

  • Accurate but resource-heavy

Sliding Window Counter

  • Balanced and efficient

Token Bucket

  • Allows bursts, easy to implement

Leaky Bucket

  • Queues requests, smooths traffic

Scaling requires a distributed design and shared or cached state storage.

All production-grade services rely on rate limiters. For instance, a user may post at most five articles per minute, or an API responds at most 500 times an hour.

Benefits of Rate Limiting

  1. Prevents excessive use of resources: High volume requests may bring down servers; a rate limiter prevents cascading failures.
  2. Prevents excessive use of external resources: Rate limits can prevent cost overruns when using paid external services.
  3. Ensures fair access for all users: No single user can monopolize resources.

Considerations When Designing a Rate Limiter

  • Anticipated number of callers
  • Average and peak requests per client
  • Average request size
  • Likelihood of traffic surges
  • Maximum acceptable latency

Responding to Rate-Limited Requests

  • Common responses include blocking excess requests with HTTP status codes 429 (Too Many Requests) or 503 (Service Unavailable).
  • Another approach is to accept excess requests but not process them, returning HTTP status code 200 (OK).

Locations for Implementing Rate Limiters

  1. Content Delivery Network (CDN): Some may be configured to rate-limit requests.
  2. Reverse Proxy: Like Nginx, which offers rate-limiting options.
  3. API Gateway: Controls access to groups of endpoints.
  4. Application: Implement custom middleware for fine-grained control.

Tracking Rate Limiting State

  • In-memory storage fits low-scale services.
  • Centralized storage introduces latency and potential bottlenecks.

Common Rate-Limiting Algorithms

Fixed Window Counter

  • Divides time into fixed windows with counters. Exceeding limits means denying excess requests until the window resets.

Sliding Window Log

  • Logs timestamps for requests to manage counts accurately.

Token Bucket Algorithm

  • Uses a bucket to track tokens for requests. Tokens are added at a predefined rate.

Conclusion

Rate limiters are essential for protecting against overloads, optimizing performance, and ensuring fair access to resources. Design decisions include setting request limits, deciding on tracking methods, and selecting the right algorithm based on application needs.