Imagine running a promotion and your backend collapsing under a flood of requests. At peak load of 10,000 rps, in-app rate limiting fails to sync between services — each microservice holds its own counter, and a burst causes limit overruns of 200%. Centralized request throttling via an API Gateway is the single control point where policies are applied to all traffic without changing business logic. We audited a fintech client: after implementing Kong rate limiting, 429 errors dropped by 90%, and backend load by 70%.
Why API Gateway Instead of In-App Rate Limiting?
An in-app solution requires changes in every service, code duplication, and coordination. A Gateway is a single layer where policies are applied uniformly. No need to touch business logic — just configure a plugin. Plus, unified monitoring, logging, and the ability to change limits on the fly without deployments.
Comparison of approaches:
| Characteristic | API Gateway | In-app | Middleware |
|---|---|---|---|
| Single policy point | Yes | No | Partial |
| Change without deploy | Yes | No | Yes |
| Load on services | Minimal | Medium | Minimal |
| Configuration complexity | Low | High | Medium |
Which Algorithm to Choose?
There are several algorithms, and the choice depends on the load profile. Token Bucket (used by Kong) allows bursts up to a configurable limit, whereas Leaky Bucket (available in APISIX) smoothens traffic. Fixed Window is simple but has edge effects; Sliding Window offers high accuracy.
| Algorithm | Behavior | Burst | Accuracy | Application |
|---|---|---|---|---|
| Token Bucket | Replenishable bucket of tokens | Yes | High | Most API Gateways (configurable burst) |
| Leaky Bucket | Leaking at constant rate | No | High | Predictable load, throttling protection |
| Fixed Window | Counter over fixed window | No | Low | Simple scenarios, but window edge effect |
| Sliding Window | Sliding window with weights | Yes | High | Critical accuracy, A/B tests |
Configuration on Specific Gateways
Configuration in Kong
Kong supports multi-level limits: global, per service, per consumer. Uses Redis for synchronization. Example via Admin API:
# Global limit (all services) — 1000 requests per minute curl -X POST http://localhost:8001/plugins \ -d "name=rate-limiting" \ -d "config.minute=1000" \ -d "config.hour=20000" \ -d "config.policy=redis" \ -d "config.redis_host=redis" \ -d "config.limit_by=ip" # Limit at service level for payments-api — 10 rps curl -X POST http://localhost:8001/services/payments-api/plugins \ -d "name=rate-limiting" \ -d "config.second=10" \ -d "config.minute=200" \ -d "config.limit_by=consumer" # Limit for consumer free-tier — 60 requests per minute curl -X POST http://localhost:8001/consumers/free-tier/plugins \ -d "name=rate-limiting" \ -d "config.minute=60" \ -d "config.hour=500" Response headers contain limit and reset time. Official documentation: Kong Rate Limiting.
Configuration in APISIX
APISIX provides three plugins: limit-count (counter), limit-req (leaky bucket), limit-conn (concurrent connections). Example via Admin API:
{ "plugins": { "limit-count": { "count": 100, "time_window": 60, "rejected_code": 429, "rejected_msg": "Too many requests", "key": "consumer_name", "policy": "redis", "redis_host": "redis", "redis_port": 6379, "redis_database": 0, "show_limit_quota_header": true }, "limit-req": { "rate": 10, "burst": 5, "key": "remote_addr", "rejected_code": 429 }, "limit-conn": { "conn": 50, "burst": 10, "key": "remote_addr", "rejected_code": 503 } } } limit-req implements Leaky Bucket — requests beyond the rate go into a burst queue, then 429. limit-conn limits concurrent connections. APISIX's limit-req plugin provides 2x better burst control than fixed window approaches.
Configuration in AWS API Gateway
AWS uses Usage Plans with throttle and quota. Example Terraform for four tiers:
resource "aws_api_gateway_usage_plan" "tiers" { for_each = { free = { rate = 10, burst = 5, quota = 1000, period = "DAY" } basic = { rate = 50, burst = 25, quota = 10000, period = "DAY" } pro = { rate = 200, burst = 100, quota = 100000, period = "DAY" } enterprise = { rate = 1000, burst = 500, quota = 0, period = "DAY" } } name = "plan-${each.key}" api_stages { api_id = aws_api_gateway_rest_api.main.id stage = "prod" } throttle_settings { rate_limit = each.value.rate burst_limit = each.value.burst } dynamic "quota_settings" { for_each = each.value.quota > 0 ? [1] : [] content { limit = each.value.quota period = each.value.period } } } Dynamic Throttling by Business Attributes
Rate limits don't always depend solely on IP or API key. Often logic is needed: user subscription, resource type, time of day. Example custom Kong plugin that fetches limits from a billing service:
local function get_rate_limit(consumer_id) local cache_key = "rate:" .. consumer_id local cached = kong.cache:get(cache_key) if cached then return cached end -- Request to billing service local client = httpc.new() local res = client:request_uri("http://billing-service/limits/" .. consumer_id) local limits = cjson.decode(res.body) kong.cache:set(cache_key, limits, 300) -- cache 5 minutes return limits end local limits = get_rate_limit(consumer_id) -- limits = { minute: 1000, hour: 10000 } Common Mistakes and How to Avoid Them
- Burst not configured: even legitimate spikes are rejected. Set burst to 50-100% of rate.
- No whitelist for internal services: monitoring fails. Use
ip-restrictionplugin with network ranges. - Using Fixed Window without considering window edge: load at window boundaries doubles. Use Sliding Window or Leaky Bucket.
- Not setting Retry-After header: clients don't know when to retry. Always include
Retry-Afterin 429 responses. - Redis without replication: if Redis goes down, limits reset. Use Redis cluster or configure a replica.
Implementation and Pricing
- Audit of current API architecture and selection of optimal Gateway.
- Designing policies: global, per service, per user.
- Setting up Redis for synchronization (if needed).
- Integration with monitoring (Prometheus, Grafana) for limit visualization.
- Documentation of the rate limiting scheme for the development team.
- Engineer training (2-3 hour workshop).
- SLA guarantee of 99.9% during operation.
Basic multi-level throttling (by IP, consumer, service) with Redis takes 2 to 5 working days depending on integration complexity. Basic setup starts at $500, including Redis sync and default policies. Dynamic user-specific limits and custom monitoring range from $2,000 to $5,000. We calculate exact cost after auditing your architecture. With 5+ years of experience and over 50 API Gateway projects completed, we guarantee a reliable solution.
Contact us for a free consultation — we'll select the optimal request throttling strategy for your load. Get backend protection starting from $500.
Note: Sliding Window algorithm provides highest accuracy for critical APIs.







