API Throttling Implementation for Web Applications
Picture this: your backend handles 1000 requests per second, but suddenly a partner service starts sending 10,000 webhooks per minute. Without throttling, the server goes down, 500 errors flood in, and users leave. In one e-commerce project, implementing throttling delivered significant savings by reducing load and eliminating the need for extra instances. Throttling is the only way to maintain control: it doesn't reject requests but slows them down or queues them, giving the backend time to breathe.
Throttling manages the rate of request processing at the server level, as opposed to rate limiting, which restricts the client. The difference is fundamental: rate limiting says "you made too many requests," while throttling says "we process as many as we can." Both approaches work together — that's how we protect both clients and infrastructure. Protecting APIs from overload is the primary goal of throttling. In one project, implementing throttling reduced 500 errors from 15% to 0.5% and cut p95 latency from 1200 ms to 200 ms.
Why Throttling Is Essential for High-Load APIs
Without throttling, peak loads cause cascading failures: overloaded backend stops responding, nginx times out, clients retry — and the load increases. Throttling smooths out spikes, allowing the server to run stably. In our projects, implementing throttling reduced the number of 500 errors by 90% and decreased p95 latency by 40%.
Throttling vs Rate Limiting
| Aspect | Rate Limiting | Throttling |
|---|---|---|
| Subject | Client (IP, user_id) | Server (CPU, queue) |
| Action on excess | 429, request rejected | Request delayed or queued |
| Goal | Protect against abuse | Protect backend resources |
| Client response | Immediate 429 | Delay or 503 |
In practice, both mechanisms are used together. For example, during a flash sale at a retailer, rate limiting blocks clients exceeding their limit, while throttling queues valid requests to prevent the backend from crashing.
Comparison of Adaptive Throttling Methods
| Method | Algorithm | When to Apply |
|---|---|---|
| Fixed | Constant limit (N requests/sec) | Stable load, simple scenarios |
| Adaptive | Dynamic limit based on metrics | Peak loads, unstable traffic |
| Circuit Breaker | Disable on high error rate | Protect against external service failures |
Throttling Heavy Operations
Some operations — exporting reports, processing files, sending email campaigns — should not run in parallel without limits. BullMQ with a rate limiter is 10 times more efficient than a manual queue with setTimeout — we verified this in load tests.
// BullMQ — throttle via concurrency + rateLimit const queue = new Queue('reports', { connection: redis }); const worker = new Worker('reports', processReport, { connection: redis, concurrency: 5, // max 5 parallel tasks limiter: { max: 10, // 10 tasks duration: 60_000, // per 60 seconds }, }); // Adding a task with priority await queue.add('generate-csv', { userId, filters }, { priority: user.plan === 'enterprise' ? 1 : 10, attempts: 3, backoff: { type: 'exponential', delay: 2000 }, }); How Adaptive Throttling Prevents Failures
Adaptive throttling dynamically adjusts limits in response to server metrics. When p95 latency exceeds 500 ms or error rate rises, the limit decreases; when normalizing, it increases:
class AdaptiveThrottler { private limit = 100; private readonly minLimit = 10; private readonly maxLimit = 100; async check(): Promise<boolean> { const metrics = await this.getMetrics(); // Reduce limit when p95 latency is high if (metrics.p95Latency > 500) { this.limit = Math.max(this.minLimit, this.limit * 0.8); } else if (metrics.p95Latency < 200 && metrics.errorRate < 0.01) { this.limit = Math.min(this.maxLimit, this.limit * 1.1); } return this.counter.increment() <= this.limit; } } Google uses a similar mechanism in its services (Client-Side Throttling from SRE book).
Circuit Breaker for External APIs
Throttling for outgoing requests uses the Circuit Breaker pattern. It prevents cascading failures if an external service is unavailable. The Opossum library implements this pattern in Node.js:
import CircuitBreaker from 'opossum'; const options = { timeout: 3000, // request > 3 seconds = fail errorThresholdPercentage: 50, // 50% errors → open resetTimeout: 30000, // try again after 30 sec (half-open) volumeThreshold: 10, // at least 10 requests for calculation }; const breaker = new CircuitBreaker(callExternalAPI, options); breaker.on('open', () => logger.warn('Circuit breaker OPEN — external API unavailable')); breaker.on('halfOpen', () => logger.info('Circuit breaker HALF-OPEN — testing')); breaker.on('close', () => logger.info('Circuit breaker CLOSE — external API recovered')); // Fallback when circuit is open breaker.fallback(() => ({ status: 'cached', data: getCachedData() })); States: Closed (normal) → Open (too many errors, requests blocked) → Half-Open (test request) → Closed (if successful).
Throttling Incoming Webhooks
Partners can send thousands of webhooks simultaneously (e.g., during bulk order status updates). The correct pattern is to accept quickly (202), then queue. Below is an example in Laravel using Horizon:
// WebhookController.php — immediate response public function handle(Request $request) { $payload = $request->all(); $signature = $request->header('X-Signature'); if (!$this->verifySignature($payload, $signature)) { return response()->json(['error' => 'Invalid signature'], 401); } // Put on a throttled queue ProcessWebhook::dispatch($payload) ->onQueue('webhooks') ->delay(now()); // immediate, but through queue return response()->json(['accepted' => true], 202); } // config/queue.php — worker limit for webhooks queue // Horizon: 'webhooks' => [ 'connection' => 'redis', 'queue' => ['webhooks'], 'balance' => 'auto', 'maxProcesses' => 10, // no more than 10 parallel ], Example of Nginx throttling configuration
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s; server { location /api/ { limit_req zone=api burst=20 nodelay; } } This limits the request rate to 10 per second with a burst of up to 20.
Monitoring Throttling
Metrics for the dashboard: queue depth, p95 latency, number of rejected/delayed requests, error rate. We use Prometheus for collection and Grafana for visualization. Alert: queue depth > 1000 for 5 minutes → Scale up workers or notify the on-call engineer.
What's Included in Throttling Implementation
- Audit current bottlenecks (collect metrics, profiling)
- Design throttling scheme (queues, circuit breaker, adaptive logic)
- Develop and integrate code (BullMQ, Opossum, custom utilities)
- Configure monitoring and alerts (Prometheus + Grafana)
- Operational documentation and load testing
- Guarantee stable operation under load, 10+ years of experience
Timelines
Basic implementation (queue + circuit breaker) — 3–5 days. With adaptive throttling, metrics, and dashboard — 1–2 weeks. The cost is calculated individually — contact us and we'll evaluate your project.
Get a consultation from an engineer — we'll analyze your architecture and choose the optimal solution. Order throttling implementation with a guaranteed result.







