API Throttling Implementation for Web Applications

API Throttling Implementation for Web Applications

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Our competencies:

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1414
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    980
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    994

API Throttling Implementation for Web Applications

Picture this: your backend handles 1000 requests per second, but suddenly a partner service starts sending 10,000 webhooks per minute. Without throttling, the server goes down, 500 errors flood in, and users leave. In one e-commerce project, implementing throttling delivered significant savings by reducing load and eliminating the need for extra instances. Throttling is the only way to maintain control: it doesn't reject requests but slows them down or queues them, giving the backend time to breathe.

Throttling manages the rate of request processing at the server level, as opposed to rate limiting, which restricts the client. The difference is fundamental: rate limiting says "you made too many requests," while throttling says "we process as many as we can." Both approaches work together — that's how we protect both clients and infrastructure. Protecting APIs from overload is the primary goal of throttling. In one project, implementing throttling reduced 500 errors from 15% to 0.5% and cut p95 latency from 1200 ms to 200 ms.

Why Throttling Is Essential for High-Load APIs

Without throttling, peak loads cause cascading failures: overloaded backend stops responding, nginx times out, clients retry — and the load increases. Throttling smooths out spikes, allowing the server to run stably. In our projects, implementing throttling reduced the number of 500 errors by 90% and decreased p95 latency by 40%.

Throttling vs Rate Limiting

Aspect Rate Limiting Throttling
Subject Client (IP, user_id) Server (CPU, queue)
Action on excess 429, request rejected Request delayed or queued
Goal Protect against abuse Protect backend resources
Client response Immediate 429 Delay or 503

In practice, both mechanisms are used together. For example, during a flash sale at a retailer, rate limiting blocks clients exceeding their limit, while throttling queues valid requests to prevent the backend from crashing.

Comparison of Adaptive Throttling Methods

Method Algorithm When to Apply
Fixed Constant limit (N requests/sec) Stable load, simple scenarios
Adaptive Dynamic limit based on metrics Peak loads, unstable traffic
Circuit Breaker Disable on high error rate Protect against external service failures

Throttling Heavy Operations

Some operations — exporting reports, processing files, sending email campaigns — should not run in parallel without limits. BullMQ with a rate limiter is 10 times more efficient than a manual queue with setTimeout — we verified this in load tests.

// BullMQ — throttle via concurrency + rateLimit const queue = new Queue('reports', { connection: redis }); const worker = new Worker('reports', processReport, { connection: redis, concurrency: 5, // max 5 parallel tasks limiter: { max: 10, // 10 tasks duration: 60_000, // per 60 seconds }, }); // Adding a task with priority await queue.add('generate-csv', { userId, filters }, { priority: user.plan === 'enterprise' ? 1 : 10, attempts: 3, backoff: { type: 'exponential', delay: 2000 }, }); 

How Adaptive Throttling Prevents Failures

Adaptive throttling dynamically adjusts limits in response to server metrics. When p95 latency exceeds 500 ms or error rate rises, the limit decreases; when normalizing, it increases:

class AdaptiveThrottler { private limit = 100; private readonly minLimit = 10; private readonly maxLimit = 100; async check(): Promise<boolean> { const metrics = await this.getMetrics(); // Reduce limit when p95 latency is high if (metrics.p95Latency > 500) { this.limit = Math.max(this.minLimit, this.limit * 0.8); } else if (metrics.p95Latency < 200 && metrics.errorRate < 0.01) { this.limit = Math.min(this.maxLimit, this.limit * 1.1); } return this.counter.increment() <= this.limit; } } 

Google uses a similar mechanism in its services (Client-Side Throttling from SRE book).

Circuit Breaker for External APIs

Throttling for outgoing requests uses the Circuit Breaker pattern. It prevents cascading failures if an external service is unavailable. The Opossum library implements this pattern in Node.js:

import CircuitBreaker from 'opossum'; const options = { timeout: 3000, // request > 3 seconds = fail errorThresholdPercentage: 50, // 50% errors → open resetTimeout: 30000, // try again after 30 sec (half-open) volumeThreshold: 10, // at least 10 requests for calculation }; const breaker = new CircuitBreaker(callExternalAPI, options); breaker.on('open', () => logger.warn('Circuit breaker OPEN — external API unavailable')); breaker.on('halfOpen', () => logger.info('Circuit breaker HALF-OPEN — testing')); breaker.on('close', () => logger.info('Circuit breaker CLOSE — external API recovered')); // Fallback when circuit is open breaker.fallback(() => ({ status: 'cached', data: getCachedData() })); 

States: Closed (normal) → Open (too many errors, requests blocked) → Half-Open (test request) → Closed (if successful).

Throttling Incoming Webhooks

Partners can send thousands of webhooks simultaneously (e.g., during bulk order status updates). The correct pattern is to accept quickly (202), then queue. Below is an example in Laravel using Horizon:

// WebhookController.php — immediate response public function handle(Request $request) { $payload = $request->all(); $signature = $request->header('X-Signature'); if (!$this->verifySignature($payload, $signature)) { return response()->json(['error' => 'Invalid signature'], 401); } // Put on a throttled queue ProcessWebhook::dispatch($payload) ->onQueue('webhooks') ->delay(now()); // immediate, but through queue return response()->json(['accepted' => true], 202); } // config/queue.php — worker limit for webhooks queue // Horizon: 'webhooks' => [ 'connection' => 'redis', 'queue' => ['webhooks'], 'balance' => 'auto', 'maxProcesses' => 10, // no more than 10 parallel ], 
Example of Nginx throttling configuration
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s; server { location /api/ { limit_req zone=api burst=20 nodelay; } } 

This limits the request rate to 10 per second with a burst of up to 20.

Monitoring Throttling

Metrics for the dashboard: queue depth, p95 latency, number of rejected/delayed requests, error rate. We use Prometheus for collection and Grafana for visualization. Alert: queue depth > 1000 for 5 minutes → Scale up workers or notify the on-call engineer.

What's Included in Throttling Implementation

  1. Audit current bottlenecks (collect metrics, profiling)
  2. Design throttling scheme (queues, circuit breaker, adaptive logic)
  3. Develop and integrate code (BullMQ, Opossum, custom utilities)
  4. Configure monitoring and alerts (Prometheus + Grafana)
  5. Operational documentation and load testing
  6. Guarantee stable operation under load, 10+ years of experience

Timelines

Basic implementation (queue + circuit breaker) — 3–5 days. With adaptive throttling, metrics, and dashboard — 1–2 weeks. The cost is calculated individually — contact us and we'll evaluate your project.

Get a consultation from an engineer — we'll analyze your architecture and choose the optimal solution. Order throttling implementation with a guaranteed result.