How to Integrate Asterisk with AI?
"We have Asterisk 18, want to add a voice assistant, but ready-made solutions bring a cloud subscription and don't let us control latency." This is a typical pain point. The default FreePBX capabilities — DTMF menus and playing pre-recorded files. AI enables a dialogue: user speaks, system understands, LLM makes a decision, TTS responds. But integration is not "install a module". It's about choosing the stack (STT: Whisper or Deepgram? LLM: LLaMA 3.1 70B or GPT-4o-mini? TTS: Piper or ElevenLabs?), about latency p99 and dialog convergence. Here's how we build such projects on-premise or hybrid.
What Problems We Solve
Problem 1: Recognition quality in noisy environments. Call centers, open-plan offices, street noise — standard STT models give WER > 30%. We use Whisper large-v3 with fine-tuning on your audio corpus (300 hours of recordings reduce WER to 10%) or Deepgram Nova-2 with noise suppression on input.
Problem 2: End-to-end latency > 2 seconds. Users won't wait longer than 1.5 s. We optimize the pipeline: streaming STT with partial results, early-exit LLM, TTS generating first packet in 150 ms. On the Whisper + vLLM + Piper stack we get p99 = 1.1 s.
Problem 3: Fault tolerance. An AI service crash should not break the call. We design a fallback: on AI node timeout, the dialplan switches to standard IVR. The ARI application monitors each component's health and on error returns control to Asterisk.
How We Do It: Stack and Configs
The choice of interface — AGI or ARI — depends on real-time requirements and dialog complexity. Comparison below.
| Characteristic | AGI (Asterisk Gateway Interface) | ARI (Asterisk REST Interface) |
|---|---|---|
| Mechanism | Script execution on event | WebSocket + REST, full duplex |
| Latency | +3-5 ms (syscall) | +1-2 ms (direct channel control) |
| Streaming support | Only batch recording | True audio streaming |
| Complexity | Low (Python script) | High (async application) |
| Fault tolerance | Built-in retry | Manual reconnect required |
| When to choose | Simple IVR, voice input/output | Real-time dialog, barge-in, topic switching |
Example AGI script (Python):
# agi_bot.py import sys from asterisk.agi import AGI agi = AGI() agi.answer() agi.set_variable("CHANNEL(audioreadformat)", "slin16") # Record audio from user agi.record_file( "/tmp/user_audio", "wav", "#", # stop key 3000, # timeout ms 0, # offset True, # beep 3 # silence threshold ) # STT + LLM + TTS in separate service import requests with open("/tmp/user_audio.wav", "rb") as f: stt_response = requests.post("http://ai-service/stt", files={"audio": f}) transcript = stt_response.json()["text"] llm_response = requests.post("http://ai-service/chat", json={"text": transcript}) response_text = llm_response.json()["response"] tts_response = requests.post("http://ai-service/tts", json={"text": response_text}) with open("/tmp/response.wav", "wb") as f: f.write(tts_response.content) agi.stream_file("/tmp/response") ARI example (asyncio):
import asyncio import aiohttp from ari_client import ARIClient async def handle_stasis(channel_id: str, ari: ARIClient): """Incoming call handler via ARI""" await ari.answer(channel_id) # Create audio snapshot of the channel await ari.channel.record( channelId=channel_id, name=f"call_{channel_id}", format="wav", terminateOn="silence", maxSilenceSeconds=2 ) Why Choose AGI Over ARI?
If your scenario is simple voice I/O (e.g., "say your name, the system will find the client"), AGI gives minimal development time. Writing the script takes 1-2 days, debugging another week. ARI requires async architecture and session management but allows handling barge-ins and partial recognition results. For contact centers with natural dialog, ARI is the only choice.
How to Achieve Latency < 1 Second?
We build the pipeline:
- STT — Whisper medium.en + ModelScope (INT8 quant) on GPU T4 — first token in 250 ms.
- LLM — vLLM with LLaMA 3.1 8B, prefill at 20 tokens, early stop generation — 400 ms. Compare: vLLM is 2.5x faster than standard Hugging Face pipeline under load.
- TTS — Piper (VITS) with first packet synthesis in 100 ms. Result: 750 ms end-to-end. Under load of 10 simultaneous calls, latency p99 stays at 1.2 s.
According to Asterisk documentation, for streaming audio it is recommended to use ARI, as AGI does not support full-duplex transmission.
LLM Comparison for Voice Scenarios
| Model | Latency (first token) | Quality (multitask) | Cost (per 1M tokens) |
|---|---|---|---|
| LLaMA 3.1 70B (on-premise) | 400–500 ms | High | ~$0.5 (electricity) |
| GPT-4o-mini (cloud) | 300–400 ms | Very high | $0.15/$0.60 (input/output) |
| Mistral 7B (on-premise) | 200–300 ms | Medium | ~$0.1 |
Triton Inference Server configuration for LLM
name: "llama_ensemble" backend: "ensemble" input [ { name: "text_input" data_type: TYPE_STRING dims: [ -1 ] } ] output [ { name: "text_output" data_type: TYPE_STRING dims: [ -1 ] } ] Process
- Analysis (1–3 days): Audit current Asterisk/FreePBX config, gather scenario requirements, measure average call duration.
- Design (2–5 days): Choose stack (models, frameworks), design integration, prototype dialog.
- PoC (1–2 weeks): Deploy on test PBX with 1 inbound number. Demo to client.
- Production (2–4 weeks): Deploy on target servers, set up monitoring (Prometheus + Grafana), integrate with CRM.
- Support (2 weeks after launch): Train operators, tweak dialog script, optimize latency.
Timelines and Cost
- Basic AGI integration (one scenario, Whisper + LLaMA + Piper): from 2 to 3 weeks.
- Full ARI implementation (real-time dialog, multilingual, fault tolerance): from 1 to 1.5 months. Cost is calculated individually after infrastructure audit. Get a consultation — we estimate your project in 1 day.
What's Included
- Deployment and operation documentation.
- Repository with configs and scripts (Git).
- Operator training (2 sessions of 1 hour).
- Two weeks of post-launch support.
- 3-month warranty on code.
We are a team with 5 years of experience in AI integrations for telephone systems. We have 12+ projects for contact centers up to 50 lines. Certified Asterisk and FreePBX engineers (Sangoma distribution). Contact us — we will send a technical and commercial proposal.







