Integrating OpenAI Realtime API for Voice AI

Integrating OpenAI Realtime API for Voice AI

AI Development Areas

Frequently Asked Questions

העבודות האחרונות

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1441
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    998
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    713
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1003

Integrating OpenAI Realtime API for Voice AI

The standard voice assistant pipeline consists of three sequential stages: speech-to-text (STT), response generation (LLM), and text-to-speech (TTS). Each stage adds latency, and the total RTT often exceeds 2–4 seconds. This severely disrupts the natural flow of conversation. OpenAI Realtime API solves this by providing a single WebSocket connection for direct voice-to-voice transmission with 200–500 ms latency. No intermediate transcription: audio goes in, audio comes out. For more details, see official documentation.

Our engineers have 5+ years of experience in voice agent development and have successfully delivered over 50 projects. We guarantee stable operation under load.

In one telemarketing project, we replaced a three-tier architecture with the API — RTT dropped from 3.2 s to 380 ms. This boosted dialogue conversion by 25% due to more natural interactions, and call center infrastructure costs were reduced by up to 50% (average monthly savings of $1,200).

How OpenAI Realtime API Processes Voice

The API opens a single WebSocket connection that simultaneously transmits audio and text messages. The client sends audio streams in PCM16 chunks; the server detects speech activity, recognizes commands (via Whisper), and generates a response. WebSocket is a protocol available in any modern programming language.

import asyncio import json import websockets import base64 async def voice_assistant(): url = "wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview" headers = { "Authorization": f"Bearer {OPENAI_API_KEY}", "OpenAI-Beta": "realtime=v1" } async with websockets.connect(url, extra_headers=headers) as ws: # Initialize session await ws.send(json.dumps({ "type": "session.update", "session": { "modalities": ["text", "audio"], "instructions": "You are a helpful voice assistant. Respond in Russian, be concise.", "voice": "alloy", "input_audio_format": "pcm16", "output_audio_format": "pcm16", "input_audio_transcription": {"model": "whisper-1"}, "turn_detection": { "type": "server_vad", "threshold": 0.5, "prefix_padding_ms": 300, "silence_duration_ms": 700 } } })) async def send_audio(audio_stream): async for chunk in audio_stream: encoded = base64.b64encode(chunk).decode() await ws.send(json.dumps({ "type": "input_audio_buffer.append", "audio": encoded })) async def receive_responses(): audio_buffer = bytearray() async for message in ws: event = json.loads(message) if event["type"] == "response.audio.delta": audio_data = base64.b64decode(event["delta"]) audio_buffer.extend(audio_data) # Play chunks as they arrive elif event["type"] == "response.audio.done": pass elif event["type"] == "conversation.item.input_audio_transcription.completed": print(f"User: {event['transcript']}") await asyncio.gather(send_audio(get_microphone_stream()), receive_responses()) 

Why OpenAI Realtime API Is Faster than Traditional Pipeline

A typical STT+LLM+TTS stack gives an RTT of 2–4 seconds. The real-time API eliminates inter-stage delays through a direct audio channel. In our projects, we achieved p99 latency of 450 ms — nearly imperceptible to the user. Compared to classical solutions, speed increases 4–8 times.

Parameter Realtime API STT+LLM+TTS
Latency (RTT) 200–500 ms 2–4 s
Number of connections 1 WebSocket 3 HTTP/gRPC
Interruption Built-in Needs workaround
Function calling Voice-driven Text-only
Voice emotions 6 built-in voices TTS-dependent

Key Features of OpenAI Realtime API

User interruption. Server-side VAD automatically detects when the user starts speaking and stops synthesis. This is critical for natural dialogue: the assistant doesn't keep talking when interrupted. Configurable parameters: threshold (sensitivity) and silence_duration (pause before processing).

Scenario Threshold Silence Duration (ms) Prefix Padding (ms)
Quiet office 0.3 500 200
Noisy call center 0.7 800 400
Smart speaker 0.5 700 300

Function calling in voice mode. The API calls custom functions directly from the voice stream. For example, the user says "Show order status #123" and the assistant executes a real CRM query.

tools = [{ "type": "function", "name": "get_order_status", "description": "Get order status by order number", "parameters": { "type": "object", "properties": { "order_id": {"type": "string", "description": "Order number"} }, "required": ["order_id"] } }] await ws.send(json.dumps({ "type": "session.update", "session": {"tools": tools, "tool_choice": "auto"} })) 
VAD Configuration Details

VAD parameters are tuned to the room acoustics: the threshold coefficient determines sensitivity to speech volume; silence_duration sets the pause to mark the end of a phrase. We recommend starting with the values from the table above and adjusting through testing.

Common Integration Mistakes

  • Incorrect VAD settings: Too low a threshold triggers on background noise; too high makes the assistant miss quiet speech. We tune parameters to your environment.
  • Lack of reconnection handling: WebSocket can drop; without auto-reconnect the assistant goes silent. Our integration includes exponential backoff reconnection.
  • Ignoring latency in function calling: If your API responds slowly, the voice agent will hang. We optimize the call chain.

Scope of Integration Work

  • Current scheme analysis — evaluate latency, audit existing STT/TTS pipeline.
  • WebSocket integration — configure connection, handle reconnection, audio compression.
  • VAD configuration — tune threshold for your noise profile.
  • Function calling implementation — connect to your CRM, API, or database.
  • Team training — handover code and documentation.
  • Post-launch support — latency monitoring, error handling, model updates.

OpenAI Realtime API Implementation Process

  1. Analysis — study your scenario and load.
  2. Design — select voice, VAD parameters, tools.
  3. Implementation — write the integration layer.
  4. Testing — measure latency in real conditions.
  5. Deployment — deploy on your infrastructure or cloud.

Timelines: basic integration — 2–3 days; production solution with business logic — 1–2 weeks. Cost is estimated individually based on complexity and scope, with integration projects typically starting at $2,500. Typical savings are $1,200 per month, reducing overall costs significantly.

What's Included in the Integration

  • Documentation of the integration architecture and setup guide.
  • Client-side WebSocket code ready for deployment.
  • One training session for your team (up to 2 hours).
  • Post-launch support for 30 days including bug fixes and latency monitoring.

Contact us for a consultation. Get a free assessment of your project — we'll help you pick the optimal configuration and launch your voice assistant within a week. Order a pilot project to test the solution on your data.