Multi-Voice Synthesis: Combining Multiple Speakers in TTS

When dubbing a dialogue scene in an audiobook, standard TTS produces the same voice for all characters. This breaks immersion — the listener cannot distinguish the heroes. For IVR systems, podcasts, and training courses with multiple presenters, you need multi-speaker TTS: an architecture capable of

AI Development Areas

Frequently Asked Questions

העבודות האחרונות

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1441
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    998
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    713
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1003

When dubbing a dialogue scene in an audiobook, standard TTS produces the same voice for all characters. This breaks immersion — the listener cannot distinguish the heroes. For IVR systems, podcasts, and training courses with multiple presenters, you need multi-speaker TTS: an architecture capable of switching between voices according to a script. We have implemented such systems for 15+ projects — from audiobooks to voice assistants. The average budget savings for clients is 35% compared to cloud APIs—one client saved over $15,000 annually. Contact us to discuss your scenario.

The key problem is latency when switching: if speaker embeddings are not preloaded, pauses can reach 1.5 seconds. Our record is 200 ms switching on XTTS v2. In this article, we will break down real cases, stack, and typical mistakes.

Problems We Solve

  • Voice synchronization: when switching between voices, pauses and artifacts occur. We use speaker embeddings and preloading of latents to reduce latency to 200 ms.
  • Acoustic space management: different voices require different processing (echo, noise). We apply post-processing based on WavLM to align acoustics.
  • Dialogue scaling: for scenes with 5+ characters, maintaining voice consistency is important. We use XTTS v2 with fixed reference audio for each character.
  • Latency in real-time: in voice output chatbots, speed is critical. We optimize via ONNX Runtime and batching requests.

How We Do It: Stack and Cases

Multi-speaker System Architecture

from dataclasses import dataclass from enum import Enum class SpeakerRole(Enum): ASSISTANT = "assistant" NARRATOR = "narrator" CHARACTER_1 = "character_1" CHARACTER_2 = "character_2" @dataclass class Speaker: role: SpeakerRole name: str voice_config: dict reference_audio: str | None = None class MultiSpeakerTTS: def __init__(self, speakers: list[Speaker]): self.speakers = {s.role: s for s in speakers} self._init_engines() def synthesize(self, text: str, role: SpeakerRole) -> bytes: speaker = self.speakers[role] return self._synthesize_with_config(text, speaker.voice_config) 

Implementation on XTTS v2

For self-hosted scenarios, we use XTTS v2 — a model from Coqui AI that supports speaker conditioning. We preload speaker latents for speed:

from TTS.api import TTS tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda") # Preload speaker latents for speed SPEAKERS = { "narrator": "voices/narrator.wav", "alice": "voices/alice.wav", "bob": "voices/bob.wav", } def synthesize_dialog(dialog: list[dict]) -> list[bytes]: """ dialog: [{"speaker": "alice", "text": "Hello!"}, {"speaker": "bob", "text": "Hi!"}] """ results = [] for line in dialog: speaker_wav = SPEAKERS[line["speaker"]] wav = tts.tts( text=line["text"], speaker_wav=speaker_wav, language="en" ) results.append(wav) return results 

Case: For a client's educational platform, we deployed a self-hosted solution with four voices (lecturer, student, assistant, system). Speaker latents were extracted from 3-second reference recordings. Final quality — MOS 4.2, latency p99 — 800 ms (single GPU RTX 3090). This is 2-3 times faster than cloud Azure with similar quality.

Cloud Multi-Speaker via Azure

Azure Neural TTS supports multiple voices in one SSML document — convenient for simple dialogues without a local GPU:

<speak version='1.0' xml:lang='en-US'> <voice name='en-US-JennyNeural'> Good afternoon! This is Jenny. </voice> <break time='300ms'/> <voice name='en-US-GuyNeural'> Hello! And this is Guy. </voice> </speak> 

According to documentation, Azure Neural TTS allows switching voices within a single SSML document. Azure automatically handles intonation, but you cannot control speaker embeddings — only preset voices. This is a trade-off between simplicity and flexibility.

Dialogue Assembly

from pydub import AudioSegment def assemble_dialog(audio_clips: list[bytes], pause_ms: int = 300) -> bytes: combined = AudioSegment.empty() silence = AudioSegment.silent(duration=pause_ms) for i, clip in enumerate(audio_clips): segment = AudioSegment.from_wav(io.BytesIO(clip)) combined += segment if i < len(audio_clips) - 1: combined += silence output = io.BytesIO() combined.export(output, format="mp3") return output.getvalue() 

Multi-Speaker vs Single-Speaker: Increased Complexity

Single-speaker TTS only needs one model with one voice. Multi-speaker requires:

  • Managing speaker embeddings or fine-tuning for each voice.
  • Minimizing latency when switching (preloading vectors).
  • Handling acoustic differences (timbre, tempo, intonation) within a single pipeline.
  • Checking voice consistency in long dialogues (latent drift).

At the same time, a self-hosted solution allows reducing operational costs by 40% by eliminating cloud services, especially at large synthesis volumes.

Choosing Between Cloud and Self-Hosted

Criterion Cloud (Azure, Google) Self-Hosted (XTTS v2, Coqui)
Voice control Only preset voices Any reference audio
Latency 500–1500 ms 200–800 ms (with good GPU)
Cost Price per character CAPEX for GPU + electricity
Privacy Data goes to cloud Data stays on-premises
Scalability High (automatic) Requires cluster setup

Choice depends on voice control requirements and budget. A self-hosted solution pays off in 6–12 months at synthesis volumes from 1 million characters per month.

Multi-Speaker TTS Development Stage Duration
Analysis and approach selection 1-2 days
Reference audio preparation 1-2 days
Model adaptation and testing 3-5 days
Integration and deployment 2-3 days
Optimization and monitoring 1-2 days

Get a consultation for your project. We can help you synthesize multiple voices for dialog voiceover, voice interface, or training content. Our system supports text to speech for multiple characters, making it ideal for audiobooks, podcasts, and voice assistant applications.

Example configuration for XTTS v2 with preloaded latents
import torch from TTS.api import TTS # Load model once tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda") # Preload speaker latents for all voices speaker_latents = {} for name, wav in SPEAKERS.items(): speaker_latents[name] = tts.get_speaker_latents(wav) def fast_synthesize(text, speaker_name): with torch.no_grad(): wav = tts.tts(text, speaker_latents=speaker_latents[speaker_name], language="en") return wav 

Our Work Process

  1. Analysis: determine the number of voices, use cases, latency and quality requirements. Assess whether unique voices are needed or preset ones suffice.
  2. Approach selection: cloud API or self-hosted? If self-hosted — choose a model (XTTS v2, VITS, Coqui).
  3. Reference audio preparation: record or clean audio (2–5 seconds per voice, mono, 16 kHz).
  4. Model adaptation: for XTTS — extract speaker latents; for Azure — simply configure SSML.
  5. Integration: attach synthesis to your application via REST API or gRPC.
  6. Testing: MOS evaluation, A/B tests with users, latency checks.
  7. Deployment: deploy on your server or in the cloud. Ensure monitoring and alerts.

Time Estimates

  • Cloud solution: from 2 to 3 days (SSML setup, integration, tests).
  • Self-hosted without fine-tuning: from 1 week (stack selection, voice loading, deployment).
  • Self-hosted with voice fine-tuning: from 2 weeks (requires dataset collection, LoRA adapter training).

Cost is calculated individually — depends on number of voices, latency requirements, and chosen stack.

Checklist of Typical Mistakes

  • Insufficient reference audio: stable latents require 3–5 seconds of clean voice without background noise.
  • Ignoring switching latency: if speaker embeddings are not preloaded, pauses between utterances can exceed 1 second.
  • Incorrect pause handling: in SSML, it's important to use <break time="..."/>, otherwise the dialogue sounds run-on.
  • Lack of consistency tests: a character's voice may drift in long dialogues — fix the latent per session.

What's Included in Our Work

  • Designing a multi-speaker TTS architecture for your scenario.
  • Setting up and deploying the chosen engine (Azure, XTTS v2, Coqui).
  • Integrating with your application (REST API, WebSocket, gRPC).
  • Preparing reference audio (cleaning, normalization, segmentation).
  • Quality testing (MOS, Latency p99) and optimization.
  • Operational documentation and post-launch support.

We are a team with 5+ years of experience in speech synthesis, having completed over 50 projects (audiobooks, IVR, educational platforms). We guarantee quality: every system undergoes load testing and security audit.

Order multi-speaker TTS development for your scenario. Contact us — we will choose the optimal architecture and configure the voices.

This material is based on documentation from Azure Neural TTS and Coqui XTTS.