Speech Synthesis: VITS and XTTS for Custom Voice
A custom TTS model gives you full control over voice, language, and style — without dependence on external APIs and recurring costs. It's relevant for creating a unique brand voice, synthesis in rare languages/dialects, and edge deployment without internet. We have trained over 15 models for clients in retail, media, and voice assistants — from short advertising jingles to full-fledged dialog systems. We guarantee achieving a MOS of at least 4.0.
Why Train a Custom TTS Model?
Ready-made cloud TTS (Google, Yandex, Amazon) impose limitations: a fixed set of voices, cost per request, internet dependency, and latency. A custom model solves these problems: you get an exclusive voice that works offline, with control over emotional tone and pace. For example, one of our clients (a delivery aggregator) saved $3,000 per month by switching from a paid API to their model trained on 8 hours of a voice actor's speech. — Client from retail Savings on API calls can reach $50,000 per year for large projects.
How to Choose the TTS Architecture?
| Model | Type | Training Data | Quality (MOS) | Inference Speed |
|---|---|---|---|---|
| VITS | End-to-end (text→audio) | 2–5 h | 4.2/5 | Realtime ×30 on GPU |
| XTTS v2 (Coqui) | Zero-shot + fine-tune | 3–6 min (few-shot) | 4.4/5 | Realtime ×10 on GPU |
| YourTTS | Multilingual VITS | 1–3 h | 4.0/5 | Realtime ×20 |
| MATCHA-TTS | Flow-matching | 2–4 h | 4.3/5 | Realtime ×50 |
| StyleTTS2 | Style-based | 1–2 h | 4.5/5 | Realtime ×15 |
For most tasks: XTTS v2 for quick startup with minimal data, VITS for full training with a clean dataset. When fine-tuned on 6 minutes of audio, XTTS v2 achieves quality comparable to full VITS on 10 hours – confirmed by our MOS measurements. With years of experience, we select the architecture best suited for your task.
Dataset Preparation
Minimum requirements for quality results:
Format: 22050 Hz, 16-bit, mono WAV Recording length: 2–15 seconds each Minimum: 1000 recordings (≈2 hours) for intelligible TTS Recommended: 3000–5000 recordings (≈8–12 hours) for high quality Text script: UTF-8, one utterance per line Dataset structure:
dataset/ ├── wavs/ │ ├── speaker_001.wav │ ├── speaker_002.wav │ └── ... ├── metadata.csv # filename|transcription └── metadata_val.csv # 10% for validation Preprocessing and normalization:
import librosa import soundfile as sf import numpy as np from pathlib import Path def preprocess_audio_for_tts( input_dir: str, output_dir: str, target_sr: int = 22050 ) -> dict: stats = {"processed": 0, "skipped": 0, "errors": []} Path(output_dir).mkdir(parents=True, exist_ok=True) for wav_path in Path(input_dir).glob("*.wav"): audio, sr = librosa.load(str(wav_path), sr=target_sr, mono=True) # Trim silence audio_trimmed, _ = librosa.effects.trim(audio, top_db=20) # Check length duration = len(audio_trimmed) / target_sr if duration < 1.5 or duration > 15.0: stats["skipped"] += 1 continue # Normalize amplitude audio_normalized = audio_trimmed / (np.max(np.abs(audio_trimmed)) + 1e-8) audio_normalized *= 0.9 # peak -0.9 dB output_path = Path(output_dir) / wav_path.name sf.write(str(output_path), audio_normalized, target_sr, subtype="PCM_16") stats["processed"] += 1 return stats VITS Training
config.json configuration for VITS (Coqui TTS):
{ "model": "vits", "run_name": "my_tts_model", "epochs": 1000, "batch_size": 32, "eval_batch_size": 16, "num_loader_workers": 4, "audio": { "sample_rate": 22050, "win_length": 1024, "hop_length": 256, "num_mels": 80, "mel_fmin": 0, "mel_fmax": null }, "datasets": [{ "name": "my_dataset", "path": "dataset/", "meta_file_train": "metadata.csv", "meta_file_val": "metadata_val.csv" }] } Launch training:
from TTS.bin.train_tts import main as train_tts from TTS.config.shared_configs import BaseDatasetConfig from TTS.tts.configs.vits_config import VitsConfig from TTS.tts.datasets import load_tts_samples from TTS.tts.models.vits import Vits, VitsAudioConfig from TTS.trainer import Trainer, TrainerArgs audio_config = VitsAudioConfig( sample_rate=22050, win_length=1024, hop_length=256, num_mels=80, mel_fmin=0, mel_fmax=None ) config = VitsConfig( audio=audio_config, run_name="brand_voice_v1", batch_size=32, eval_batch_size=16, epochs=1000, text_cleaner="phoneme_cleaners", use_phonemes=True, phoneme_language="ru-ru", phoneme_cache_path="phoneme_cache/", output_path="checkpoints/", datasets=[BaseDatasetConfig( formatter="ljspeech", meta_file_train="metadata.csv", path="dataset/" )] ) train_samples, eval_samples = load_tts_samples( config.datasets, eval_split=True, eval_split_size=0.1 ) model = Vits(config, ap=None, tokenizer=None, speaker_manager=None) trainer = Trainer( TrainerArgs(), config, output_path="checkpoints/", model=model, train_samples=train_samples, eval_samples=eval_samples ) trainer.fit() XTTS v2 Fine-Tuning (Few-Shot)
XTTS v2 supports fine-tuning with 3–6 minutes of audio:
from TTS.demos.xtts_ft_demo.xtts_demo import train_gpt # Dataset: at least 100 recordings, each 2–6 seconds long train_gpt( language="ru", num_epochs=6, batch_size=4, grad_acumm=1, train_csv="dataset/metadata_train.csv", eval_csv="dataset/metadata_eval.csv", output_path="xtts_ft_checkpoints/" ) Inference with custom voice after fine-tuning:
from TTS.api import TTS tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2") tts.tts_to_file( text="Welcome to our company.", speaker_wav="reference_voice.wav", # 3–10 sec reference audio language="ru", file_path="output.wav", model_path="xtts_ft_checkpoints/best_model.pth" ) Our Approach to TTS Model Training
Our process includes five stages:
- Analytics: determine the target audience for the voice, requirements for language, emotions, speed. Select architecture (VITS, XTTS, YourTTS) for the task.
- Dataset collection and preparation: record voice actor in studio or clean existing recordings. Remove noise, silence, normalize. Transcribe texts.
- Model training: run on GPU cluster, monitor metrics (train/val loss, KL loss, grad_norm). Use early stopping and checkpoints.
- Quality evaluation: listen to synthesis every 100 epochs, compare to reference. Achieve MOS of at least 4.0.
- Deployment and integration: convert to ONNX for edge or deploy as gRPC/REST API. Provide documentation and support.
Training Metric Monitoring
Key metrics in tensorboard: - loss/train_loss: should decrease monotonically - loss/val_loss: parallel to train, no divergence - loss/kl_loss: KL divergence of latent space - loss/disc_loss: discriminator (GAN component) - grad_norm: should be < 10, otherwise gradient explosionTraining Infrastructure
| GPU | Training Time (1000 epochs, VITS) | VRAM |
|---|---|---|
| RTX 3090 (24 GB) | ~12 hours | 18 GB |
| A100 (40 GB) | ~5 hours | 22 GB |
| 2× A10G | ~3 hours | 2×24 GB |
| CPU (no GPU) | Not recommended | — |
Cloud options: RunPod ($1.5/h for A100), Lambda Cloud ($1.1/h), Vast.ai (~$0.5–0.8/h for A100).
Post-Training: Model Deployment
# ONNX export for edge deployment from TTS.utils.synthesizer import Synthesizer synthesizer = Synthesizer( tts_checkpoint="checkpoints/best_model.pth", tts_config_path="checkpoints/config.json" ) # Inference wav = synthesizer.tts("Test phrase for synthesis") synthesizer.save_wav(wav, "test_output.wav") What's Included in the Work
- Trained model (VITS, XTTS, or YourTTS) with achieved quality of at least MOS 4.0.
- Clean dataset with transcriptions and preprocessing scripts.
- Configuration files and code to reproduce training.
- Inference scripts for local and server use.
- API wrapper (FastAPI/gRPC) for integration into your service.
- Documentation for setup and operation.
- Support for 2 weeks after delivery.
Timeline: dataset preparation (recording + transcription) — 2–4 weeks. VITS model training — 1–2 weeks (GPU). Integration into production service with API — 1 week. Full cycle from scratch to brand voice — 4–6 weeks. Get a consultation from our AI engineer — we will select the optimal architecture and calculate precise deadlines. Order TTS model training for your project.







