Research Demo Page

Zero-Shot Lombard Speech Synthesis
with Controllable Style Embeddings

Seymanur Akti1 • Alexander Waibel1, 2
1 Karlsruhe Institute of Technology (KIT) 2 Carnegie Mellon University (CMU)

Synthesizing Lombard-like speech via latent space manipulation without requiring Lombard-specific training data.

Abstract

The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable TTS (TTS) system capable of synthesizing Lombard-like speech in a zero-shot manner without requiring Lombard-specific training data. Our approach extends TTS architectures with a learned style embedding representation and analyzes the resulting latent space using principal component analysis (PCA) to identify directions associated with Lombard-related attributes. By manipulating these directions, we obtain interpretable control over vocal effort and articulation and generate speech at different Lombardness levels. Experimental results show that the proposed method preserves speaker identity and naturalness, improves intelligibility under noisy conditions, and generalizes to previously unseen speakers. These findings demonstrate that style-embedding manipulation provides an effective and scalable framework for controllable zero-shot Lombard speech synthesis.

Audio Samples (GT vs TTS)

Side-by-side audio comparison across clean, loud, and very Loud conditions.

Sample #1
Clean Clean
Ground Truth
TTS
Loud SNR 10
Ground Truth
TTS
Very Loud SNR 5
Ground Truth
TTS
Sample #2
Clean Clean
Ground Truth
TTS
Loud SNR 10
Ground Truth
TTS
Very Loud SNR 5
Ground Truth
TTS
Sample #3
Clean Clean
Ground Truth
TTS
Loud SNR 10
Ground Truth
TTS
Very Loud SNR 5
Ground Truth
TTS
Sample #4
Clean Clean
Ground Truth
TTS
Loud SNR 10
Ground Truth
TTS
Very Loud SNR 5
Ground Truth
TTS
Sample #5
Clean Clean
Ground Truth
TTS
Loud SNR 10
Ground Truth
TTS
Very Loud SNR 5
Ground Truth
TTS
Sample #6
Clean Clean
Ground Truth
TTS
Loud SNR 10
Ground Truth
TTS
Very Loud SNR 5
Ground Truth
TTS