AccentDrift: Real-time Streaming Accent Conversion via
Sparse Speech Tokenization
Abstract
Although recent speech AI advances have enabled real-time speech understanding and generation, real-time accent conversion (AC) for second-language (L2) speakers in real-world interactions remains underexplored. In this paper, we present AccentDrift, a real-time streaming AC system that preserves linguistic content and speaker identity while transforming speech into a target accent with minimal latency. From an information-theoretic perspective, we adopt sparse speech tokenization to extract linguistic information from continuous speech. An accent adapter injects accent style into sparse semantic tokens, followed by a timbre adapter that hierarchically generates speaker-specific acoustics. To support streaming, we adopt a cache-aware streaming architecture and parallel streams, achieving 520 ms latency. AccentDrift outperforms prior parallel AC models, and demonstrates the feasibility of generating native-like accents for L2 speakers.
Accent Conversion Comparison
We used the same input samples for L2 speaker (L2-ARCTIC) from the demo page of Streaming Baseline [28]
[28] Nguyen, Tuan-Nam, et al. "Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement." Proc. Interspeech 2025. 2025.
| Input (Arabic Accent) |
AccentDrift (Streaming, 0.5s) |
AccentDrift (Streaming, 0.9s) |
Streaming Baseline (Streaming, 0.8s) |
Vevo-Style (Non-Streaming) |
|---|---|---|---|---|
| Input (Chinese Accent) |
AccentDrift (Streaming, 0.5s) |
AccentDrift (Streaming, 0.9s) |
Streaming Baseline (Streaming, 0.8s) |
Vevo-Style (Non-Streaming) |
| Input (Vietnamese Accent) |
AccentDrift (Streaming, 0.5s) |
AccentDrift (Streaming, 0.9s) |
Streaming Baseline (Streaming, 0.8s) |
Vevo-Style (Non-Streaming) |
| Input (Indian Accent) |
AccentDrift (Streaming, 0.5s) |
AccentDrift (Streaming, 0.9s) |
Streaming Baseline (Streaming, 0.8s) |
Vevo-Style (Non-Streaming) |
| Input (Korean Accent) |
AccentDrift (Streaming, 0.5s) |
AccentDrift (Streaming, 0.9s) |
Streaming Baseline (Streaming, 0.8s) |
Vevo-Style (Non-Streaming) |
Sparse Semantic Tokenization and Hierarchical Style Adaptation
We compare a dense semantic tokenizer (CosyVoice 3) with our sparse semantic tokenizer to demonstrate the necessity of a sparse information bottleneck for accent conversion. Furthermore, the hierarchical style adaptation framework enables the preservation of the source speech timbre.
| Source (Indian Accent) |
Target (US Accent) |
AccentDrift (AC) |
Vevo-Style (AC) |
Vevo-Voice (AC + VC) |
CosyVoice3 (VC) |
CosyVoice3 (Recon.) |
|---|---|---|---|---|---|---|