AccentDrift: Real-time Streaming Accent Conversion via
Sparse Speech Tokenization

Anonymous Authors
Anonymous Groups

Abstract

Although recent speech AI advances have enabled real-time speech understanding and generation, real-time accent conversion (AC) for second-language (L2) speakers in real-world interactions remains underexplored. In this paper, we present AccentDrift, a real-time streaming AC system that preserves linguistic content and speaker identity while transforming speech into a target accent with minimal latency. From an information-theoretic perspective, we adopt sparse speech tokenization to extract linguistic information from continuous speech. An accent adapter injects accent style into sparse semantic tokens, followed by a timbre adapter that hierarchically generates speaker-specific acoustics. To support streaming, we adopt a cache-aware streaming architecture and parallel streams, achieving 520 ms latency. AccentDrift outperforms prior parallel AC models, and demonstrates the feasibility of generating native-like accents for L2 speakers.


Accent Conversion Comparison

We used the same input samples for L2 speaker (L2-ARCTIC) from the demo page of Streaming Baseline [28]

[28] Nguyen, Tuan-Nam, et al. "Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement." Proc. Interspeech 2025. 2025.

Input
(Arabic Accent)
AccentDrift
(Streaming, 0.5s)
AccentDrift
(Streaming, 0.9s)
Streaming Baseline
(Streaming, 0.8s)
Vevo-Style
(Non-Streaming)
Input
(Chinese Accent)
AccentDrift
(Streaming, 0.5s)
AccentDrift
(Streaming, 0.9s)
Streaming Baseline
(Streaming, 0.8s)
Vevo-Style
(Non-Streaming)
Input
(Vietnamese Accent)
AccentDrift
(Streaming, 0.5s)
AccentDrift
(Streaming, 0.9s)
Streaming Baseline
(Streaming, 0.8s)
Vevo-Style
(Non-Streaming)
Input
(Indian Accent)
AccentDrift
(Streaming, 0.5s)
AccentDrift
(Streaming, 0.9s)
Streaming Baseline
(Streaming, 0.8s)
Vevo-Style
(Non-Streaming)
Input
(Korean Accent)
AccentDrift
(Streaming, 0.5s)
AccentDrift
(Streaming, 0.9s)
Streaming Baseline
(Streaming, 0.8s)
Vevo-Style
(Non-Streaming)

Sparse Semantic Tokenization and Hierarchical Style Adaptation

We compare a dense semantic tokenizer (CosyVoice 3) with our sparse semantic tokenizer to demonstrate the necessity of a sparse information bottleneck for accent conversion. Furthermore, the hierarchical style adaptation framework enables the preservation of the source speech timbre.

Source
(Indian Accent)
Target
(US Accent)
AccentDrift
(AC)
Vevo-Style
(AC)
Vevo-Voice
(AC + VC)
CosyVoice3
(VC)
CosyVoice3
(Recon.)