Improving Model Expressivity and Speaker Matching in Low-Latency Voice Conversion Models [paper (pending)] [code (pending)]

Interspeech 2026 (Submission)

Abstract

Real-time voice conversion systems struggle to achieve high speaker similarity in zero-shot scenarios, especially under lightweight computational constraints. This challenge stems from lack in model expressivity and the difficulty of capturing comprehensive speaker-relevant information within the strict latency and model size requirements. We propose a real-time voice conversion framework that addresses this limitation by fusing complementary information into conventional global speaker embeddings. To ensure robust feature disentanglement, we furthermore employ an encoder-specific information perturbation strategy during training. Our approach maintains causal inference requirements and preserves prosodic and linguistic content while operating under low-latency constraints. Experimental evaluations demonstrate improvements speaker-matching metrics over to state-of-the-art real-time baselines.
Pipeline Diagram

Audio Comparison

Source (Unseen)
Target (Unseen)
RT-VC
StreamVC
Proposed

Extra: Singing Voice and Multilingual Conversion

Source
Target
Proposed


Video Demonstration

The following video shows the real-time voice conversion framework in action, demonstrating speaker matching and latency performance.


Ethics Statement

The ethical considerations associated with real-time voice conversion stem from broader concerns surrounding voice conversion and generative speech technologies, particularly their potential to enable impersonation and infringe on individual privacy. To address these risks, training scripts and general source code will not be released as open source. Rather we provide a Torchscript export, a conversion demonstration script and evaluation scripts for reproducibility. We do not condone or accept any misuse of VC technology. The present research is conducted to advance the state of the art and contribute new knowledge to the field.