Research manuscript · Streaming speech privacy

CloneBlock-RT

Streaming Voice Protection Against Zero-Shot Voice Cloning in Live Telephony

Protect speech as it arrives.
Not after the recording becomes a clone.

Yuxuan Liu, Kunyang Zhang, Peihong Zhang, Yiqiang Cai, Yizhou Tan, Shengchen Li

Xi’an Jiaotong-Liverpool University, Suzhou, China · Author affiliation as in the manuscript

Real-time privacy protection against voice cloning through zero-shot text-to-speech (TTS): preserve speech content during a call while reducing the recording’s usefulness as an unauthorized voice reference.

Single passCausal inference; no online optimization
32 bandsInput-dependent, bounded spectral gains
20 msInitial analysis buffer, not call latency

00 / THE TTS CONNECTION

Give text a voice.
Keep voice reuse consensual.

TTS turns text into speech. Reference-conditioned TTS can also use a short recording to generate new content in a target speaker’s voice—the form of voice cloning studied here. It can enable consensual personalized speech; our research addresses unauthorized reuse of recordings.

FROM THE REAL-WORLD RISK TO THE TTS TASK

New text + reference recording of a target speaker → new speech in that speaker’s voice

CloneBlock-RT protects the upstream reference recording. It is neither a new TTS generator nor a detector applied to already-generated fake speech.

Zero-shot TTS / voice cloning

In this setup, the TTS model is not fine-tuned for the target speaker; reference audio conditions inference. Here, zero-shot does not mean that no reference recording is provided.

Multi-reference zero-shot TTS

The paper supplies multiple clips from the same call to XTTS-v2 without target-speaker adaptation training. We therefore call this multi-reference zero-shot TTS: increasing the number of reference clips does not by itself establish a few-shot learning setup.

Why not simply label it zero-shot / few-shot TTS?

Few-shot TTS may describe adaptation from a small speaker-specific dataset or a particular in-context example protocol, depending on the work. This paper does not evaluate few-shot target-speaker adaptation. It is relevant to personalized TTS using limited reference audio, but its evaluated setting remains zero-shot / multi-reference zero-shot TTS.

01 / THE PROBLEM

The call ends.
The recording does not.

An attacker can retain a call, then choose and combine clips for zero-shot voice cloning. The protector must act in real time, before it knows which segments will later be selected.

Live speech passes through CloneBlock-RT before the telephone channel. The received recording may later be assembled into a cloning reference.
Figure 1 · Protect before transmission; construct cloning references after capture. Illustrative signals, not audio samples. Click to open the full-resolution image.
01

More than fast inference

Protection should survive telephone processing, not merely change the pre-transmission signal.

02

More than one reference clip

A call supplies multiple, potentially overlapping windows. One-clip results do not establish protection under multi-reference collection.

03

More than an embedding proxy

Alongside embedding pooling, XTTS-v2 synthesis tests whether the reduction persists in cloned speech.

02 / ONLINE PROTECTION

A bounded adjustment.
For every arriving frame.

Rather than repeat a fixed perturbation, the causal predictor uses current and past speech context to produce 32 band gains, modify complex STFT coefficients, and reconstruct a protected waveform.

Incoming speech to non-centered STFT; a two-layer 64-unit causal GRU predicts 32 bounded gains, applied before inverse STFT and overlap-add. Buffers and hidden state persist.
Figure 2 · The deployed path uses only arrived speech. Positive gains preserve the phase of the modified STFT coefficients. Waveforms, spectra and gain traces are schematic.

Bounded does not mean inaudible

Gb,t = 1 + 0.18 tanh(Ĝb,t)

Gains lie in [0.82, 1.18]. This bounds per-coefficient magnitude modulation; speech utility must still be measured separately.

Causal does not mean zero latency

16-kHz mono, a 20-ms Hann window, a 10-ms hop and a 512-point FFT. Analysis initially buffers 20 ms; the predictor has no future-frame look-ahead.

Analysis buffering, GRU hidden state and synthesis overlap state persist across chunks.

03 / ARCHITECTURE & OFFLINE TRAINING

We train the protector.
Not the voice cloner.

CloneBlock-RT contains a trainable causal gain predictor within Dφ. It learns frame-wise spectral adjustments that lower post-telephone speaker similarity while constraining distortion. Parameters are optimized offline; a live call uses a single forward pass without parameter updates.

INPUT → LEARNING → OUTPUT

Speech in; protected speech out

The system accepts 16-kHz mono speech x. After non-centered STFT analysis, the causal predictor uses present and past context. Two GRU layers with 64 units each feed a head producing 32 frame-wise gains for mel-spaced bands over 0–8 kHz.

After tanh bounding, gains are mapped to frequency bins and multiplied with the complex spectrum. iSTFT and overlap-add reconstruct x̃. The head predicts gains—not text, speaker labels, or cloned speech.

Input-feature boundary: the manuscript does not specify whether the GRU receives magnitudes, log spectra, or other features, nor its input dimension.

DATA & OPTIMIZATION · §3.1–3.2

Speaker-disjoint offline training

Corpus
VCTK v0.92 · 16 kHz · mono
Train / validation / test
78 / 10 / 20 disjoint speakers
Each batch
16 crops × 4 seconds
Optimizer / learning rate
Adam / 0.001
Training length
50,000 updates—not epochs
Model selection
Lowest validation mean similarity subject to channel-free STOI ≥ 0.98

The 30-second pseudo-calls and 96 windows per session belong to evaluation, not training batches. LibriSpeech supplies evaluation texts and utility-test speech, not the reported protector training set.

Which components learn, and which supply the training signal?
ComponentParameter updatesRole
Protector: GRU + gain headYes; the learned modulePredict bounded contextual gains; retained at deployment
STFT / iSTFT / overlap-addNo; signal processingSpectral analysis and waveform reconstruction; retained at deployment
Matched telephone transform T(·; θ)Sampled conditions, not a learned channel networkSimulate transmission and pass gradients during training
ECAPA-TDNNNo; pretrained weights frozenSpeechBrain spkrec-ecapa-voxceleb supplies training and evaluation embeddings
XTTS-v2Not part of protector trainingSynthesize cloned speech afterward to evaluate downstream protection

What happens in one training update?

  1. Sample speech and protect it.Take 4-second training crops x and compute x̃ = Dφ(x). Learn an input-dependent transformation rather than separately optimizing a perturbation for each utterance.
  2. Apply one matched telephone condition.Sample θ and apply matched band-limiting, resampling, smooth AGC, and Gaussian noise to x and x̃, including shared noise. Otherwise channel differences could themselves change similarity.
  3. Measure identity and preservation losses.Pass both received waveforms through frozen ECAPA-TDNN, L2-normalize their embeddings, and compute cosine similarity. Also compare pre-channel spectra and RMS energy of x and x̃, following Eq. 2.
  4. Backpropagate; update only the protector.Gradients pass through the frozen encoder’s input, differentiable channel, and reconstruction to the GRU and gain head. Adam updates φ; freezing encoder weights does not detach its input gradients.
Protected and unprotected branches share telephone settings and noise. A frozen ECAPA-TDNN encoder produces normalized embeddings; the post-channel cosine loss updates only the protector.
Figure 3 · A frozen encoder still passes input gradients to the protector. Only protector parameters change. Preservation constraints follow Method Eq. 2.

The losses: reduce identity similarity without simply destroying speech

1 / IDENTITY · POST-CHANNEL

Lspk = Eθ[cos(h(T(x; θ)), h(T(Dφ(x); θ)))]

Minimize identity similarity between original and protected branches after transmission. Average scalar losses over sampled channels—not waveforms or embeddings. The objective specifies no target speaker to imitate.

2 / RECONSTRUCTION · PRE-CHANNEL

Lrec = Σr∈R ‖|STFTr(x)| − |STFTr(x̃)|‖1

Constrain magnitude-spectrum differences at multiple STFT resolutions to limit excessive structural changes. The manuscript does not enumerate the window, hop, and FFT settings in R.

3 / ENERGY · AS DEFINED IN EQ. 2

Lenergy = |RMS(x̃) − RMS(x)|

Constrain the change in overall RMS amplitude. Both preservation terms are computed before telephone processing, so modifications removed by the channel are still penalized.

Minimize (Eq. 2): L = λspkLspk + λrecLrec + λenergyLenergy

Matched channels, averaged conditions

Training randomizes band-limiting, resampling, smooth AGC and Gaussian noise, shared across both branches. The objective minimizes expected cosine similarity over sampled conditions—not a worst-case guarantee for every channel.

Choose protection under a utility constraint

α is selected from {0.10, 0.14, 0.18, 0.22} by minimizing validation similarity subject to channel-free STOI ≥ 0.98, yielding α = 0.18. Opus and G.711 are evaluation-only, not training transforms.

Which reproduction details are still unspecified?

Beyond the third-loss discrepancy, reproduction needs the GRU input features and normalization, gain-head layers and band interpolation, the multi-resolution STFT set R, channel sampling ranges and distributions, exact speaker IDs, parameter initialization, and hidden-state handling during chunk training. The paper establishes that the protector is trained, but does not uniquely specify the full implementation.

Source: supplied manuscript §§2.1–2.3 and §§3.1–3.2 (pages 2–4). Multi-reference selection and XTTS synthesis are evaluation procedures, not training-loss components.

04 / SESSION-LEVEL EVALUATION

The attacker can look back.
The protector cannot.

Twenty held-out VCTK speakers yield 100 pseudo-calls of 30 seconds, with 96 candidate windows per session. Greedy selection targets the captured audio’s own observed centroid; the unprotected session centroid is used only for scoring.

Recorded clips are selected using their observed centroid. Separate paths test embedding pooling and XTTS-v2 synthesis; measured XTTS similarities decrease from 0.70/0.88/0.92 to 0.58/0.72/0.78 at K=1/16/64.
Figure 4 · Evaluation procedure above; reported Table 1 measurements below. Windows may overlap, so K is not unique speech duration; cloning also imposes reference-length limits.

05 / MEASURED EVIDENCE

The improvement.
And its boundaries.

All values below come from the supplied manuscript, not an inference run on this website. Similarity reduction is not an impersonation success-rate measurement.

Table 1 · Cosine similarity (lower is better); all reported baselines
SystemReference embeddingsXTTS cloned output
K=1K=16Mean over KK=1K=16K=64
No defense0.72020.98610.88300.700.880.92
ClearMask0.65930.93260.85580.620.790.85
Enkidu0.64330.90180.85190.570.740.81
CloneBlock-RT0.67330.92030.82450.580.720.78

Bold marks column minima. Reference mean equally weights K={1,2,4,8,16}, not an integrated area; reference K=1 values were reconstructed from rounded summaries in the paper. CloneBlock-RT does not outperform Enkidu at every K or on every metric.

Table 2 · Speech utility over 200 utterances, mean ± standard deviation
ConditionSystemWER (%) ↓STOI ↑
NoneNo defense52.84 ± 24.881.0000 ± 0.0000
CloneBlock-RT52.14 ± 23.810.9887 ± 0.0047
Opus 16 kb/sNo defense54.27 ± 25.531.0000 ± 0.0000
CloneBlock-RT54.68 ± 25.890.9242 ± 0.0410
G.711No defense53.76 ± 24.521.0000 ± 0.0000
CloneBlock-RT53.09 ± 24.470.9823 ± 0.0119

WER changes remain within 0.70 percentage points on an already high-error ASR pipeline, not evidence of recognition equivalence. STOI drops most under Opus. Unit STOI baselines follow from condition-matched unprotected references—not distortion-free transmission.

STREAMING EXECUTION

Real-time throughput.
Not end-to-end call latency.

GPU: NVIDIA A100, FP32, batch one. Timing includes analysis, gain prediction, synthesis and GPU transfers; excludes codecs and file I/O. The manuscript does not specify the CPU model.

RTF = processing time / audio duration
ChunkCPU RTFGPU RTF
20 ms0.0360.019
40 ms0.0250.013

06 / PAPER & FIGURES

From manuscript
to inspectable design.

The technical illustrations are standalone local SVG and high-resolution PNG files, not diagrams assembled in the webpage. This is a research explainer, not an online protection or cloning service.