AAAI Workshop 2026 · Voice protection

Keep a public voice from becoming free training data.

A widely recorded public figure can be imitated from collected speech clips and repurposed into misleading synthetic audio. SceneGuard protects recordings before release by adding audible background sound that fits the scene and is difficult to remove cleanly.

Rui Sang · Yuxuan Liu

5.5%
training-time speaker-similarity degradation
0.986
STOI for protected speech
−6.7 pp
zero-shot attack success
Threat story in motion

The same public voice, two very different training paths.

Animated reading: the upper lane shows an unprotected collection-and-cloning risk; the lower lane inserts SceneGuard before release and weakens the identity signal learned by the attacker.

SceneGuard illustrated through a public-figure voice-cloning scenario Donald Trump is used only as an illustrative example of a widely recorded public figure. Unprotected recordings can be collected for unauthorized voice cloning and fabricated messages. SceneGuard instead mixes scene-consistent audible noise using an optimized temporal mask and gain, preserving intelligibility while weakening speaker identity learned from protected recordings. ILLUSTRATIVE PUBLIC-FIGURE SCENARIO NO REAL QUOTE · NO SYNTHETIC AUDIO PUBLIC VOICE SOURCE Donald Trump PUBLIC-FIGURE EXAMPLE Speeches · interviews MANY PUBLIC RECORDINGS WITHOUT PROTECTION Open clip archive Collected from public media Unauthorized voice cloning FABRICATED MESSAGE Misleading audio ILLUSTRATIVE RISK · NOT A QUOTE SCENEGUARD BEFORE RELEASE Scene-consistent audible protection 1 · scene label 2 · matching ambience 3 · optimize m(t) + γ PROTECTED TRAINING DATA Identity match weakens Speaker SIM 1.000 → 0.945 STOI 0.986 · CLEAR SPEECH SceneGuard preserves access to the message while changing the identity evidence available to an unauthorized training pipeline.
Figure · A public-voice threat story with SceneGuard inserted before releaseDonald Trump is an illustrative public-figure example; the figure contains no real quotation or generated audio.
01 · Question

What if the protection sounds like it belongs?

Imperceptible perturbations can be removed by compression or denoising. Scene-consistent audible sound is harder to separate because it behaves like a plausible part of the recording.

Threat

Unauthorized training

An attacker collects protected recordings and fine-tunes or conditions a voice-cloning system.

Fragility

Preprocessing can erase defenses

Common countermeasures target small, imperceptible perturbations before model training.

Trade-off

Audibility is deliberate

The defense accepts contextual noise and measures intelligibility rather than claiming pristine audio.

02 · Method

Optimize a scene-aware mixture, not a generic noise floor.

Select

Scene-conditioned library

PANNs provides an acoustic-scene label, which indexes a library built from ten urban scene classes.

Protect

Speaker-aware objective

ECAPA-TDNN embeddings guide the optimization toward lower speaker similarity, with smoothness and energy regularization.

Gate

Usability + robustness

Protected audio is checked with Whisper, STOI, PESQ, and five preprocessing countermeasures before attack evaluation.

03 · Evidence

Protection survives the preprocessing tested in the paper.

Training attack

1.000 → 0.945

Speaker similarity after training on protected rather than clean data; p < 10⁻¹⁵ and Cohen’s d = 2.18.

Zero-shot attack

20.0% → 13.3%

Attack success rate, a 6.7 percentage-point reduction and 33.5% relative reduction.

Countermeasures

0.901 → 0.688

Speaker similarity ranges from MP3 128 kbps to 8 kHz downsampling; lower is stronger protection.

Training Table 1 reports WER 2.77%, PESQ 2.22, and STOI 0.99. The separate protected-speech quality evaluation reports WER 3.60%, PESQ 2.034, and STOI 0.986; these protocols are intentionally not conflated here.

04 · Boundary

The limitation is part of the design, not fine print.

Evaluation frame

  • LibriTTS: 100 training and 40 test samples for training-attack evaluation.
  • Noise library: 50,000 three-second clips from TAU Urban Acoustic Scenes 2022.
  • MP3, spectral subtraction, low-pass filtering, and downsampling countermeasures.

Claim boundary

  • Audible noise may be unsuitable for pristine studio recordings.
  • PESQ 2.03 is below the paper’s ideal threshold of 3.0.
  • Adaptive attacks designed specifically for scene-aware noise were not tested.
Method figure · original paper

Follow the defense from the scene label to the attacker’s training loop.

The published figure begins with a contextual noise asset, then uses direct mixing or an optimized temporal mask and strength. Its objective makes the trade-off explicit: weaken speaker similarity while preserving signal quality and temporal smoothness.

Published SceneGuard Figure 1 showing scene-aware audible-noise generation and evaluation against training-time and zero-shot voice cloning.
Figure 1 · SceneGuard method overview, reproduced from the paper.The upper band builds usable protected speech; the lower branches evaluate two cloning settings.
Scenes & assets contextual noise libraryDefense generation direct mix or optimized mask + strengthEvaluation training-time and zero-shot cloning

Review the threat model, trade-offs, and ablations.