Keep a public voice from becoming free training data.
A widely recorded public figure can be imitated from collected speech clips and repurposed into misleading synthetic audio. SceneGuard protects recordings before release by adding audible background sound that fits the scene and is difficult to remove cleanly.
Rui Sang · Yuxuan Liu
5.5%
training-time speaker-similarity degradation
0.986
STOI for protected speech
−6.7 pp
zero-shot attack success
Threat story in motion
The same public voice, two very different training paths.
Animated reading: the upper lane shows an unprotected collection-and-cloning risk; the lower lane inserts SceneGuard before release and weakens the identity signal learned by the attacker.
Figure · A public-voice threat story with SceneGuard inserted before releaseDonald Trump is an illustrative public-figure example; the figure contains no real quotation or generated audio.
01 · Question
What if the protection sounds like it belongs?
Imperceptible perturbations can be removed by compression or denoising. Scene-consistent audible sound is harder to separate because it behaves like a plausible part of the recording.
Threat
Unauthorized training
An attacker collects protected recordings and fine-tunes or conditions a voice-cloning system.
Fragility
Preprocessing can erase defenses
Common countermeasures target small, imperceptible perturbations before model training.
Trade-off
Audibility is deliberate
The defense accepts contextual noise and measures intelligibility rather than claiming pristine audio.
02 · Method
Optimize a scene-aware mixture, not a generic noise floor.
Select
Scene-conditioned library
PANNs provides an acoustic-scene label, which indexes a library built from ten urban scene classes.
Protect
Speaker-aware objective
ECAPA-TDNN embeddings guide the optimization toward lower speaker similarity, with smoothness and energy regularization.
Gate
Usability + robustness
Protected audio is checked with Whisper, STOI, PESQ, and five preprocessing countermeasures before attack evaluation.
03 · Evidence
Protection survives the preprocessing tested in the paper.
Training attack
1.000 → 0.945
Speaker similarity after training on protected rather than clean data; p < 10⁻¹⁵ and Cohen’s d = 2.18.
Zero-shot attack
20.0% → 13.3%
Attack success rate, a 6.7 percentage-point reduction and 33.5% relative reduction.
Countermeasures
0.901 → 0.688
Speaker similarity ranges from MP3 128 kbps to 8 kHz downsampling; lower is stronger protection.
Training Table 1 reports WER 2.77%, PESQ 2.22, and STOI 0.99. The separate protected-speech quality evaluation reports WER 3.60%, PESQ 2.034, and STOI 0.986; these protocols are intentionally not conflated here.
04 · Boundary
The limitation is part of the design, not fine print.
Evaluation frame
LibriTTS: 100 training and 40 test samples for training-attack evaluation.
Noise library: 50,000 three-second clips from TAU Urban Acoustic Scenes 2022.
MP3, spectral subtraction, low-pass filtering, and downsampling countermeasures.
Claim boundary
Audible noise may be unsuitable for pristine studio recordings.
PESQ 2.03 is below the paper’s ideal threshold of 3.0.
Adaptive attacks designed specifically for scene-aware noise were not tested.
Method figure · original paper
Follow the defense from the scene label to the attacker’s training loop.
The published figure begins with a contextual noise asset, then uses direct mixing or an optimized temporal mask and strength. Its objective makes the trade-off explicit: weaken speaker similarity while preserving signal quality and temporal smoothness.
Figure 1 · SceneGuard method overview, reproduced from the paper.The upper band builds usable protected speech; the lower branches evaluate two cloning settings.
Scenes & assets contextual noise libraryDefense generation direct mix or optimized mask + strengthEvaluation training-time and zero-shot cloning
Review the threat model, trade-offs, and ablations.