ISMIR 2025Peer-reviewed research

MAIA

Music Adversarial Inpainting Attack

Locate influential musical regions, regenerate only those regions, and evaluate security impact together with perceptual quality.

Yuxuan Liu, Peihong Zhang, Rui Sang, Zhixin Li, Shengchen Li

Research question

Can we induce an MIR model to misclassify by regenerating only the musical segments it relies on most?

Rather than dispersing dense additive perturbations across the waveform, MAIA first identifies the musical regions that contribute most to the model’s current decision. It then uses surrounding musical context to regenerate only those bounded regions, with guidance from the attack objective. The central question is whether this localized, model-aware generative intervention can achieve an effective attack while preserving musical coherence and perceptual quality.

01

Locate model-critical regions

Identify the time-frequency regions that contribute most to the target model’s current decision, using gradient-based analysis or query-based coarse-to-fine localization.

02

Regenerate with musical context

Use a conditional music inpainting model to reconstruct only the selected regions from their surrounding context instead of perturbing the entire track.

03

Balance attack and fidelity

Reduce confidence in the correct class while jointly evaluating attack effectiveness, musical coherence and perceptual quality.

Threat model

Two access settings, one bounded-edit principle

Gradients available

White-box MAIA

Grad-CAM localizes influential regions. Reconstruction and untargeted attack losses guide GACELA for at most 10 optimization iterations, with loss weights selected from {0.5, 1.0, 2.0}.

Query feedback only

Black-box MAIA

A 0.5-second coarse-to-fine masking procedure estimates region importance. CMA-ES then searches GACELA's latent space under a 1,000-query budget.

The paper evaluates cover-song identification with CoverHunter on SHS100K and music-genre classification with IDS-NMR on GTZAN.

Why local regeneration

Change less, but change what matters

MAIA does not edit the whole waveform. It first ranks influential regions, then constrains generative changes to a small set of selected intervals. This makes the intervention interpretable in time and gives the inpainting model surrounding context for musically plausible reconstruction.

Model-aware localizationBounded generative editsSecurity–quality evaluation

How MAIA Works

A three-step importance-driven adversarial inpainting framework

01

Importance Analysis

Identify critical time-frequency regions that most influence the model's decision

White-boxClass gradients reveal positive evidence. Black-boxRepeated masks reveal influential intervals.
Grad-CAM heatmap on a mel-spectrogram from the MAIA paper
Paper figureThe paper's Grad-CAM heatmap makes the default method view concrete: brighter time-frequency regions contribute more strongly to the model's current decision.
Mechanism schematics

Technical deep dive: How are the critical regions located?

Open either access setting when you want the computation, formula, and evidence boundary.

White-box localization: Grad-CAMExpand only when needed
1. Input and class scoreThe waveform is converted into a time-frequency representation and passed through the target MIR model. Grad-CAM starts from the score of the model’s current class.
2. Weight feature maps by class gradientsGradients of the selected class score are propagated to the target convolutional layer. Spatially averaged gradients assign a class-specific importance weight to each feature-map channel.
Grad-CAM class-gradient weighting: the current class score is back-propagated to convolutional feature maps, whose spatially averaged gradients produce the channel weight alpha k superscript c Mechanism schematic
Grad-CAM channel weight: alpha subscript k superscript c equals one over Z times the sum over u and v of the partial derivative of y-hat superscript c with respect to F subscript k superscript l at u, v.
3. Map positive evidence back to the spectrogramThe weighted feature maps are aggregated and passed through ReLU. After normalization and coordinate mapping, high-intensity areas become candidate adversarial regions.
Mc(u,v) = ReLU(∑k αkc Fkl(u,v))
Grad-CAM heatmap on mel-spectrogram from the MAIA paper Paper figure

Grad-CAM is an existing localization method. MAIA’s contribution is not the invention of Grad-CAM, but its use within an importance-driven music inpainting attack pipeline.

Black-box localization: Coarse-to-fine queriesExpand only when needed
1. Partition and maskThe audio is divided into non-overlapping coarse chunks. Each chunk is zero-masked in turn, with a short Tukey taper applied at its boundaries.
Zero mask+Tukey boundary taper
2. Query the model and rank the chunksA chunk is considered important when masking it produces a large change in the loss associated with the original label. Division by duration keeps scores comparable across different segment lengths.
I(Ci) = [L(M(x̃−Ci), y) − L(M(x), y)] / duration(Ci)
3. Refine until the desired granularityThe highest-scoring segment is recursively subdivided and evaluated again. Refinement continues until the desired granularity is reached or the query budget is exhausted, after which the top-ranked regions are selected.
Illustrative mechanism — not a measured sample

This procedure does not require model parameters or gradients, but it does require prediction feedback from repeated model queries. The paper’s 1,000-query cap applies to the overall black-box attack, not to each localization round.

02

Segment Selection

Select the highest-ranked regions for bounded generative modification

Priority-based Selection

Rank candidate regions by model-derived importance and retain the top regions under the edit budget.

Refinement Strategy

In the black-box setting, recursively split the highest-scoring coarse region to obtain finer localization.

Selected important segments for inpainting

Selected important segments for inpainting

03

Adversarial Inpainting

Reconstruct selected segments with GACELA and task-specific adversarial guidance

Mechanism schematic

Context-conditioned local regeneration

GACELA does not generate an isolated segment from scratch. Its generator is conditioned on the audio surrounding the masked interval, allowing the regenerated gap to connect with the musical content before and after it. MAIA keeps the outside context unchanged and modifies only the selected region.

Original audio xApply mask m
Unchanged left contextMasked regionUnchanged right context
Context representation / log-magnitude mel spectrogram+z→Context-conditioned generationPretrained GACELA generator Gθ
Unchanged left contextGenerated local segmentUnchanged right context

Only the selected interval is replaced; audio outside the mask comes directly from the original input.

Before Inpainting

Before Inpainting

→
After Inpainting

After Inpainting

The paper figure illustrates context-conditioned local music reconstruction.

White-box optimization: reconstruction, attack, and re-inpaintingExpand only when needed
ℒ = λrec ℒrec(xinp(k), x) + λatt ℒattack(M(xinp(k)), y)

ℒrecKeeps the local reconstruction consistent with the musical context and perceptual quality.

ℒattackReduces confidence in the original correct label y for an untargeted attack.

λrec, λattControl the fidelity–attack trade-off; both are grid-searched over {0.5, 1.0, 2.0}.

IterationsThe white-box experiment performs at most 10 optimization iterations.

xinp(k+1) ← xinp(k) − α · sign(∇xinpℒ ⊙ m)

The mask confines the update to the selected region.

xinp(k+1) = x ⊙ (1−m) + Gθ(xinp(k+1) ⊙ m, x ⊙ (1−m)) ⊙ m
Forward evaluation→Masked gradient update→Context-conditioned re-inpainting→Evaluate again

The pretrained GACELA model is used as a generative prior. MAIA iteratively updates the selected audio region and re-projects it through the inpainting process; it does not jointly retrain GACELA with the target MIR model.

Black-box optimization: latent-space search with CMA-ESExpand only when needed
z(k+1) = CMA-ES(z(k), F(M, xinp(k)))
x̂inp = Gθ(ẑ, x ⊙ (1−m))
  1. Process candidate regions from highest to lowest importance.
  2. Sample latent candidates from the current CMA-ES distribution.
  3. Use GACELA to generate candidate fills under the same surrounding context.
  4. Query the target model for prediction-label or confidence feedback.
  5. Favor candidates that improve the attack objective and update the latent distribution.
  6. Re-inpaint while preserving context continuity.
  7. Stop after misclassification or when the overall query budget is exhausted.
White-boxUses gradients to update the selected region.Black-boxUses model feedback to search the GACELA latent variable.Shared constraintBoth modify only localized regions and condition on the unchanged context.

Verified interactive sample

Real model-based regeneration at MAIA-selected regions

This interactive example isolates the generative inpainting component of MAIA. The displayed audio is produced by an actual music inpainting model at the selected regions. Paper-level attack performance is reported separately and is not inferred from this individual example.

Source sampleGoldberg Variation 20 — 30 s excerpt
1

Locate the influential region

The highlighted centers preserve the three region positions used by the earlier MAIA page. They are provenance-preserved interface metadata, not a new target-model localization run.

Original spectrogram with three selected regions
Demonstrated component: music audio inpaintingDuration: 30.0 s
2

Regenerate only the selected region

Each view uses the same 30-second time axis. GACELA's 32-frame checkpoint alignment is recorded in the manifest, and the original region centers are preserved.

Original
Original audio spectrogram
Masked
Masked audio spectrogram
Regenerated
GACELA regenerated audio spectrogram
Inspect the measured difference spectrogram
Difference spectrogram calculated from the original and regenerated audio
Computed from the two WAV files. Changes outside the effective regions are zero in the validation output.
3

Listen with synchronized A/B comparison

All three sources share one timeline. Switching tracks preserves the current position; region looping constrains playback to the selected interval.

0:00.00:30.0

Ready. Original selected.

4

Verify how this sample was generated

Generated offline with the stated model and served as a verified static artifact for reliable playback.

Real model outputDeterministic seedVerified asset
Open provenance record
Model
Checkpoint
Seed
Effective regions
Target model evaluated
Generation date
Input SHA-256
Output SHA-256
GACELA code commit
Boundary crossfade

Aggregate paper results

Evidence and limitations

These values are aggregate results from Table 1 of the paper. They are not predictions or metrics for the interactive sample above.

MAIA-WB / CSI92.8%ASR · MOS 4.0
MAIA-WB / MGC93.5%ASR · MOS 3.8
MAIA-BB / CSI80.1%ASR · MOS 3.6
MAIA-BB / MGC77.9%ASR · MOS 3.3
Table 1 excerpt — MAIA rows reported for two tasks and two access settings
SettingTaskASRPost-attack scoreFADLSDMOS
MAIA-WBCSI92.8%mAP 0.48811.251.584.0
MAIA-WBMGC93.5%ACC 0.46613.851.943.8
MAIA-BBCSI80.1%mAP 0.59412.561.903.6
MAIA-BBMGC77.9%ACC 0.60114.681.853.3

Scope of evaluation

The paper studies two MIR tasks and two target system–dataset pairs; broader generalization requires additional targets and music domains.

Access assumptions

White-box MAIA requires gradients and model access. Black-box MAIA relies on query feedback and uses up to 1,000 queries.

The paper evaluates ASR, post-attack mAP or accuracy, FAD using MERT-V0, LSD, and a five-point subjective listening study with 100 participants.

Evaluation roadmap

Beyond attack success: transferability and transformation robustness

Not evaluated in the current paper

Mechanism-based hypotheses and an evaluation roadmap—not additional results reported by the MAIA paper.

PropertyAdditive perturbation attacksMAIA mechanism hypothesisEvidence status
Cross-model transferabilityPerturbations optimized for one source model may transfer when different models share similar decision boundaries, but transfer is not guaranteed.MAIA’s localization and optimization are also source-model-aware. Local content-level regeneration may affect features shared by multiple models, but it may also overfit the source model’s important regions.Open question—requires source-to-target transfer experiments.
Lossy transcoding and compressionFine-grained, low-amplitude perturbations can be attenuated or distorted by MP3/AAC encoding, resampling, and bitrate changes.Because MAIA replaces a local region with plausible generated musical content, its modification may be less dependent on exact waveform-level noise. This suggests possible persistence after transcoding, but it was not tested.Hypothesis only—no codec-robustness result is reported.
Denoising and filteringNoise-like components may be reduced by denoisers or frequency filtering, although adaptive attacks can be designed against a known defense.Content-like inpainted segments may not be recognized as noise by a denoiser, but denoising can still alter generated details and reduce attack effectiveness.Open question—requires post-denoising attack evaluation.
Temporal and playback transformationsCropping, time shifting, resampling, loudness normalization, or playback-recording can disrupt perturbations that depend on precise alignment.MAIA modifies bounded musical content, but its importance mask and target-model effect may still depend on timing and alignment.Not established by the current experiments.
01

Cross-model transfer matrix

Craft attacks on one MIR model and evaluate them directly on different architectures and training datasets without re-optimizing the adversarial sample.

02

Audio transformation suite

Evaluate MP3/AAC transcoding at multiple bitrates, resampling, denoising, filtering, loudness normalization, time shifting, and playback-recording transformations.

03

Retained effectiveness and utility

Measure post-transformation attack success, confidence change, perceptual distance, and listening quality instead of reporting attack success alone.

MAIA demonstrates a stronger attack–quality trade-off than the evaluated PGD, C&W, NES, and ZOO baselines in the paper’s original settings. Whether this advantage extends to cross-model transfer or real audio-processing pipelines remains an empirical question.

Methodological bridge

From Localized Music Attacks to Multimodal Agent Safety

A methodological transfer, not an experimental result of MAIA.

MAIALocate influential audio regions
Method transfer
Agent SafetyIdentify high-risk observations, context, memory, or actions
MAIAApply bounded generative edits
Method transfer
Agent SafetyConstruct controlled multimodal counterfactual inputs
MAIAEvaluate white-box and black-box attacks
Method transfer
Agent SafetyPerform gradient-based and query-based red teaming
MAIAMeasure attack effectiveness and perceptual quality together
Method transfer
Agent SafetyEvaluate security impact, task utility, side effects, and human perception
MAIA does not directly solve prompt injection, jailbreak, tool authorization, or trusted execution. The bridge identifies a transferable evaluation methodology: locate influential inputs, apply bounded counterfactual changes, and jointly measure attack effectiveness and utility preservation.

Paper and resources

Read, reproduce, and inspect

Citation

The official English title is retained in the formal citation.

@inproceedings{liu2025maia,
  title={MAIA: An Inpainting-Based Approach for Music Adversarial Attacks},
  author={Liu, Yuxuan and Zhang, Peihong and Sang, Rui and Li, Zhixin and Li, Shengchen},
  booktitle={Proceedings of the 26th International Society for Music Information Retrieval Conference},
  year={2025}
}