EAIM @ AAAI 2026 · PMLR 303:1–15

TS-RaMIA

TS-RaMIA: Membership Inference Attacks for Symbolic Music Generation Models

A structure-aware, debiased audit that asks whether a symbolic music work was used to train a generation model—without turning confounders into false evidence.

Yuxuan Liu, Rui Sang, Peihong Zhang, Zhixin Li, Kunyang Zhang, Shengyuan He, Ye Li, Kaiyi Xu, Shengchen Li

First author: Yuxuan Liu

01

Membership auditing

Can a creator test whether a piece was used to train a music generation model?

02

Gray-box · forward-pass only

What the auditor can access—and what remains a deployment gap

Allowed

  • Submit a candidate symbolic-music sequence
  • Obtain per-token log-probabilities through teacher forcing
  • Know the documented tokenization scheme

Not required

  • No gradients
  • No model weights
  • No optimizer state
  • No access to the target training set

Deployment gap

  • Some commercial APIs do not expose per-token probabilities.
  • Generation-only interfaces need sampling-based approximations.
  • Those approximations increase query cost and are not the paper’s primary evaluation.

03

The evaluation trap

Why naïve sequence loss can look convincing for the wrong reason

Raw baseline AUC—Before structural-length control
Length-matched AUC—313 nearest-neighbor pairs
Conditionally calibrated AUC—Residualized against log structural length
  • LengthLonger works contain more opportunities for extreme loss values.
  • Structure countBar, position, and tempo counts can correlate with the split.
  • Event densityDense passages change the loss distribution independently of membership.
Without confounder control, an apparently successful attack may only be recognizing structural complexity.

04

Core insight

Leakage is not uniformly distributed across musical tokens

REMI

BarPositionTempoPitchVelocity

Bar, Position, and Tempo form the structural mask; note, pitch, and velocity tokens have different functions.

ABC

Header|:[]A–G

Headers before the first body line are excluded. Body-level meter, key, or voice changes remain under the paper’s rules; normalized newlines are not structural tokens.

Observation

Per-token NLL is the base observation. Tail statistics surface sparse local anomalies that a full-sequence average can dilute.

Decision rule

A directed, standardized meta-attacker produces the final score under cross-validation. A single raw NLL is not a membership probability.

05

TS-RaMIA pipeline

Five stages, one controlled audit score

Select a stage to inspect it on the unchanged Figure 1 from the paper.

Tokenizer and structural masking

Expose Bar, Position, and Tempo in REMI or body-level structural characters in ABC while excluding formatting artifacts.

Published Figure 1, extracted from the final PMLR paper without redrawing. The transparent outline is a webpage annotation only.
Technical details
Likelihood extraction

Teacher forcing yields per-token NLL. Non-overlapping chunks exclude each chunk’s initial token because it lacks within-chunk context.

Debiasing

Nearest-neighbor length matching and conditional calibration isolate membership evidence from structural length.

Tail statistics

Three structural top-k scores use k ∈ {32, 64, 128}; three windowed p95 statistics capture local peaks.

Meta-attacker

Optional reverse/hierarchical features complete a nine-dimensional vector when available. Each fold applies z-score normalization before L2-regularized, class-weighted logistic regression.

Evaluation

Composer-stratified five-fold cross-validation produces out-of-fold predictions, reducing stylistic leakage between training and evaluation folds.

06

Published evidence

Inspect the controlled Table 1 results

This explorer exposes only aggregate results published in PMLR 303. It never produces a sample-level probability.

TPR@1%FPR—

AUC—

TPR@1%FPR—

TPR@5%FPR—

TPR@10%FPR—

—

Controlled evaluation view; not a per-sample probability.

Table 1 · REMI Transformer results. AUC includes 95% confidence intervals.
ViewMethodAUC [95% CI]TPR@1%FPRTPR@5%FPRTPR@10%FPR
Published Figure 2Rendered from the final paper and cropped only around the figure boundary. It is not a reconstructed curve and no values are digitized from it. Final Table 1 is the numerical reporting source.

07

Deployment-relevant evaluation

Why low false-positive rates matter

1,000

non-member works audited

1% FPR≈ 10false alarms

False accusations are costly, so rights-holder audits need evidence at a strict false-positive budget—not AUC alone.

At the 1% FPR operating point, the length-matched fusion result detects 14.6% of true members. That is recall at a fixed error budget, not the false-positive rate and not detection of most training samples.

The decision threshold must be selected on a development split, never tuned on the test set.

08

Ablations and falsification

What failed—and what that taught us

Table 2 · Length-matched ablations
MethodAUC ± std.TPR@1%FPR
Top-64 is a trade-off

It balances variance against signal dilution; it was not chosen arbitrarily.

Windowed p95 adds a small lift

Local peaks help, but they do not replace structural top-k aggregation.

Equal-N preserves the trend

Performance near the full-piece result argues against piece length being the only signal.

Note-only attack

AUC —

Score inversion can recover approximately 0.68, but only if orientation is known in advance. It is not a usable attack as-is.

EVT tail modeling

AUC —

With 314 non-member samples, fitting was unstable and did not outperform simpler top-k aggregation.

09

Cross-representation transfer

Does the trend survive a different symbolic representation?

Main representationREMI Transformer
Conversion pathMAESTRO MIDI → MusicXML → ABC— · successful
Transfer representationABC / NotaGen
Raw AUC—
TPR@1%FPR—
Length-matched AUC——

10

Privacy–utility checkpoint scan

Membership risk accumulates as training progresses

Figure 2(d), cropped from the final paper without changing the plotted content.

The same audit pipeline is applied to intermediate checkpoints. Attack AUC rises as training continues, indicating increasing membership risk.

The scan can inform selection of a lower-risk checkpoint when utility is adequate. The page does not invent checkpoint values that the text does not report.

11

Limitations and responsible use

Evidence the paper supports—and claims it does not establish

Supported by this paper

  • Forward-pass membership auditing
  • Symbolic-music structural tokens
  • Low-FPR controlled evaluation
  • REMI main evaluation
  • ABC/NotaGen transfer evidence

Not established by this paper

  • Generation-only commercial API performance
  • Proof of copyright infringement
  • Prompt-injection or jailbreak defense
  • End-to-end AI Agent security
  • Universal transfer to arbitrary music models
  • Large-scale production latency
TS-RaMIA provides statistical auditing evidence and should be combined with provenance, licensing records, metadata, and other supporting evidence.

12

Conceptual transfer / future direction — not evaluated in this paper

From structured music auditing to structured Agent auditing

Structural music tokensBar / Position / Tempo

Risk-critical Agent eventsTool calls / permission changes / external transmission / file writes

Sequence confoundersStructural length and event density

Trajectory confoundersTrajectory length / tool count / task difficulty / user type

Group-aware evaluationComposer-stratified folds

Agent evaluationUser-, task-, domain-, or attack-family-stratified splits

Strict operating pointTPR@1%FPR

Safety operating pointDetect unsafe actions without excessively blocking benign users

What transfers

  • Audit the security-relevant parts of a structured sequence.
  • Control confounders before claiming an attack works.
  • Evaluate at deployment-relevant low false-positive rates.
  • Separate exploratory scores from calibrated decisions.
  • Use group-aware splits to prevent evaluation leakage.

What does not transfer directly

  • Membership inference is not prompt-injection detection.
  • Symbolic-music structure is not identical to Agent semantics.
  • The paper does not evaluate Tool Use, RAG, Memory, or mobile-device actions.
  • Production systems may not expose token log-probabilities.
  • The proposed Agent extension is future work.

13

Paper and resources

Verify the primary sources

BibTeX

@InProceedings{pmlr-v303-liu26a,
  title = {TS-RaMIA: Membership Inference Attacks for Symbolic Music Generation Models},
  author = {Liu, Yuxuan and Sang, Rui and Zhang, Peihong and Li, Zhixin and Zhang, Kunyang and He, Shengyuan and Li, Ye and Xu, Kaiyi and Li, Shengchen},
  booktitle = {Proceedings of Machine Learning Research},
  pages = {1--15},
  year = {2026},
  volume = {303},
  publisher = {PMLR},
  url = {https://proceedings.mlr.press/v303/liu26a.html}
}

Enlarged paper figure