A symbolic sequence under audit
EAIM @ AAAI 2026 · PMLR 303:1–15
TS-RaMIA
TS-RaMIA: Membership Inference Attacks for Symbolic Music Generation Models
A structure-aware, debiased audit that asks whether a symbolic music work was used to train a generation model—without turning confounders into false evidence.
01
Membership auditing
Can a creator test whether a piece was used to train a music generation model?
A statistical membership-risk score
H₀: x ∉ Dtrain · H₁: x ∈ Dtrain
Then evaluated on a held-out test set
02
Gray-box · forward-pass only
What the auditor can access—and what remains a deployment gap
Allowed
- Submit a candidate symbolic-music sequence
- Obtain per-token log-probabilities through teacher forcing
- Know the documented tokenization scheme
Not required
- No gradients
- No model weights
- No optimizer state
- No access to the target training set
Deployment gap
- Some commercial APIs do not expose per-token probabilities.
- Generation-only interfaces need sampling-based approximations.
- Those approximations increase query cost and are not the paper’s primary evaluation.
03
The evaluation trap
Why naïve sequence loss can look convincing for the wrong reason
- LengthLonger works contain more opportunities for extreme loss values.
- Structure countBar, position, and tempo counts can correlate with the split.
- Event densityDense passages change the loss distribution independently of membership.
Without confounder control, an apparently successful attack may only be recognizing structural complexity.
04
Core insight
Leakage is not uniformly distributed across musical tokens
REMI
Bar, Position, and Tempo form the structural mask; note, pitch, and velocity tokens have different functions.
ABC
Headers before the first body line are excluded. Body-level meter, key, or voice changes remain under the paper’s rules; normalized newlines are not structural tokens.
Observation
Per-token NLL is the base observation. Tail statistics surface sparse local anomalies that a full-sequence average can dilute.
Decision rule
A directed, standardized meta-attacker produces the final score under cross-validation. A single raw NLL is not a membership probability.
05
TS-RaMIA pipeline
Five stages, one controlled audit score
Select a stage to inspect it on the unchanged Figure 1 from the paper.
Expose Bar, Position, and Tempo in REMI or body-level structural characters in ABC while excluding formatting artifacts.
Technical details
Teacher forcing yields per-token NLL. Non-overlapping chunks exclude each chunk’s initial token because it lacks within-chunk context.
Nearest-neighbor length matching and conditional calibration isolate membership evidence from structural length.
Three structural top-k scores use k ∈ {32, 64, 128}; three windowed p95 statistics capture local peaks.
Optional reverse/hierarchical features complete a nine-dimensional vector when available. Each fold applies z-score normalization before L2-regularized, class-weighted logistic regression.
Composer-stratified five-fold cross-validation produces out-of-fold predictions, reducing stylistic leakage between training and evaluation folds.
06
Published evidence
Inspect the controlled Table 1 results
This explorer exposes only aggregate results published in PMLR 303. It never produces a sample-level probability.
AUC—
TPR@1%FPR—
TPR@5%FPR—
TPR@10%FPR—
—
Controlled evaluation view; not a per-sample probability.
| View | Method | AUC [95% CI] | TPR@1%FPR | TPR@5%FPR | TPR@10%FPR |
|---|
07
Deployment-relevant evaluation
Why low false-positive rates matter
1,000
non-member works audited
False accusations are costly, so rights-holder audits need evidence at a strict false-positive budget—not AUC alone.
At the 1% FPR operating point, the length-matched fusion result detects 14.6% of true members. That is recall at a fixed error budget, not the false-positive rate and not detection of most training samples.
The decision threshold must be selected on a development split, never tuned on the test set.
08
Ablations and falsification
What failed—and what that taught us
| Method | AUC ± std. | TPR@1%FPR |
|---|
It balances variance against signal dilution; it was not chosen arbitrarily.
Local peaks help, but they do not replace structural top-k aggregation.
Performance near the full-piece result argues against piece length being the only signal.
Note-only attack
AUC —Score inversion can recover approximately 0.68, but only if orientation is known in advance. It is not a usable attack as-is.
EVT tail modeling
AUC —With 314 non-member samples, fitting was unstable and did not outperform simpler top-k aggregation.
09
Cross-representation transfer
Does the trend survive a different symbolic representation?
10
Privacy–utility checkpoint scan
Membership risk accumulates as training progresses
The same audit pipeline is applied to intermediate checkpoints. Attack AUC rises as training continues, indicating increasing membership risk.
The scan can inform selection of a lower-risk checkpoint when utility is adequate. The page does not invent checkpoint values that the text does not report.
11
Limitations and responsible use
Evidence the paper supports—and claims it does not establish
Supported by this paper
- Forward-pass membership auditing
- Symbolic-music structural tokens
- Low-FPR controlled evaluation
- REMI main evaluation
- ABC/NotaGen transfer evidence
Not established by this paper
- Generation-only commercial API performance
- Proof of copyright infringement
- Prompt-injection or jailbreak defense
- End-to-end AI Agent security
- Universal transfer to arbitrary music models
- Large-scale production latency
TS-RaMIA provides statistical auditing evidence and should be combined with provenance, licensing records, metadata, and other supporting evidence.
12
Conceptual transfer / future direction — not evaluated in this paper
From structured music auditing to structured Agent auditing
Structural music tokensBar / Position / Tempo
Risk-critical Agent eventsTool calls / permission changes / external transmission / file writes
Sequence confoundersStructural length and event density
Trajectory confoundersTrajectory length / tool count / task difficulty / user type
Group-aware evaluationComposer-stratified folds
Agent evaluationUser-, task-, domain-, or attack-family-stratified splits
Strict operating pointTPR@1%FPR
Safety operating pointDetect unsafe actions without excessively blocking benign users
What transfers
- Audit the security-relevant parts of a structured sequence.
- Control confounders before claiming an attack works.
- Evaluate at deployment-relevant low false-positive rates.
- Separate exploratory scores from calibrated decisions.
- Use group-aware splits to prevent evaluation leakage.
What does not transfer directly
- Membership inference is not prompt-injection detection.
- Symbolic-music structure is not identical to Agent semantics.
- The paper does not evaluate Tool Use, RAG, Memory, or mobile-device actions.
- Production systems may not expose token log-probabilities.
- The proposed Agent extension is future work.
13
Paper and resources
Verify the primary sources
BibTeX
@InProceedings{pmlr-v303-liu26a,
title = {TS-RaMIA: Membership Inference Attacks for Symbolic Music Generation Models},
author = {Liu, Yuxuan and Sang, Rui and Zhang, Peihong and Li, Zhixin and Zhang, Kunyang and He, Shengyuan and Li, Ye and Xu, Kaiyi and Li, Shengchen},
booktitle = {Proceedings of Machine Learning Research},
pages = {1--15},
year = {2026},
volume = {303},
publisher = {PMLR},
url = {https://proceedings.mlr.press/v303/liu26a.html}
}The audit question
Can a creator test training-set membership?
The candidate was not used for training.
The candidate was used for training.
A statistical signal, calibrated on validation data.
Gray-box access
Forward-pass evidence—without gradients or weights
Teacher forcing under a known tokenizer.
Sampling approximations add query cost and are not the paper’s primary result.
The evaluation trap
A strong raw AUC may be measuring structural complexity
Without confounder control, an apparently successful attack may only be recognizing structural complexity.
Core insight
Membership leakage is localized; uniform averages dilute it
REMI STRUCTURAL MASK
ABC STRUCTURAL MASK
Base observationPer-token NLL
Localized signalStructural tail statistics
Final decisionDirected, standardized out-of-fold meta-attacker
TS-RaMIA method
Mask → NLL → Debias → Tail → Meta-fusion
- Top-k captures sparse leakage pockets.
- Length matching and calibration remove spurious correlations.
- Composer-stratified out-of-fold evaluation reduces stylistic leakage.
Controlled evidence
Performance at the strict low-FPR operating point
— TPR@1%FPR
— TPR@1%FPR
— TPR@1%FPR
Falsification, transfer, and limits
What survived—and what did not
- Top-64 is the best trade-off.
- Windowed p95 adds a small lift.
- Equal-N retains the trend.
- Note-only is unusable without score inversion.
- EVT adds complexity without gain.
— AUC · — TPR@1%FPR
NotaGen under distribution shift—not a direct replication.The method requires per-token probabilities; generation-only API performance is not established.
Conceptual transfer / future direction — not evaluated in this paper
From structured music auditing to structured Agent auditing
Music structureBar / Position / Tempo
Risk-critical Agent eventsTool calls / permission changes / external transmission / file writes
ConfoundersLength / density / composer
Agent confoundersTrajectory length / tool count / task difficulty / user
Strict evaluationTPR@1%FPR
Safety objectiveDetect unsafe actions without overblocking benign users
What transfers
Structure-aware auditing · confounder control · low-FPR evaluation · calibrated decisions · group-aware splits.
What does not transfer directly
Membership inference is not prompt-injection detection. The paper does not evaluate Tool Use, RAG, Memory, or mobile-device actions, and production systems may not expose token log-probabilities.