This two-stage curriculum uses uncertainty about device identity as a practical proxy for domain invariance, then reintroduces device-specific examples in a controlled mixture.
Peihong Zhang* · Yuxuan Liu* · Zhixin Li · Rui Sang · Yiqiang Cai · Yizhou Tan · Shengchen Li · *equal contribution
+2.6 pp
Cai_XJTLU unseen devices · 5% labels
50 / 50
median split for curriculum construction
80 / 20
invariant / specific ratio in stage two
Method in motion
Device uncertainty becomes a teaching order.
Animated reading: flat device posteriors produce high entropy, pass the median split as invariant examples, and enter training before the controlled 80:20 refinement stage.
Figure · Entropy-guided curriculum from proxy to scheduleThe bars visualize posterior shape; the split and 80:20 ratio match the paper’s fixed protocol.
01 · Question
Which examples teach a model to generalize across devices?
If an auxiliary classifier cannot confidently identify the recorder from an example, that example may carry more scene information than device identity.
Proxy
Entropy as uncertainty
Shannon entropy over device predictions turns a qualitative idea—domain invariance—into a sortable training signal.
Balance
Do not discard hard cases
Device-specific examples return after the invariant stage, preventing the curriculum from becoming a permanent filter.
Compatibility
Keep the classifier intact
The strategy changes data scheduling rather than the acoustic-scene architecture or its inference path.
02 · Method
Separate curriculum construction from scene learning.
Phase 1
Estimate device posterior
A lightweight two-layer MLP reads frozen encoder features and predicts device identity.
Partition
Median entropy split
The top 50% high-entropy examples form the domain-invariant subset; the rest form the specific subset.
Phase 2
Stage the exposure
Train first on invariant examples, then switch after a validation-loss plateau to fixed 80:20 mixed mini-batches.
03 · Evidence
Three architectures improve at the lowest label budget.
DCASE baseline
44.00 → 46.30
Overall accuracy at 5% labels; unseen devices improve from 42.4 to 44.0.
Cai_XJTLU
48.91 → 51.50
Overall accuracy at 5% labels; unseen devices improve from 46.7 to 49.3.
Han_SJTUTHU
54.35 → 56.60
Overall accuracy at 5% labels; unseen devices improve from 52.7 to 55.2.
Accuracy (%) reported in Table 2. Improvements diminish as labeled data becomes abundant, matching the intended low-resource setting.
04 · Boundary
A transparent proxy with explicit simplifications.
Evaluation frame
DCASE 2024 Task 1, 230,350 one-second clips across ten scenes.
Five label budgets from 5% to 100%.
Seen and unseen-device subsets are reported separately at 5%.
Claim boundary
The 50% split and 80:20 mixture are fixed design choices.
The paper identifies adaptive or weighted schedules as future work.
Entropy is a useful proxy here, not a universal proof of domain invariance.
Method figure · original paper
Turn device uncertainty into a staged learning plan.
The paper’s framework separates curriculum construction from scene learning: a lightweight device classifier produces an entropy ranking, then the ASC model learns first from the high-entropy subset before controlled refinement.
Figure 2 · Entropy-guided curriculum construction, reproduced from the paper.High-entropy examples form the domain-invariant subset used to anchor Stage 1.
Phase 1 predict device identityPartition median split of entropyPhase 2 invariant-first, then 80:20 refinement
Inspect the curriculum and complete benchmark table.