DCASE 2025 · Co-first author

Begin with examples that forget the recorder.

This two-stage curriculum uses uncertainty about device identity as a practical proxy for domain invariance, then reintroduces device-specific examples in a controlled mixture.

Peihong Zhang* · Yuxuan Liu* · Zhixin Li · Rui Sang · Yiqiang Cai · Yizhou Tan · Shengchen Li · *equal contribution

+2.6 pp
Cai_XJTLU unseen devices · 5% labels
50 / 50
median split for curriculum construction
80 / 20
invariant / specific ratio in stage two
Method in motion

Device uncertainty becomes a teaching order.

Animated reading: flat device posteriors produce high entropy, pass the median split as invariant examples, and enter training before the controlled 80:20 refinement stage.

Entropy-guided curriculum constructionA frozen encoder and auxiliary device classifier estimate entropy, divide examples into invariant and specific subsets, and train the scene classifier in two stages.1 · ESTIMATE DEVICE UNCERTAINTYAuxiliary device classifierFrozen audio features · 2-layer MLPHIGH ENTROPY · INVARIANTLOW ENTROPY · DEVICE-SPECIFIC2 · MEDIAN SPLITRank H(x)TOP 50% · XINVBOTTOM 50% · XSPEC3 · STAGED TRAININGScene classifierSame architecture · different exposureSTAGE 1 · XINV ONLYLEARN GENERALIZABLE FEATURESSTAGE 2 · 80:20 MIXNovelty: entropy supplies a domain-invariance proxy without modifying the deployed ASC model.
Figure · Entropy-guided curriculum from proxy to scheduleThe bars visualize posterior shape; the split and 80:20 ratio match the paper’s fixed protocol.
01 · Question

Which examples teach a model to generalize across devices?

If an auxiliary classifier cannot confidently identify the recorder from an example, that example may carry more scene information than device identity.

Proxy

Entropy as uncertainty

Shannon entropy over device predictions turns a qualitative idea—domain invariance—into a sortable training signal.

Balance

Do not discard hard cases

Device-specific examples return after the invariant stage, preventing the curriculum from becoming a permanent filter.

Compatibility

Keep the classifier intact

The strategy changes data scheduling rather than the acoustic-scene architecture or its inference path.

02 · Method

Separate curriculum construction from scene learning.

Phase 1

Estimate device posterior

A lightweight two-layer MLP reads frozen encoder features and predicts device identity.

Partition

Median entropy split

The top 50% high-entropy examples form the domain-invariant subset; the rest form the specific subset.

Phase 2

Stage the exposure

Train first on invariant examples, then switch after a validation-loss plateau to fixed 80:20 mixed mini-batches.

03 · Evidence

Three architectures improve at the lowest label budget.

DCASE baseline

44.00 → 46.30

Overall accuracy at 5% labels; unseen devices improve from 42.4 to 44.0.

Cai_XJTLU

48.91 → 51.50

Overall accuracy at 5% labels; unseen devices improve from 46.7 to 49.3.

Han_SJTUTHU

54.35 → 56.60

Overall accuracy at 5% labels; unseen devices improve from 52.7 to 55.2.

Accuracy (%) reported in Table 2. Improvements diminish as labeled data becomes abundant, matching the intended low-resource setting.

04 · Boundary

A transparent proxy with explicit simplifications.

Evaluation frame

  • DCASE 2024 Task 1, 230,350 one-second clips across ten scenes.
  • Five label budgets from 5% to 100%.
  • Seen and unseen-device subsets are reported separately at 5%.

Claim boundary

  • The 50% split and 80:20 mixture are fixed design choices.
  • The paper identifies adaptive or weighted schedules as future work.
  • Entropy is a useful proxy here, not a universal proof of domain invariance.
Method figure · original paper

Turn device uncertainty into a staged learning plan.

The paper’s framework separates curriculum construction from scene learning: a lightweight device classifier produces an entropy ranking, then the ASC model learns first from the high-entropy subset before controlled refinement.

Published entropy-guided curriculum Figure 2 showing device posterior entropy, sample partitioning, and two-stage acoustic scene classification training.
Figure 2 · Entropy-guided curriculum construction, reproduced from the paper.High-entropy examples form the domain-invariant subset used to anchor Stage 1.
Phase 1 predict device identityPartition median split of entropyPhase 2 invariant-first, then 80:20 refinement

Inspect the curriculum and complete benchmark table.