ICASSP 2026 · Acoustic scene classification

Teach with the signal that matters now.

DDSC replaces a static easy-to-hard curriculum with two signals whose influence changes over training: device invariance first, learning potential later.

Peihong Zhang · Yuxuan Liu · Rui Sang · Zhixin Li · Yiqiang Cai · Yizhou Tan · Shengchen Li

+4.2 pp
mean overall gain · 5% labels · four systems
+3.9 pp
mean unseen-device gain · 5% labels
0
extra inference modules
Method in motion

The curriculum changes what “useful” means over time.

Animated reading: two sample-level signals flow into a cosine-controlled fusion; the focal weight shifts from invariance toward learning progress.

Dynamic dual-signal curriculumPrototype entropy and smoothed loss change are fused by a cosine schedule that changes emphasis between early and late training.TWO SIGNALSTIME-VARYING FUSIONSAMPLE PRIORITYDomain invariancePrototype-posterior entropyHIGH H(x)Learning progressEMA · |loss change|Cosine fusionwᵢ(t) = α(t) · Hᵢ + β(t) · ΔℓᵢEARLY · INVARIANCELATE · PROGRESSWeighted batchArchitecture unchangedNovelty: sample utility is recomputed from the model’s current state—not fixed before training.
Figure · DDSC dynamic weighting mechanismMotion encodes the changing fusion weight; it does not represent additional inference computation.
01 · Question

Why should curriculum order stay fixed?

Under device shift, an example that helps learn transferable structure early may not be the example with the greatest learning value later.

Domain shift

Recorders leave a fingerprint.

Scenes are observed through real and simulated devices, including device types absent from training.

Low labels

Every example must earn its place.

The study tests five label budgets from 5% to 100%, where poor ordering is most costly at the low end.

Static curricula

Difficulty is not utility.

A fixed ranking cannot reflect the model’s changing state or distinguish transferability from current learning progress.

02 · Method

A curriculum that changes its reason for selecting a sample.

Signal A

Estimate invariance

Online device prototypes produce posterior distributions. Higher entropy indicates weaker device specificity.

Signal B

Track progress

An exponential moving average of absolute per-sample loss change estimates learning potential without a second model.

Schedule

Change the balance

A cosine schedule favors invariant examples earlier and progressively admits higher-potential, device-specific cases.

03 · Evidence

The strongest gains appear where the method is intended to help.

At the 5% label budget, all four evaluated systems improve; results are reported over five independent experiments for DDSC variants.

DCASE baseline

48.17%

Overall accuracy with DDSC at 5% labels; unseen-device accuracy is 46.10%.

DS-FlexiNet

55.56%

Overall accuracy with DDSC at 5% labels; unseen-device accuracy is 52.75%.

Han et al.

57.86%

Overall accuracy with DDSC at 5% labels; unseen-device accuracy is 56.42%.

Values reported in Table 2 of the paper. “pp” denotes percentage points.

04 · Boundary

What the experiment establishes—and what it does not.

Evaluation frame

  • DCASE 2024 Task 1, 230,350 one-second clips across ten scenes.
  • Training uses real A/B/C and simulated S1–S3 devices.
  • Testing includes unseen simulated devices S4–S6.

Claim boundary

  • The evidence is acoustic-scene classification under device shift.
  • DDSC changes training weights, not the inference architecture.
  • Generalization beyond this task requires a new controlled study.
Method figure · original paper

Trace both signals at the point where they meet.

The published overview makes the research contribution legible in one view: online prototype entropy measures device invariance, smoothed loss change tracks learning progress, and a scheduler fuses them into per-example weights.

Published DDSC Figure 2: domain invariance and learning-progress signals feed adaptive signal fusion and the weighted objective.
Figure 2 · Overview of DDSC, reproduced from the paper.The scheduler is training-time only; the inference architecture remains unchanged.
Left device prototypes → entropyRight loss change → difficultyCenter cosine schedule → dynamic weights

Inspect the equations, protocol, and full result table.