DN-MOPDQwen3.5-9B, 4B and 2BSeptember 2026

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

TL;DR Label routing decides which teacher supervises a prompt; DN-MOPD also controls how strongly each teacher's feedback counts.

Xin Li1Hao Jiang1Xin Gao2Annan Wang1Yuchen Xie1Jinghao Guo1Xingwei Qu3Yichi ZhangChau Yuen1

1Nanyang Technological University2Yale University3University of Manchester

Schematic of DN-MOPD: three teachers' feedback rescaled onto a common scale for one shared student Three teacher stations, math, code and instruction following (IF), emit token-level teacher–student log-ratios on different scales: small for math, moderate for code, large and jittery for IF. The signals flow into a DN-MOPD normalizer that applies w_d = clip(σ_all / σ_d, 0.25, 4) per domain: math is amplified, code stays near 1, and IF is attenuated down to the 0.25 floor, so it is bounded rather than fully equalized. The rescaled signals keep their signs and feed one shared student. Schematic, not to scale. raw log-ratios rescaled Σ Math Code IF DN-MOPD wd=clip( σall σd , 0.25, 4) amplify ×1.5–2.1 rescale ≈1 attenuate 0.25 floor 0.2514wd Shared student domain spreadσd pooled spreadσall rescaled advantageÃt=wd·At clip bounds0.25 ≤ wd ≤ 4
Schematic, not to scale
  • 3 → 1specialists (math, code, IF) teach one shared student
  • 3 sizesQwen3.5-9B, 4B and 2B, each with its own expert pool
  • 6 benchmarksAIME25, AIME26 · LiveCodeBench v5, v6 · IFEval, IFBench
  • 80 updatesper student, with an 8,192-token training response cap
Figure 1. Panel (a): teacher–student log-ratios for math, code and IF drawn on different scales, with IF the most spread out. Panel (b): DN-MOPD box that amplifies math, rescales code and attenuates IF with w_d = clip(σ_all/σ_d, 0.25, 4) before feeding a shared student. Panel (c): horizontal bars of the six-task Total gain over label routing: 9B +2.47 at 8K and +1.17 at 16K; 4B +3.08 and +2.24; 2B +2.92 and +2.36.
Figure 1 Label routing decides which teacher supervises a prompt; DN-MOPD also controls how strongly each teacher's feedback counts. (a) Teacher–student log-ratios differ in scale across domains (schematic). (b) DN-MOPD rescales each domain's feedback with a bounded, sign-preserving multiplier (schematic). (c) Six-task Total gain of DN-MOPD over label routing. Full size

01 Overview

Routing determines where feedback comes from, but not how much it counts.

Multi-teacher on-policy distillation (MOPD) routes each prompt to the specialist of its domain. DN-MOPD keeps that routing and puts every teacher's feedback on a common scale.

The problem

The feedback is unbalanced

2.3–4.4× Instruction-following log-ratios are 2.3–4.4 times as dispersed as the pooled signal in the first training batch, and mathematics log-ratios about half as dispersed.

Share of the combined gradient supplied by the IF loss, initial 4B student

Equal weights 94%
DN-MOPD first-batch weights 64%

about 1%of response tokens come from IF

up to halfof the pooled log-ratio variance comes from IF

The one-line fix

Rescale each domain by its measured spread

wd = clip(σall / σd, 0.25, 4)
Ãt = wd · At

Estimated on every batch. It keeps label routing and the sign of every advantage; MOPD is the special case in which every multiplier is one.

Cost

No extra teacher calls

No additional teacher call, teacher model or learned router: the operation reuses the rollout log-ratios, computes one pooled std and one per domain, and scales the existing advantages.

  • Same teacher per prompt
  • Same prompts per domain
  • Signs preserved

The result

Six-task Total over MOPD, 16K evaluation cap

+1.17 to +2.36pp

  • 9B+1.17
  • 4B+2.24
  • 2B+2.36

DN-MOPD − Label at student seed 42, with paired 95% bootstrap intervals: every interval is above zero.

At an 8K evaluation cap the gain is +2.47 to +3.08 pp.

Robustness

3 seeds, every size

Three-seed mean gain over Label at 16K, with 95% interval (student seeds 42, 43, 44).

9B+1.12[+0.62, +1.62]
4B+1.97[+1.39, +2.56]
2B+2.34[+1.86, +2.84]

Every seed-matched comparison favors DN-MOPD.

Mathematics

Math recovered

MATH-500 at 16K, DN-MOPD minus Label (points).

9B+0.74
4B+1.25
2B+5.95

At 16K, Label shows no math gain over the initial student at any size; DN-MOPD improves math at every size.

Abstract

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.

Teacher assignment alone does not transfer the specialists' skills.

  • MOPD with label routing does not outperform the strongest single-teacher student at any of three Qwen3.5 sizes.
  • It transfers little of the mathematics expert's gain.
  • Its domains give feedback on unequal scales. In the first training batch, instruction-following log-ratios are 2.3–4.4 times as dispersed as the pooled signal, and mathematics log-ratios about half as dispersed.
  • For the initial 4B student, the instruction-following loss supplies 94% of the combined gradient.

DN-MOPD puts every teacher's feedback on a common scale.

  • It rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread, estimated on every batch.
  • It keeps label routing and the sign of every advantage.
  • MOPD is the special case in which every multiplier is one.

Calibrating feedback scale consistently improves on MOPD.

  • DN-MOPD improves the six-task average over MOPD at every size and both evaluation budgets.
  • Every paired interval is above zero, and the gain holds across three student seeds.
  • It also exceeds the strongest single-teacher student and recovers most of the mathematics gain that MOPD loses.

02 Method

DN-MOPD puts every teacher's feedback on a common scale.

One initial model; three copies are RL-trained into specialists for math, code and IF, and a fourth copy is the student πu. Every training prompt x carries a domain label d ∈ {math, code, IF}, and Td is that domain's specialist.

Scroll through one training iteration; the diagram follows each step.

One DN-MOPD training iteration, in five steps Step 1, rollout: labelled prompts (x, d) for math, code and IF go to the student, which samples a response to each; the rollout log-probabilities are cached. Step 2, teacher scoring: each response is routed by its label to the frozen math, code or IF teacher, giving teacher log-probabilities for the same response. Step 3, scale estimation: r is the teacher log-probability minus the cached rollout log-probability; its spread is measured per domain and pooled, and each domain multiplier is the clipped ratio of pooled to domain spread. In the first batch of the 4B DN-MOPD run (seed 42) the spreads are 0.306 for math, 0.756 for code, 1.798 for IF and 0.614 pooled, giving multipliers 2.01, 0.81 and 0.34. Step 4, rescale: the actor-side advantage is multiplied by its domain multiplier, which keeps every sign. Step 5, update: the clipped OPD objective updates the student and the next iteration begins; label routing is the case where every multiplier is 1. Tokens, tick marks and bar heights are schematic. DN-MOPD 1ROLLOUT Labelled prompts (x, d) d = math d = code d = IF Student πu yi ~ πu(· | xi) ℓroll cached 2TEACHER SCORING Math teacher Code teacher IF teacher frozen · one per response Scored batch · same responses ℓT ℓroll 3SCALE ESTIMATION r = ℓT − ℓroll raw ±σd · dashed ±σall after wd: ±σall σd wd Math 0.306 Code 0.756 IF 1.798 Pooled 0.614 σall wd = clip(σall / σd, 0.25, 4) first batch · 4B DN-MOPD run · seed 42 ×2.01 ×0.81 ×0.34 4RESCALE A = ℓT − ℓactor from actor re-evaluation à = stopgrad(wd · A) before: A, every wd = 1 dashed A · solid à signs kept · bars schematic 5STUDENT UPDATE Clipped OPD objective −min{ ρ Ã, clip(ρ, 1 − η−, 1 + η+) à } Update student πu Label routing is the special case: every wd = 1 Next iteration
One training iteration in five steps Schematic, except σ and w in step 3: first batch, 4B DN-MOPD run, seed 42.
  1. 1Rollout

    The student answers labelled prompts

    The student πu samples a response yi to each prompt xi, and each prompt keeps its domain label di. The teachers will give feedback on these responses, the student's own tokens.

    The rollout log-probabilities ℓrolli,t are cached with each response; step 3 reuses them.

    Algorithm 1 · line 1

    1. 1
      Sample responses yi ~ πu(· | xi).
  2. 2Teacher scoring

    Each response is scored by the frozen teacher of its domain

    Label routing decides which teacher supervises a prompt: a math response goes to the math specialist, and so on. That teacher's token-level log-probabilities ℓTi,t on the student's own response give a dense distillation signal, the token advantage

    At = log pT_d(yt | ht) − log πu(yt | ht),
    ht = (x, y<t)

    A positive At encourages the sampled token; a negative one discourages it. Label-routed MOPD stops here: routing determines where feedback comes from, but not how much it counts.

    Algorithm 1 · line 2

    1. 2
      Score each yi with its domain teacher Td_i.
  3. 3Scale estimationAdded by DN-MOPD

    Measure how spread out each domain's feedback is

    On valid response tokens, ri,t = ℓTi,t − ℓrolli,t. σd is the population std of r over the valid response tokens of domain d in the batch, and σall pools all valid tokens across domains:

    wd = clip(σall / σd, 0.25, 4)
    First batch, 4B DN-MOPD run, seed 42
    Domainσdσall/σdwd
    Math0.3062.012.01
    Code0.7560.810.81
    IF1.7980.340.34
    Pooled0.614––

    For a nondegenerate, unclipped domain, Std(wd rd) = σall: in the diagram each band settles on the dashed pooled spread. wd = 1 if a std is zero or has fewer than two observations.

    No 4B multiplier is clipped in this batch; in the 9B batch the IF ratio (0.23) is raised to the 0.25 floor. The IF multiplier sits at the 0.25 floor in most batches, so clipping bounds rather than equalizes that domain's scale. No additional teacher call, teacher model or learned router is needed: one pooled std and one per domain.

    Algorithm 1 · lines 3–5

    1. 3
      ri,t ← ℓTi,t − ℓrolli,t on valid tokens.
    2. 4
      σall ← Std({ri,t}).
    3. 5
      For each domain d:
      • σd ← Std({ri,t : di = d})
      • wd ← clip(σall/σd, 0.25, 4)
      • use wd = 1 if either statistic is degenerate.
  4. 4RescaleAdded by DN-MOPD

    Rescale each domain's advantages, keeping every sign

    The advantage is recomputed from the actor log-probs and multiplied by its domain's weight:

    Ai,t = ℓTi,t − ℓactori,t,
    Ãi,t = stopgrad(wd_i Ai,t)

    The rule amplifies feedback with a smaller spread and attenuates feedback with a larger spread; in this batch math is scaled by ×2.01 and IF by ×0.34. It is a positive rescaling with no mean subtraction, so it keeps the sign of every advantage.

    Algorithm 1 · lines 6–7

    1. 6
      Recompute Ai,t ← ℓTi,t − ℓactori,t.
    2. 7
      Ãi,t ← stopgrad(wd_i Ai,t).
  5. 5Update

    Update the student with the clipped OPD objective

    Ãi,t is used in the baseline clipped OPD loss:

    Li,t(θ) = −min{ ρi,t(θ) Ãi,t, clip(ρi,t(θ), 1−η−, 1+η+) Ãi,t }
    ρi,t(θ) = πθ(yi,t | hi,t) / exp(ℓactori,t)

    Setting every wd = 1 recovers Label, so label routing is the special case. DN-MOPD changes neither which teacher supervises a prompt nor how many prompts each domain receives. The next iteration samples a new batch, and the multipliers are estimated again on every batch.

    Algorithm 1 · line 8

    1. 8
      Update πu on à with the clipped OPD objective.

What the rule does in practice

Math

×1.5–2.1

The multiplier amplifies math feedback by about 1.5–2.1 at every size.

Code

≈ 1

It keeps code near 1 at 4B and 2B.

IF

0.25 floor

The IF multiplier sits at the 0.25 floor in most batches, so clipping bounds rather than equalizes that domain's scale.

Figure 2. Four stages of one DN-MOPD iteration: (1) the student samples a response to a labeled prompt; (2) the math, code or IF teacher of the prompt's domain scores it, giving teacher and rollout log-probabilities; (3) inside the DN-MOPD box, r = ℓ_T − ℓ_roll gives σ_d per domain and σ_all over all valid tokens, which set w_d = clip(σ_all/σ_d, 0.25, 4), and the actor advantage A = ℓ_T − ℓ_actor is rescaled to à = stopgrad(w_d A); (4) the clipped OPD objective updates the student, and the loop repeats.
Full diagram · Figure 2 One DN-MOPD training iteration. (1) The student samples responses to labeled prompts. (2) Each response is scored by the frozen teacher of its domain. (3) Each domain's multiplier is estimated from cached rollout log-ratios and applied to the actor-side advantages. (4) The student is updated with the clipped OPD objective. Label routing is the special case in which every multiplier is 1. Full size

03 Results

Calibrating feedback scale consistently improves on MOPD.

DN-MOPD improves the six-task average over MOPD at every size and both evaluation budgets: +1.17 to +2.36 pp at 16K and +2.47 to +3.08 pp at 8K. Every paired interval is above zero.

Setup

  • Models Qwen3.5-9B, 4B and 2B, each with an independently trained expert pool (math, code, IF); students learn from teachers of their own size.
  • Training 80 updates, student seed 42 (seeds 43 and 44 are controls), and an 8,192-token training response cap.
  • Evaluation Total is the mean of the six task scores. Main tables use a 16,384-token evaluation cap; the 8K results are in the paper appendix. Paired 95% intervals come from a question-level bootstrap, conditional on the student seed and teacher pool.

Benchmarks

DomainBenchmarksAnswers per question
MathAIME25, AIME2664 (avg@64)
CodeLiveCodeBench v5, v6 (167/175 disjoint problems)6 (avg@6)
IFIFEval, IFBench (strict prompt accuracy)16 (avg@16)

Baselines

  • The initial student and the three RL experts; single-teacher OPD students.
  • Uniform pool (averages teacher probabilities); Dynamic router (selects a teacher from the student's response); Label-routed MOPD.
  • SeqKD-SFT (offline distillation on teacher answers); ParamMerge-Avg and ParamMerge-TA (task arithmetic, λ = 1).

DN-MOPD − Label

Six-task Total gain over label-routed MOPD, per size

Student seed 42 · paired 95% bootstrap intervals (questions resampled within each suite, B = 10,000)

9B16K+1.17[+0.28, +2.03]
8K+2.47[+1.65, +3.27]
4B16K+2.24[+1.30, +3.20]
8K+3.08[+2.12, +4.06]
2B16K+2.36[+1.58, +3.17]
8K+2.92[+2.02, +3.85]
Show as table
SizeEval capDN-MOPDLabelΔ (pp)95% interval
9B16K59.5658.39+1.17[+0.28, +2.03]
9B8K55.3152.84+2.47[+1.65, +3.27]
4B16K52.5450.30+2.24[+1.30, +3.20]
4B8K48.5645.48+3.08[+2.12, +4.06]
2B16K28.9626.60+2.36[+1.58, +3.17]
2B8K27.8724.96+2.92[+2.02, +3.85]
Three seeds

DN-MOPD − Label per student seed

SizeSeed 42Seed 43Seed 44Mean95% CI
16K evaluation cap
9B+1.17+1.08+1.10+1.12 [+0.62, +1.62]
4B+2.24+1.85+1.82+1.97 [+1.39, +2.56]
2B+2.36+2.24+2.43+2.34 [+1.86, +2.84]
8K evaluation cap
9B+2.47+2.62+3.05+2.71 [+2.16, +3.29]
4B+3.08+2.74+3.01+2.94 [+2.37, +3.53]
2B+2.92+2.88+2.83+2.87 [+2.36, +3.40]
DN-MOPD minus Label, six-task Total (pp), per student seed (42, 43, 44) and the three-seed mean with a two-level bootstrap 95% interval.
MATH-500

Mathematics carries the largest gains

Size16K cap8K cap
9B+0.74 [+0.31, +1.18]+1.45 [+0.92, +2.00]
4B+1.25 [+0.67, +1.85]+1.66 [+1.05, +2.29]
2B+5.95 [+4.95, +6.96]+6.85 [+5.73, +7.99]
MATH-500 (16 answers per question), DN-MOPD minus Label (pp) with paired 95% intervals. At 16K, Label shows no math gain over the initial student at any size, whereas DN-MOPD improves math at every size.

The single-teacher reference

Label never exceeds the strongest single-teacher student. DN-MOPD exceeds it at every size, but this lead is smaller than the gain over Label, and some intervals include zero. At 9B the lead grows to +2.57 [+1.65, +3.46] when both continue to 160 updates.

Not the best recipe overall

SeqKD-SFT and task arithmetic (ParamMerge-TA) retain higher Totals under their own training recipes (Tables 1–2). Bold marks the best result within the multi-teacher OPD block only.

Table 1

Capability integration at 9B

MethodMath avg@64Code avg@6IF avg@16Total
AIME25AIME26LCB v5LCB v6IFEvalIFBench
Initial student57.762.654.951.482.433.857.1
RL experts
Math expert59.769.453.449.482.334.958.2
Code expert58.565.458.554.382.334.658.9
IF expert54.660.554.051.586.542.158.2
Offline distillation
SeqKD-SFT61.367.662.756.086.042.462.6
Parameter merging
ParamMerge-Avg59.468.354.050.984.136.458.9
ParamMerge-TA59.168.655.250.986.143.360.5
Single-teacher OPD
Math teacher58.969.553.749.482.434.758.1
Code teacher58.366.257.052.682.634.258.5
IF teacher56.861.952.149.284.738.957.3
Multi-teacher OPD
Uniform pool57.165.953.652.383.335.758.0
Dynamic router56.463.555.451.584.337.758.1
Label-routed MOPD55.363.256.852.684.038.458.4
DN-MOPD Ours58.967.756.351.484.538.759.6
Capability integration at 9B. Qwen3.5-9B scores (%) on six public tasks, 16K evaluation cap. Total averages the six task scores. Bold: best within the multi-teacher OPD block.
Table 2

Capability integration at smaller scales

MethodQwen3.5-4BQwen3.5-2B
MathCodeIFTotalMathCodeIFTotal
Initial student52.238.152.847.717.611.343.324.0
RL experts
Math expert54.239.754.049.322.113.444.726.7
Code expert51.647.453.650.820.721.644.428.9
IF expert50.738.160.649.815.612.552.827.0
Offline distillation
SeqKD-SFT55.652.359.755.923.123.951.632.9
Parameter merging
ParamMerge-Avg54.543.156.251.220.716.347.528.2
ParamMerge-TA53.844.461.453.224.420.453.332.7
Single-teacher OPD
Math teacher53.939.953.649.122.114.144.226.8
Code teacher52.246.853.550.920.421.244.928.8
IF teacher50.339.257.649.016.511.749.225.8
Multi-teacher OPD
Uniform pool52.341.555.049.619.914.646.226.9
Dynamic router49.944.757.150.617.313.848.526.5
Label-routed MOPD50.243.357.450.317.013.449.526.6
DN-MOPD Ours54.045.458.252.520.716.349.929.0
Capability integration at smaller scales. Independently trained Qwen3.5-4B and 2B expert pools, 16K evaluation cap. Each domain averages its two tasks; Total weights the three domains equally. Bold: best within the multi-teacher OPD block at each size.
Table 3

Changes to label-routed OPD

VariantTeacher ruleSignal change9B4B2B
Uniform poolMixtureUnchanged-0.41-0.69+0.30
Dynamic routerResponseUnchanged-0.25+0.28-0.07
Label-routedDomainUnchanged+0.00+0.00+0.00
Annealed injectionDomainEarly imitation-0.24+1.11+0.30
Math ×2 onlyDomainMath scale—+1.20+0.27
IF ×0.25 onlyDomainIF scale—+1.79+2.19
DN-MOPDDomainDomain scale+1.17+2.24+2.36
Changes to label-routed OPD. Total point differences (pp) relative to Label, evaluated at 16K. Pooling and Dynamic change the teacher assignment; annealed injection and DN-MOPD retain it and change the learning signal. The two fixed-weight rows set one domain's weight (mathematics 2 or IF 0.25, others 1) and were run at 4B and 2B only.

04 Analysis

Where the gain comes from

The experts specialize, their feedback arrives on different scales, and controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone.

Figure 3. Three heatmaps, one per Qwen3.5 size (9B, 4B, 2B), with rows for the math, code and IF experts and columns for the math, code and IF evaluation domains; the diagonal cells, where each expert is evaluated in its own domain, are outlined and show that expert's gain over the shared initialization.
Figure 3 Expert specialization across domains. Each cell is the gain over the shared initialization (pp) at 16K, averaged over the two tasks in the evaluation domain. Rows identify experts and columns identify evaluation domains. Full size
Figure 4. (a) Relative log-ratio spread per domain at each size: IF far above the pooled spread and math below it. (b) Bars of the DN-MOPD minus Label domain accuracy at 16K for math, code and IF at 9B, 4B and 2B, with math the largest at every size. (c) Multipliers over training for each size: math above one, code near one, and IF held at the 0.25 floor with its unclipped value below the floor.
Figure 4 Feedback scales, domain gains and training multipliers. (a) Median relative log-ratio spread with interquartile ranges over the recorded batches of the DN-MOPD runs. (b) Domain accuracy differences between DN-MOPD and Label at 16K. (c) Raw multipliers (faint), trailing 5-update means (solid), and IF before clipping (dotted). Full size

Multipliers during training

Math is amplified; IF sits at the 0.25 floor

Per-update multipliers wd of the seed-42 DN-MOPD runs, log scale. The main runs end at update 80; later updates are the continuation.

Drag across the chart, or focus it and use the arrow keys, to scrub through the updates.

Show as table
UpdatewmathwcodewIFIF before clippingIF at floor
Qwen3.5-9B · 125 updates
11.770.890.250.23yes
801.011.410.250.15yes
1251.481.670.250.14yes
Qwen3.5-4B · 144 updates
12.010.810.340.34no
802.030.980.250.16yes
1441.401.000.250.18yes
Qwen3.5-2B · 160 updates
12.100.820.430.43no
802.220.850.250.15yes
1601.820.910.250.20yes

Selected updates (first, 80th, last). Every update is in multipliers.json.

Mechanism

IF supplies about 1% of response tokens but up to half of the pooled log-ratio variance. For the initial 4B student, IF accounts for 94% of the combined gradient under equal weights and 64% under DN-MOPD's first-batch weights.

Fixed-weight controls

  1. Teacher assignment

    Changing teacher assignment (Uniform pool, Dynamic router) brings no consistent gain.

  2. Amplify math alone

    Amplifying math alone (Math ×2) recovers about half of DN-MOPD's improvement at 4B and little at 2B.

  3. Reduce IF alone

    Reducing IF alone (IF ×0.25) recovers most of it and raises math by about three points.

  4. Fixed weights near DN-MOPD's

    Weights fixed at DN-MOPD's first-batch multipliers, or at a global (2, 1, 0.25), show no detectable difference from DN-MOPD at 9B and 4B.

  5. Per-batch estimation

    At 2B, per-batch estimation outperforms first-batch weights by 1.10 points.

Table 4 · Label vs DN-MOPD

Shorter math answers, fewer responses at the cap

The math gains come with shorter answers and fewer responses hitting the generation cap. Mean math response tokens, and the share of all responses reaching the 16K cap.

Math tokens mean per response

At the 16K cap % of responses

Show as table
MethodTotalMath tokensAt cap (%)
Qwen3.5-9B
Label-routed58.48,6379.8
DN-MOPD59.67,0325.2
Qwen3.5-4B
Label-routed50.38,65312.4
DN-MOPD52.56,1465.3
Qwen3.5-2B
Label-routed26.610,24620.2
DN-MOPD29.07,44413.5

The full Table 4, with SeqKD-SFT, ParamMerge-TA and code and IF lengths, is below.

Table 4

Shorter answers

The math gains come with shorter answers and fewer responses hitting the generation cap. At 9B, math tokens drop from 8,637 to 7,032 and the at-cap share from 9.8% to 5.2%.

MethodTotalMath tokensCode tokensIF tokensAt cap (%)
Qwen3.5-9B
SeqKD-SFT62.65,5475,8964221.5
ParamMerge-TA60.55,0835,4794301.7
Label-routed58.48,6376,8464509.8
DN-MOPD59.67,0326,4274415.2
Qwen3.5-4B
SeqKD-SFT55.95,5347,3894403.2
ParamMerge-TA53.24,8656,3324942.9
Label-routed50.38,6538,23852312.4
DN-MOPD52.56,1467,6495225.3
Qwen3.5-2B
SeqKD-SFT32.96,6709,4198228.9
ParamMerge-TA32.75,3838,3159308.9
Label-routed26.610,24610,04774420.2
DN-MOPD29.07,44410,05470613.5
Capability and generated length. Mean response tokens by domain and the share of responses reaching the 16K cap. Scores and lengths come from the same evaluation outputs.
Table 5

Longer training

At 160 updates DN-MOPD remains ahead of Label at every size.

Method80 updates160 updatesΔ Total
Qwen3.5-9B
Single teacher (IF)57.357.6+0.4
Label-routed58.459.4+1.0
DN-MOPD59.660.7+1.2
Qwen3.5-4B
Single teacher (code)50.951.2+0.4
Label-routed50.351.6+1.3
DN-MOPD52.553.1+0.5
Qwen3.5-2B
Single teacher (code)28.828.8-0.1
Label-routed26.628.3+1.7
DN-MOPD29.030.0+1.1
Training-duration comparison. Six-task Total at fixed 80- and 160-update endpoints, evaluated at 16K. Changes use unrounded scores.

05 Limitations

What the evidence does not show

Scope of the claims

  • DN-MOPD is not the best integration recipe overall: SeqKD-SFT and task arithmetic (ParamMerge-TA) retain higher Totals under their own training recipes (Tables 1–2).
  • At 9B, code does not improve at 16K. Gains in code and IF are smaller and less consistent.
  • The advantage over the strongest single-teacher student has some intervals that include zero.

Limitations

  • There is one expert pool per size within one model family, and the student seeds do not cover retraining experts.
  • DN-MOPD was selected on an earlier Qwen3 development instrument, where an earlier Qwen3-4B comparison under a different setup found no clear gain, so benefits depend on the teacher–student configuration.
  • The clipping bounds were not tuned per size, and the IF multiplier usually sits at the lower bound.
  • Fixed weights near DN-MOPD's multipliers perform comparably at 9B and 4B, so per-batch re-estimation adds to the gain only at 2B.

06 Cite

Citation

li2026dnmopd.bib
@article{li2026dnmopd,
  title   = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
  author  = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
  journal = {arXiv preprint arXiv:2609.35347},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.35347}
}

Thumbnail of the first page of the DN-MOPD paper. Paper · PDF Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation Read the paper