The problem
The feedback is unbalanced
Share of the combined gradient supplied by the IF loss, initial 4B student
about 1%of response tokens come from IF
up to halfof the pooled log-ratio variance comes from IF
DN-MOPDQwen3.5-9B, 4B and 2BSeptember 2026
TL;DR Label routing decides which teacher supervises a prompt; DN-MOPD also controls how strongly each teacher's feedback counts.
1Nanyang Technological University2Yale University3University of Manchester
01 Overview
Multi-teacher on-policy distillation (MOPD) routes each prompt to the specialist of its domain. DN-MOPD keeps that routing and puts every teacher's feedback on a common scale.
The problem
Share of the combined gradient supplied by the IF loss, initial 4B student
about 1%of response tokens come from IF
up to halfof the pooled log-ratio variance comes from IF
The one-line fix
Estimated on every batch. It keeps label routing and the sign of every advantage; MOPD is the special case in which every multiplier is one.
Cost
No additional teacher call, teacher model or learned router: the operation reuses the rollout log-ratios, computes one pooled std and one per domain, and scales the existing advantages.
The result
+1.17 to +2.36pp
DN-MOPD − Label at student seed 42, with paired 95% bootstrap intervals: every interval is above zero.
At an 8K evaluation cap the gain is +2.47 to +3.08 pp.
Robustness
Three-seed mean gain over Label at 16K, with 95% interval (student seeds 42, 43, 44).
Every seed-matched comparison favors DN-MOPD.
Mathematics
MATH-500 at 16K, DN-MOPD minus Label (points).
At 16K, Label shows no math gain over the initial student at any size; DN-MOPD improves math at every size.
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
02 Method
One initial model; three copies are RL-trained into specialists for math, code and IF, and a fourth copy is the student πu. Every training prompt x carries a domain label d ∈ {math, code, IF}, and Td is that domain's specialist.
Scroll through one training iteration; the diagram follows each step.
1Rollout
The student πu samples a response yi to each prompt xi, and each prompt keeps its domain label di. The teachers will give feedback on these responses, the student's own tokens.
The rollout log-probabilities ℓrolli,t are cached with each response; step 3 reuses them.
Algorithm 1 · line 1
2Teacher scoring
Label routing decides which teacher supervises a prompt: a math response goes to the math specialist, and so on. That teacher's token-level log-probabilities ℓTi,t on the student's own response give a dense distillation signal, the token advantage
A positive At encourages the sampled token; a negative one discourages it. Label-routed MOPD stops here: routing determines where feedback comes from, but not how much it counts.
Algorithm 1 · line 2
3Scale estimationAdded by DN-MOPD
On valid response tokens, ri,t = ℓTi,t − ℓrolli,t. σd is the population std of r over the valid response tokens of domain d in the batch, and σall pools all valid tokens across domains:
| Domain | σd | σall/σd | wd |
|---|---|---|---|
| Math | 0.306 | 2.01 | 2.01 |
| Code | 0.756 | 0.81 | 0.81 |
| IF | 1.798 | 0.34 | 0.34 |
| Pooled | 0.614 | – | – |
For a nondegenerate, unclipped domain, Std(wd rd) = σall: in the diagram each band settles on the dashed pooled spread. wd = 1 if a std is zero or has fewer than two observations.
No 4B multiplier is clipped in this batch; in the 9B batch the IF ratio (0.23) is raised to the 0.25 floor. The IF multiplier sits at the 0.25 floor in most batches, so clipping bounds rather than equalizes that domain's scale. No additional teacher call, teacher model or learned router is needed: one pooled std and one per domain.
Algorithm 1 · lines 3–5
4RescaleAdded by DN-MOPD
The advantage is recomputed from the actor log-probs and multiplied by its domain's weight:
The rule amplifies feedback with a smaller spread and attenuates feedback with a larger spread; in this batch math is scaled by ×2.01 and IF by ×0.34. It is a positive rescaling with no mean subtraction, so it keeps the sign of every advantage.
Algorithm 1 · lines 6–7
5Update
Ãi,t is used in the baseline clipped OPD loss:
Setting every wd = 1 recovers Label, so label routing is the special case. DN-MOPD changes neither which teacher supervises a prompt nor how many prompts each domain receives. The next iteration samples a new batch, and the multipliers are estimated again on every batch.
Algorithm 1 · line 8
Math
×1.5–2.1
The multiplier amplifies math feedback by about 1.5–2.1 at every size.
Code
≈ 1
It keeps code near 1 at 4B and 2B.
IF
0.25 floor
The IF multiplier sits at the 0.25 floor in most batches, so clipping bounds rather than equalizes that domain's scale.
03 Results
DN-MOPD improves the six-task average over MOPD at every size and both evaluation budgets: +1.17 to +2.36 pp at 16K and +2.47 to +3.08 pp at 8K. Every paired interval is above zero.
| Domain | Benchmarks | Answers per question |
|---|---|---|
| Math | AIME25, AIME26 | 64 (avg@64) |
| Code | LiveCodeBench v5, v6 (167/175 disjoint problems) | 6 (avg@6) |
| IF | IFEval, IFBench (strict prompt accuracy) | 16 (avg@16) |
DN-MOPD − Label
Student seed 42 · paired 95% bootstrap intervals (questions resampled within each suite, B = 10,000)
| Size | Eval cap | DN-MOPD | Label | Δ (pp) | 95% interval |
|---|---|---|---|---|---|
| 9B | 16K | 59.56 | 58.39 | +1.17 | [+0.28, +2.03] |
| 9B | 8K | 55.31 | 52.84 | +2.47 | [+1.65, +3.27] |
| 4B | 16K | 52.54 | 50.30 | +2.24 | [+1.30, +3.20] |
| 4B | 8K | 48.56 | 45.48 | +3.08 | [+2.12, +4.06] |
| 2B | 16K | 28.96 | 26.60 | +2.36 | [+1.58, +3.17] |
| 2B | 8K | 27.87 | 24.96 | +2.92 | [+2.02, +3.85] |
| Size | Seed 42 | Seed 43 | Seed 44 | Mean95% CI |
|---|---|---|---|---|
| 16K evaluation cap | ||||
| 9B | +1.17 | +1.08 | +1.10 | +1.12 [+0.62, +1.62] |
| 4B | +2.24 | +1.85 | +1.82 | +1.97 [+1.39, +2.56] |
| 2B | +2.36 | +2.24 | +2.43 | +2.34 [+1.86, +2.84] |
| 8K evaluation cap | ||||
| 9B | +2.47 | +2.62 | +3.05 | +2.71 [+2.16, +3.29] |
| 4B | +3.08 | +2.74 | +3.01 | +2.94 [+2.37, +3.53] |
| 2B | +2.92 | +2.88 | +2.83 | +2.87 [+2.36, +3.40] |
| Size | 16K cap | 8K cap |
|---|---|---|
| 9B | +0.74 [+0.31, +1.18] | +1.45 [+0.92, +2.00] |
| 4B | +1.25 [+0.67, +1.85] | +1.66 [+1.05, +2.29] |
| 2B | +5.95 [+4.95, +6.96] | +6.85 [+5.73, +7.99] |
Label never exceeds the strongest single-teacher student. DN-MOPD exceeds it at every size, but this lead is smaller than the gain over Label, and some intervals include zero. At 9B the lead grows to +2.57 [+1.65, +3.46] when both continue to 160 updates.
SeqKD-SFT and task arithmetic (ParamMerge-TA) retain higher Totals under their own training recipes (Tables 1–2). Bold marks the best result within the multi-teacher OPD block only.
| Method | Math avg@64 | Code avg@6 | IF avg@16 | Total | |||
|---|---|---|---|---|---|---|---|
| AIME25 | AIME26 | LCB v5 | LCB v6 | IFEval | IFBench | ||
| Initial student | 57.7 | 62.6 | 54.9 | 51.4 | 82.4 | 33.8 | 57.1 |
| RL experts | |||||||
| Math expert | 59.7 | 69.4 | 53.4 | 49.4 | 82.3 | 34.9 | 58.2 |
| Code expert | 58.5 | 65.4 | 58.5 | 54.3 | 82.3 | 34.6 | 58.9 |
| IF expert | 54.6 | 60.5 | 54.0 | 51.5 | 86.5 | 42.1 | 58.2 |
| Offline distillation | |||||||
| SeqKD-SFT | 61.3 | 67.6 | 62.7 | 56.0 | 86.0 | 42.4 | 62.6 |
| Parameter merging | |||||||
| ParamMerge-Avg | 59.4 | 68.3 | 54.0 | 50.9 | 84.1 | 36.4 | 58.9 |
| ParamMerge-TA | 59.1 | 68.6 | 55.2 | 50.9 | 86.1 | 43.3 | 60.5 |
| Single-teacher OPD | |||||||
| Math teacher | 58.9 | 69.5 | 53.7 | 49.4 | 82.4 | 34.7 | 58.1 |
| Code teacher | 58.3 | 66.2 | 57.0 | 52.6 | 82.6 | 34.2 | 58.5 |
| IF teacher | 56.8 | 61.9 | 52.1 | 49.2 | 84.7 | 38.9 | 57.3 |
| Multi-teacher OPD | |||||||
| Uniform pool | 57.1 | 65.9 | 53.6 | 52.3 | 83.3 | 35.7 | 58.0 |
| Dynamic router | 56.4 | 63.5 | 55.4 | 51.5 | 84.3 | 37.7 | 58.1 |
| Label-routed MOPD | 55.3 | 63.2 | 56.8 | 52.6 | 84.0 | 38.4 | 58.4 |
| DN-MOPD Ours | 58.9 | 67.7 | 56.3 | 51.4 | 84.5 | 38.7 | 59.6 |
| Method | Qwen3.5-4B | Qwen3.5-2B | ||||||
|---|---|---|---|---|---|---|---|---|
| Math | Code | IF | Total | Math | Code | IF | Total | |
| Initial student | 52.2 | 38.1 | 52.8 | 47.7 | 17.6 | 11.3 | 43.3 | 24.0 |
| RL experts | ||||||||
| Math expert | 54.2 | 39.7 | 54.0 | 49.3 | 22.1 | 13.4 | 44.7 | 26.7 |
| Code expert | 51.6 | 47.4 | 53.6 | 50.8 | 20.7 | 21.6 | 44.4 | 28.9 |
| IF expert | 50.7 | 38.1 | 60.6 | 49.8 | 15.6 | 12.5 | 52.8 | 27.0 |
| Offline distillation | ||||||||
| SeqKD-SFT | 55.6 | 52.3 | 59.7 | 55.9 | 23.1 | 23.9 | 51.6 | 32.9 |
| Parameter merging | ||||||||
| ParamMerge-Avg | 54.5 | 43.1 | 56.2 | 51.2 | 20.7 | 16.3 | 47.5 | 28.2 |
| ParamMerge-TA | 53.8 | 44.4 | 61.4 | 53.2 | 24.4 | 20.4 | 53.3 | 32.7 |
| Single-teacher OPD | ||||||||
| Math teacher | 53.9 | 39.9 | 53.6 | 49.1 | 22.1 | 14.1 | 44.2 | 26.8 |
| Code teacher | 52.2 | 46.8 | 53.5 | 50.9 | 20.4 | 21.2 | 44.9 | 28.8 |
| IF teacher | 50.3 | 39.2 | 57.6 | 49.0 | 16.5 | 11.7 | 49.2 | 25.8 |
| Multi-teacher OPD | ||||||||
| Uniform pool | 52.3 | 41.5 | 55.0 | 49.6 | 19.9 | 14.6 | 46.2 | 26.9 |
| Dynamic router | 49.9 | 44.7 | 57.1 | 50.6 | 17.3 | 13.8 | 48.5 | 26.5 |
| Label-routed MOPD | 50.2 | 43.3 | 57.4 | 50.3 | 17.0 | 13.4 | 49.5 | 26.6 |
| DN-MOPD Ours | 54.0 | 45.4 | 58.2 | 52.5 | 20.7 | 16.3 | 49.9 | 29.0 |
| Variant | Teacher rule | Signal change | 9B | 4B | 2B |
|---|---|---|---|---|---|
| Uniform pool | Mixture | Unchanged | -0.41 | -0.69 | +0.30 |
| Dynamic router | Response | Unchanged | -0.25 | +0.28 | -0.07 |
| Label-routed | Domain | Unchanged | +0.00 | +0.00 | +0.00 |
| Annealed injection | Domain | Early imitation | -0.24 | +1.11 | +0.30 |
| Math ×2 only | Domain | Math scale | — | +1.20 | +0.27 |
| IF ×0.25 only | Domain | IF scale | — | +1.79 | +2.19 |
| DN-MOPD | Domain | Domain scale | +1.17 | +2.24 | +2.36 |
04 Analysis
The experts specialize, their feedback arrives on different scales, and controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone.
Multipliers during training
Per-update multipliers wd of the seed-42 DN-MOPD runs, log scale. The main runs end at update 80; later updates are the continuation.
Drag across the chart, or focus it and use the arrow keys, to scrub through the updates.
| Update | wmath | wcode | wIF | IF before clipping | IF at floor |
|---|---|---|---|---|---|
| Qwen3.5-9B · 125 updates | |||||
| 1 | 1.77 | 0.89 | 0.25 | 0.23 | yes |
| 80 | 1.01 | 1.41 | 0.25 | 0.15 | yes |
| 125 | 1.48 | 1.67 | 0.25 | 0.14 | yes |
| Qwen3.5-4B · 144 updates | |||||
| 1 | 2.01 | 0.81 | 0.34 | 0.34 | no |
| 80 | 2.03 | 0.98 | 0.25 | 0.16 | yes |
| 144 | 1.40 | 1.00 | 0.25 | 0.18 | yes |
| Qwen3.5-2B · 160 updates | |||||
| 1 | 2.10 | 0.82 | 0.43 | 0.43 | no |
| 80 | 2.22 | 0.85 | 0.25 | 0.15 | yes |
| 160 | 1.82 | 0.91 | 0.25 | 0.20 | yes |
Selected updates (first, 80th, last). Every update is in multipliers.json.
Mechanism
IF supplies about 1% of response tokens but up to half of the pooled log-ratio variance. For the initial 4B student, IF accounts for 94% of the combined gradient under equal weights and 64% under DN-MOPD's first-batch weights.
Changing teacher assignment (Uniform pool, Dynamic router) brings no consistent gain.
Amplifying math alone (Math ×2) recovers about half of DN-MOPD's improvement at 4B and little at 2B.
Reducing IF alone (IF ×0.25) recovers most of it and raises math by about three points.
Weights fixed at DN-MOPD's first-batch multipliers, or at a global (2, 1, 0.25), show no detectable difference from DN-MOPD at 9B and 4B.
At 2B, per-batch estimation outperforms first-batch weights by 1.10 points.
Table 4 · Label vs DN-MOPD
The math gains come with shorter answers and fewer responses hitting the generation cap. Mean math response tokens, and the share of all responses reaching the 16K cap.
| Method | Total | Math tokens | At cap (%) |
|---|---|---|---|
| Qwen3.5-9B | |||
| Label-routed | 58.4 | 8,637 | 9.8 |
| DN-MOPD | 59.6 | 7,032 | 5.2 |
| Qwen3.5-4B | |||
| Label-routed | 50.3 | 8,653 | 12.4 |
| DN-MOPD | 52.5 | 6,146 | 5.3 |
| Qwen3.5-2B | |||
| Label-routed | 26.6 | 10,246 | 20.2 |
| DN-MOPD | 29.0 | 7,444 | 13.5 |
The full Table 4, with SeqKD-SFT, ParamMerge-TA and code and IF lengths, is below.
The math gains come with shorter answers and fewer responses hitting the generation cap. At 9B, math tokens drop from 8,637 to 7,032 and the at-cap share from 9.8% to 5.2%.
| Method | Total | Math tokens | Code tokens | IF tokens | At cap (%) |
|---|---|---|---|---|---|
| Qwen3.5-9B | |||||
| SeqKD-SFT | 62.6 | 5,547 | 5,896 | 422 | 1.5 |
| ParamMerge-TA | 60.5 | 5,083 | 5,479 | 430 | 1.7 |
| Label-routed | 58.4 | 8,637 | 6,846 | 450 | 9.8 |
| DN-MOPD | 59.6 | 7,032 | 6,427 | 441 | 5.2 |
| Qwen3.5-4B | |||||
| SeqKD-SFT | 55.9 | 5,534 | 7,389 | 440 | 3.2 |
| ParamMerge-TA | 53.2 | 4,865 | 6,332 | 494 | 2.9 |
| Label-routed | 50.3 | 8,653 | 8,238 | 523 | 12.4 |
| DN-MOPD | 52.5 | 6,146 | 7,649 | 522 | 5.3 |
| Qwen3.5-2B | |||||
| SeqKD-SFT | 32.9 | 6,670 | 9,419 | 822 | 8.9 |
| ParamMerge-TA | 32.7 | 5,383 | 8,315 | 930 | 8.9 |
| Label-routed | 26.6 | 10,246 | 10,047 | 744 | 20.2 |
| DN-MOPD | 29.0 | 7,444 | 10,054 | 706 | 13.5 |
At 160 updates DN-MOPD remains ahead of Label at every size.
| Method | 80 updates | 160 updates | Δ Total |
|---|---|---|---|
| Qwen3.5-9B | |||
| Single teacher (IF) | 57.3 | 57.6 | +0.4 |
| Label-routed | 58.4 | 59.4 | +1.0 |
| DN-MOPD | 59.6 | 60.7 | +1.2 |
| Qwen3.5-4B | |||
| Single teacher (code) | 50.9 | 51.2 | +0.4 |
| Label-routed | 50.3 | 51.6 | +1.3 |
| DN-MOPD | 52.5 | 53.1 | +0.5 |
| Qwen3.5-2B | |||
| Single teacher (code) | 28.8 | 28.8 | -0.1 |
| Label-routed | 26.6 | 28.3 | +1.7 |
| DN-MOPD | 29.0 | 30.0 | +1.1 |
05 Limitations
06 Cite
@article{li2026dnmopd,
title = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
author = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
journal = {arXiv preprint arXiv:2609.35347},
year = {2026},
url = {https://arxiv.org/abs/2609.35347}
}
