Combine RL and distillation to consolidate capabilities — on-policy distillation (per-token teacher feedback on the student's own rollouts, reverse-KL), multi-teacher on-policy distillation (MOPD), off-policy SFT distillation trade-offs, self-distillation variants, and combining dense distillation signal with verifier rewards. Use when merging multiple RL domain experts into one model, choosing SFT-distillation vs on-policy distillation, cutting RL cost, or recovering capabilities lost during domain adaptation.