ATraining's method for merging three RL specialist teachers (STEM, agentic, helpfulness/safety) into a single model via trace-distillation SFT then a final lightweight RL stage. Use when combining multiple domain-specialist models into one, balancing sample vs token weighting across capabilities, or running a final RL pass that preserves reasoning while improving safety/style.