Practical, budget-friendly post-training of small open models (≈0.1–10B) — SFT (full or LoRA/PEFT) then RL with GRPO/DAPO/GSPO using TRL, verifiable+format reward functions, and multi-stage RL to avoid forgetting. Use when fine-tuning a small model on limited GPUs, reproducing a reasoning recipe (SmolLM/SmolLM3/Granite-style), writing TRL GRPOTrainer reward functions, choosing LoRA configs, or sequencing SFT→RL for a compact model.