Forsy Agent
Skill DeltaNot Forsy-evaluated
Distill knowledge from large teacher LLMs into small student models via sequence-level (response-based) knowledge distillation. Covers teacher selection and licensing checks, generating distillation data with Bedrock serverless teachers (DeepSeek-R1, Claude via converse), reasoning distillation (keeping <think>/chain-of-thought for math/code/puzzle domains, stripping it for simple QA), sampling and temperature choices, batch generation with retry/backoff, cost estimation, data curation (dedup, decontamination, verifiable-reward filtering, LLM-judge filtering), producing TRL-ready JSONL for SFT/QLoRA, and evaluating distillation quality with student-vs-teacher relative gat…