Quantize a model for on-device (mobile / NPU) deployment and verify it still meets quality — choose fp16 vs int8 (dynamic-range PTQ, full-int8 PTQ with a representative dataset, or QAT), pick per-channel weights and input symmetry, decide which layers stay float, and gate on an accuracy-vs-latency budget.