Guide the agent through implementing a Triton kernel that unpacks and dequantizes a quantized weight tensor (int4 or int8) into fp16 or bf16. This is the standalone building block underneath W4A16 / W8A16 schemes (AWQ, GPTQ, SqueezeLLM, bitsandbytes NF4, int8 per-channel). Covers bit-unpacking, per-group scale/zero arithmetic, NF4 codebook lookup, and — critically — when *not* to write a standalone dequant kernel because the operation should be fused into the matmul instead.