You MUST consult this skill when choosing parallelism strategies (TP vs PP vs EP vs DP) for model serving, deciding whether speculative decoding helps a workload, designing KV cache management (in-GPU, disaggregated, offloaded), architecting prefill-decode disaggregation, building agentic or multi-turn inference pipelines, optimizing serving cost (quantization, MoE routing, CPU offload), or scaling Ray Serve / vLLM deployments. Also trigger when evaluating hardware for inference (NVLink domains, GB300 racks, CPU-GPU ratios) or debugging serving throughput bottlenecks. NOT for model training, fine-tuning, or RLHF — only the serving and inference path.