LLM/model inference serving architecture: the engine → serving → orchestration
layering (vLLM/SGLang/TensorRT-LLM, Triton, KServe/Ray Serve), KV-cache &
continuous batching, prefill-decode disaggregation, and scaling. Architect-level
topology, not model training.
USE WHEN: designing model/LLM serving infra, "vLLM", "SGLang", "TensorRT-LLM",
"Triton", "KServe", "Ray Serve", "continuous batching", "KV cache", "prefill
decode", "TTFT", multi-GPU/multi-model serving, inference autoscaling.
DO NOT USE FOR: on-device (use `edge-inference`); provider routing (use
`model-gateway-routing`); RAG app logic (use rag/rag-frameworks skills).