Guide the agent through implementing a Triton kernel that writes newly computed K and V tensors into a pre-allocated KV cache during LLM inference. This covers two cache layouts (contiguous and paged / vLLM-style PagedAttention), unified prefill and decode handling via a `slot_mapping` tensor, GQA/MQA where the cache stores fewer heads than Q, optional fp8/int8 quantized KV cache with scaling, coalesced versus scattered write patterns, and the boundary checks required to avoid corrupting other requests' cache regions. The kernel runs once per layer per forward step; correctness is the dominant concern, throughput is secondary because the data volume per call is small.