Use when the user wants to design, review, refactor, optimize, or productionize high-performance compute kernels such as attention, KV-cache, CUDA, ROCm, Triton, CUTLASS, or fused tensor kernels.