Guide the agent through implementing a single Triton kernel that computes `y = rmsnorm(x + residual)` while also writing back `x + residual` for the next transformer block's residual stream. This fusion is the dominant pattern in LLaMA, Mistral, Qwen, and similar decoder blocks: every attention sub-block and every MLP sub-block ends with `residual_add -> rmsnorm`. Done correctly, the kernel saves one full read+write pass over the activation tensor compared to a naive `add` kernel followed by an `rmsnorm` kernel, and removes one launch.