Build LLM streaming UX (SSE/WebSocket) that feels fast, supports cancellation, surfaces tool calls cleanly, and handles partial JSON. Use when wiring an LLM backend to a chat UI, IDE assistant, or any UX where perceived latency matters.