Live-system triage — read-only evidence first, establish what/since-when/how-wide, correlate onset with changes, trace one failing request, mitigate reversibly before root-causing. Use when production or a shared environment is broken now — outages, error-rate or latency spikes, "site is down", pager alerts, users reporting errors — and before restarting, redeploying, flushing, or deleting anything as a first move. Not for reproducible bugs in a dev checkout (see root-cause-debugging).