Build evaluation suites for LLM apps: test sets, graders, regressions, and metrics. Use when measuring prompt or model changes.