AI & Agent WorkflowsOpen accessPublished 3 Oct 2026
Establish whether an agent actually works, and whether it still works, using a recorded eval set rather than ad hoc chats. Use when an agent is about to ship and the only testing was somebody typing a few questions into the test pane, when an agent that used to answer correctly now does not and nobody can say when it broke, when a knowledge source or model version changed and you need to know what it affected, when a stakeholder asks how accurate it is and there is no number to give them, when you cannot decide what "correct" even means for a generative answer, or when the agent works for the builder and fails for a user with different permissions. For a one-off read of a…