Use when building an LLM agent that makes repeated decisions with measurable outcomes (trading, deploys, content picks, lead scoring, moderation) and should improve from its own results — especially when the agent repeats past mistakes, ignores what worked before, or its prompt grows unbounded with history.