Measure and tune a decision-machine-1 integration: build a golden set of real labelled examples, score it in batches, sweep thresholds for precision and recall, pick one cut-off per action, and monitor drift with the response headers. Use before changing any label wording or threshold, and whenever someone wants to publish an accuracy or margin number.