Use when designing or auditing the empirical core of a COLM paper — contamination analysis for evaluation data, fair baselines under matched prompting and compute, pinned model versions and decoding parameters, uncertainty over runs and samples, scaling coverage, and honest reporting of API-model comparisons.