Run one analytics question several **independent** times and report whether the answer is stable (every run agrees) or drifting (runs disagree because the question is under-defined). Stability is necessary, not sufficient: a wrong query is perfectly stable. This check needs no ground truth.