Print fairtask's headline comparison (one-prompt baseline vs the final agent pipeline) and the all-systems table over the same 30 human-annotated development-set cases, generated by the repository's own report script from the committed results. Use when asked for the results, the headline numbers, the comparison table, or baseline versus final. No model calls; the first run may shallow-clone the 30 evaluation repositories (about 1.2 GB) to re-verify cited evidence, and says so before doing it.