Evaluates a Databricks Genie space on SQL semantic correctness, table appropriateness, time-window, aggregation, follow-up, and refusal by bootstrapping an eval set from the certified-questions panel and registering Genie-aware `mlflow.genai.scorer` callables. Use when the agent emits a `genie_query` span and certified questions has ≥10 rows. Do NOT use when the question is generic SQL evaluation outside the Genie semantic model — defer to upstream Databricks SQL skills.