Use when deciding between a strong hosted model (GPT-4o) and a small self-hosted VLM/LLM for a structured-extraction task, and preparing the pipeline to swap to the self-hosted one.