Route each LLM call to the cheapest model that can handle it, with confidence-based fallback to larger models. Use when one frontier model serves all traffic, when cost-per-call is uniform across vastly different task difficulties, or when designing a new GenAI feature for cost-efficient scale.