Every request is classified before it runs. Simple work is answered at the cheapest tier; only hard, correctness-critical work escalates.
Concurrency, security, proofs, high-stakes judgment — these jump straight to the top tier, bypassing the small model's blind spots.
The gateway logs every decision: downgrade rate, tier mix, and estimated savings vs. running everything at the top tier.
Point any tool's base URL at the Cheaper gateway — it speaks both the Anthropic and the OpenAI-compatible API, so it routes far beyond Claude. See how each one connects →