Financial infrastructure for AI inference.
tiers is an LLM routing and governance layer — one endpoint between your application and every model provider. It classifies each call, routes to the cheapest model that clears the quality bar, enforces policy, verifies the output, and returns a signed receipt.
Switch to Ungoverned. The requests stay the same. The quality bar stays the same. Only the economics change.
Some work never needs a model. tiers resolves it locally, with no token bill and no provider exposure.
A router chooses a model. tiers governs the outcome.
Every production AI call creates four questions: where should it run, is it allowed, did it work, and can you prove what happened? Gateways answer the first. Dashboards arrive after the fact. tiers answers all four inside the call path.
Observability tells you what happened after the call. A gateway moves traffic and walks away. tiers prices, inspects, verifies, and settles the call before the output reaches your system. That is the difference between moving traffic and governing production.
Four capabilities. One governed call.
Each one solves a different failure mode. Together they turn inference into infrastructure.
Gate
Finds the lowest-cost execution path that meets the task's quality, latency, and policy requirements.
Aiglos
Enforces policy inside the session, before tools, network requests, and subprocesses execute.
Forge
Checks the result against the task, schema, and rolling quality baseline before release.
Ulmo
Settles the call into one tamper-evident record: route, cost, policy, verification, and outcome.
Every call makes the system better.
Routing without verification is guesswork. Verification without settlement is a dashboard. tiers connects the signals, so every governed outcome improves the next decision.
The first call is governed by policy. The thousandth is governed by everything the system learned before it. On the reference workload, governed spend fell from $47.00 to $6.10 per day while maintaining the quality bar.1
The same work. 87% less spend.
tiers did not reach this result by lowering the quality bar. It reached it by paying the right price for each unit of work: smaller models where they were sufficient, deterministic execution where a model was unnecessary, and verification before an output entered production.1
The point is not a smaller AI budget. It is more useful work from the same one. Every dollar removed from waste becomes another task, another workflow, or another team that can use intelligence.
One endpoint. A live market behind it.
Models differ by capability, latency, geography, price, and risk. List prices across production models span a 643x range.2 tiers clears each task against those constraints, then records how the decision was made.
tiers has no model to favor and no token markup to protect. We operate no foundation models, sell no tokens, retain no prompt or output content, and are paid to govern the call, not steer it to a provider.
Every agent action contains two transactions: a cognition transaction that decides what to do, and a value transaction that does it. The second has financial infrastructure. tiers is building it for the first.
The market has reached the same conclusion.
The largest AI buyers are asking for lower cost, clearer control, and an independent exchange between applications and models.3
“There’s going to be a hot new company that’s going to come along and say this: I’m going to sit between you and Anthropic and OpenAI, and I’m going to make sure that you only need their tokens when you actually need them.”
Marc Benioff, Chair and CEO, Salesforce · All-In, May 2026
Change one line. Govern every call.
Point your existing SDK at tiers. Keep your provider relationships, your keys, and your application code.
Simulated session using verified reference-workload figures.1 When the proxy is live, the same component runs against your tiers endpoint and nothing changes except the transport. Open the full demo.
Aligned by design.
tiers charges 2% of governed spend. No markup on model costs. No seats. No minimums. When your inference bill falls, our fee falls with it.
Built for the people accountable for AI.
Engineering gets control. Finance gets attribution. Security gets enforcement. Leadership gets proof.
Put every AI call under control.
One endpoint across every model. Lower cost, enforced policy, verified outputs, and a signed receipt for every decision. And the same budget operates 7.7x the intelligence.1
Common questions.
What is AI inference governance?
How do I reduce AI API costs?
What is the difference between an AI gateway and tiers?
Why use tiers instead of calling model APIs directly?
How does tiers route AI calls to cheaper models?
What is a signed settlement receipt?
How many providers does tiers support?
base_url to the tiers endpoint and pass your tiers key. Your provider keys and contracts stay yours.How much does tiers cost?
How do I integrate tiers into an existing application?
base_url of your OpenAI or Anthropic SDK client to the tiers endpoint and replace your provider API key with your tiers key. Everything else — your prompts, application logic, and provider relationships — stays the same. For OpenAI: base_url="https://YOUR_TIERS_ENDPOINT/v1". For Anthropic: base_url="https://YOUR_TIERS_ENDPOINT". For direct HTTP, POST to https://YOUR_TIERS_ENDPOINT/v1/chat/completions with Authorization: Bearer your_tiers_key.1. Verified against a signed, immutable 14-day shadow baseline running the identical workload through the enterprise default model and through tiers: $47.00 per day default versus $6.10 per day governed on the reference workload, an 87% reduction at the maintained quality bar. Platform average: 83.08% across measured cohorts; optimized cohorts: 87.9%. The 90-day chart shows the production learning curve; the final 14-day comparison is the signed counterfactual. Category shares and console rows illustrate the reference workload. Results vary by workload.
2. Provider list-price snapshot, , across the 56 models and 14 providers in the tiers catalog. Prices, availability, and model coverage change over time.
3. All quotations are public statements about the market, reproduced verbatim with source and date. Ellipses mark omitted words. None are endorsements of tiers.