Every prompt paysthe right price.

Steeros routes every prompt to the cheapest model that can handle it. Up to 85% lower spend, same quality. One proxy, no code changes.

Claude CodeCursorGitHub CopilotWindsurfOpenAIAnthropicGeminiLlamaDeepSeekMistralPerplexityOllamaHugging FaceNVIDIAAWSAzureDatabricksIBMAlibabaReplitCloudflareByteDanceBaiduLM StudioSTEEROS
87.7%
cost cut on our held-out eval
139 prompts, router vs all-flagship
95%
flagship quality kept
RouteLLM study, LMSYS 2024
92.4%
lower than Opus-only routing
our traffic, 1,011 requests
$0.004
average light-tier request
vs $0.068 unsteered

Pick a flagship and the door flies open for everything. Teams burn 60-70% of their spend on prompts a small model handles perfectly.

Like taking a Ferrari to buy milk. Every single time.

Four moves between you and the right model.

1Classify

Score every request before a single token is spent.

A hybrid classifier reads intent, context, tools, and risk. Rules first, embeddings second, an LLM only when the case is genuinely hard.

docscomplexity 2/5risk lowdev-tools
2Route

The cheapest model that clears the bar wins.

Four modes, one policy: cheap, fast, balanced, quality. You set per-tier cost ceilings; the router spends inside them, every time.

light$0.003
balancedchosen$0.012
flagship$0.20
3Convert

Anthropic format in, native format out.

Token caps get clamped, thinking blocks stripped, mid-chat system roles demoted. Each transform is logged so nothing surprises you.

max_tokens clamped to 8192
thinking block stripped
system message demoted to user
4Account

Every cent accounted, to six decimals.

Provider spend, cache reads and writes, classifier calls, compaction. Measured against a fair baseline, session and all-time.

req #48211 deepseek-chat $0.000041
cache read 14.2k tok ×0.1 $0.00002
baseline: same tokens on a fixed model

Three rungs. One ladder.

The smallest model that clears the bar wins. You set the ceilings; the router spends inside them. One proxy in front of DeepSeek, Anthropic, OpenAI, and Gemini.

The errand tier.

Light tier
cost ceiling per request
$0.003
  • git messages
  • README edits
  • typo fixes
  • test stubs

51% of a typical mix lands here

The daily driver.

Balanced tier
cost ceiling per request
$0.012
  • features
  • refactors
  • code review
  • most debugging

32% of a typical mix lands here

When it actually earns it.

Flagship tier
cost ceiling per request
$0.20
  • architecture
  • hard bugs
  • novel design
  • long-horizon planning

17% of a typical mix lands here

The boring details, handled.

Cache-aware pricing

Every request is priced the way the provider actually bills it: cache reads, cache writes, full inputs. Our own traffic banked $22.79 in cache savings.

anthropic read
×0.1
anthropic write
×2.0
deepseek cache hit
~3% of input
deepseek full input
×1.0

Fair baselines

Savings are measured against the same token bundle on a fixed model, at that model’s own cache rates. No flattering math.

1,011 requests, actual: $22.41deepseek-only: $24.31 · opus-only: $294.68

Benchmarks

Held-out eval across chat, math, mcq, and code, scored blind against a judge and ground truth. Small samples on the non-chat categories.

chat 100 · math 80 · mcq 100 · code 5087.7% cheaper than all-flagship

Carbon

Spend and carbon fall together. Fewer flagship tokens means fewer GPU-hours.

66%

fewer flagship tokens, same output

Compaction

Long sessions get compacted; past answers get reused, so repeat prompts skip straight to the right tier. 399 calls on our traffic, $0.12 total.

summary: 3.4k tokens → 412route: remembered → light

Guardrails

Monthly caps that degrade to the light tier. Automatic escalation when a cheap model fails. PII scrubbing on the roadmap.

over_cap_policy: degrade_to_light

Live dashboard

Watch every route, token, and cent in real time, on your own machine at localhost:3000.

tokens, cost, routes

Do the math on your burn.

Slide your team size and see spend before Steeros, spend after, and what stays in your budget.

10

1 to 100 seats

60

10 to 300 prompts

without Steeros
$857/mo
12,600 requests
all at flagship prices
with Steeros
$299/mo
−65%
Kept in your budget$558every month

Same output, roughly 65% fewer flagship tokens, so energy falls with the bill.

Modeled on our pilot mix: 51% light at $0.004, 32% balanced at $0.020, 17% flagship at $0.090 per request. Your distribution will differ.

Free to pilot. Built for fleets.

Steeros Free25kreq/mo
  • Working config template, routes.free.yaml
  • One-line localhost proxy
  • Cost ceilings you edit by hand
  • Enough for a solo dev or a small team pilot
EnterpriseCustom pricing
  • Every agent: Claude Code, Cursor, Copilot, Gemini CLI
  • PII scrubbing and redaction
  • Self-tuning classifier on your prompt mix
  • Unlimited seats, real support

We only use this to reply. No drip campaigns.