AI News HubLIVE
In-site rewrite5 min read

Ramp AI Router: Choose the Right Model for Every Request, Cut Costs by 30%

Ramp released its AI Router, which automatically selects the best language model for each request, reducing LLM costs by 30% while maintaining performance. Previously used internally for over 100 AI use cases, the tool is now available to the public.

SourceHacker News AIAuthor: robbiet480

Ramp Router

The right model forevery request.

We built Router to keep 100+ AI use cases at Ramp on the right model. It cut our LLM costs by 30% while making our features smarter and faster. Now we’re opening it up to everyone.

Request access

01 / PROMPT

“Extract line items from these invoices”

Structured extraction / high volume

02 / PROMPT

“Classify this support ticket”

Fast / lowest cost

03 / PROMPT

“Summarize this board deck”

Long context

04 / PROMPT

“Review this contract for renewal risk”

High accuracy

05 / PROMPT

“Explain this transaction anomaly”

Complex reasoning

06 / PROMPT

“Generate SQL for this question”

Technical accuracy

07 / PROMPT

“Debug this failed API request”

Coding

08 / PROMPT

“Translate this customer document”

Multilingual

09 / PROMPT

“Draft a personalized sales email”

Tone and creativity

10 / PROMPT

“Moderate this user message”

Low latency

01 / PROMPT

“Extract line items from these invoices”

Structured extraction / high volume

02 / PROMPT

“Classify this support ticket”

Fast / lowest cost

03 / PROMPT

“Summarize this board deck”

Long context

04 / PROMPT

“Review this contract for renewal risk”

High accuracy

05 / PROMPT

“Explain this transaction anomaly”

Complex reasoning

06 / PROMPT

“Generate SQL for this question”

Technical accuracy

07 / PROMPT

“Debug this failed API request”

Coding

08 / PROMPT

“Translate this customer document”

Multilingual

09 / PROMPT

“Draft a personalized sales email”

Tone and creativity

10 / PROMPT

“Moderate this user message”

Low latency

01 / PROMPT

“Extract line items from these invoices”

Structured extraction / high volume

02 / PROMPT

“Classify this support ticket”

Fast / lowest cost

03 / PROMPT

“Summarize this board deck”

Long context

04 / PROMPT

“Review this contract for renewal risk”

High accuracy

05 / PROMPT

“Explain this transaction anomaly”

Complex reasoning

06 / PROMPT

“Generate SQL for this question”

Technical accuracy

07 / PROMPT

“Debug this failed API request”

Coding

08 / PROMPT

“Translate this customer document”

Multilingual

09 / PROMPT

“Draft a personalized sales email”

Tone and creativity

10 / PROMPT

“Moderate this user message”

Low latency

Thinking

Gemini 3 Flash

M24

Preview experimentation with fast agentic and multimodal workflows.

Selected route

Claude Haiku 4.5

M20

Lightweight Anthropic workloads where speed and vendor consistency matter.

Evaluated

Claude Sonnet 5

M09

Everyday agentic coding with a balance of capability and cost.

Evaluated

Gemini 3.5 Flash

M22

Fast agentic, coding and long-context workflows at production scale.

Evaluated

Claude Opus 4.8

M07

Focused complex fixes that need frontier quality with faster execution.

Evaluated

Grok 4.5

M01

Strong coding quality at reasonable cost when latency is less important.

Evaluated

Claude Opus 4.6

M08

A proven general-purpose option for difficult coding work.

Evaluated

Claude Fable 5

M05

The hardest, highest-value tasks where success matters more than cost or speed.

Evaluated

Gemini 3.1 Pro

M15

Complex, long-context or multimodal tasks within the Google ecosystem.

Evaluated

Gemini 3.1 Flash Lite

M23

Simple, high-volume tasks optimized for speed and minimal cost.

Evaluated

Gemini 3 Flash

M24

Preview experimentation with fast agentic and multimodal workflows.

Selected route

Claude Haiku 4.5

M20

Lightweight Anthropic workloads where speed and vendor consistency matter.

Evaluated

Claude Sonnet 5

M09

Everyday agentic coding with a balance of capability and cost.

Evaluated

Gemini 3.5 Flash

M22

Fast agentic, coding and long-context workflows at production scale.

Evaluated

Claude Opus 4.8

M07

Focused complex fixes that need frontier quality with faster execution.

Evaluated

Grok 4.5

M01

Strong coding quality at reasonable cost when latency is less important.

Evaluated

Claude Opus 4.6

M08

A proven general-purpose option for difficult coding work.

Evaluated

Claude Fable 5

M05

The hardest, highest-value tasks where success matters more than cost or speed.

Evaluated

Gemini 3.1 Pro

M15

Complex, long-context or multimodal tasks within the Google ecosystem.

Evaluated

Gemini 3.1 Flash Lite

M23

Simple, high-volume tasks optimized for speed and minimal cost.

Evaluated

Gemini 3 Flash

M24

Preview experimentation with fast agentic and multimodal workflows.

Selected route

Claude Haiku 4.5

M20

Lightweight Anthropic workloads where speed and vendor consistency matter.

Evaluated

Claude Sonnet 5

M09

Everyday agentic coding with a balance of capability and cost.

Evaluated

Gemini 3.5 Flash

M22

Fast agentic, coding and long-context workflows at production scale.

Evaluated

Claude Opus 4.8

M07

Focused complex fixes that need frontier quality with faster execution.

Evaluated

Grok 4.5

M01

Strong coding quality at reasonable cost when latency is less important.

Evaluated

Claude Opus 4.6

M08

A proven general-purpose option for difficult coding work.

Evaluated

Claude Fable 5

M05

The hardest, highest-value tasks where success matters more than cost or speed.

Evaluated

Gemini 3.1 Pro

M15

Complex, long-context or multimodal tasks within the Google ecosystem.

Evaluated

Gemini 3.1 Flash Lite

M23

Simple, high-volume tasks optimized for speed and minimal cost.

Evaluated

01 / PROMPT

“Extract line items from these invoices”

Structured extraction / high volume

02 / PROMPT

“Classify this support ticket”

Fast / lowest cost

03 / PROMPT

“Summarize this board deck”

Long context

04 / PROMPT

“Review this contract for renewal risk”

High accuracy

05 / PROMPT

“Explain this transaction anomaly”

Complex reasoning

06 / PROMPT

“Generate SQL for this question”

Technical accuracy

07 / PROMPT

“Debug this failed API request”

Coding

08 / PROMPT

“Translate this customer document”

Multilingual

09 / PROMPT

“Draft a personalized sales email”

Tone and creativity

10 / PROMPT

“Moderate this user message”

Low latency

01 / PROMPT

“Extract line items from these invoices”

Structured extraction / high volume

02 / PROMPT

“Classify this support ticket”

Fast / lowest cost

03 / PROMPT

“Summarize this board deck”

Long context

04 / PROMPT

“Review this contract for renewal risk”

High accuracy

05 / PROMPT

“Explain this transaction anomaly”

Complex reasoning

06 / PROMPT

“Generate SQL for this question”

Technical accuracy

07 / PROMPT

“Debug this failed API request”

Coding

08 / PROMPT

“Translate this customer document”

Multilingual

09 / PROMPT

“Draft a personalized sales email”

Tone and creativity

10 / PROMPT

“Moderate this user message”

Low latency

01 / PROMPT

“Extract line items from these invoices”

Structured extraction / high volume

02 / PROMPT

“Classify this support ticket”

Fast / lowest cost

03 / PROMPT

“Summarize this board deck”

Long context

04 / PROMPT

“Review this contract for renewal risk”

High accuracy

05 / PROMPT

“Explain this transaction anomaly”

Complex reasoning

06 / PROMPT

“Generate SQL for this question”

Technical accuracy

07 / PROMPT

“Debug this failed API request”

Coding

08 / PROMPT

“Translate this customer document”

Multilingual

09 / PROMPT

“Draft a personalized sales email”

Tone and creativity

10 / PROMPT

“Moderate this user message”

Low latency

Thinking

M24 / Google

Gemini 3 Flash

Selected route

M20 / Anthropic

Claude Haiku 4.5

Lightweight Anthropic workloads where speed and vendor consistency matter.

M09 / Anthropic

Claude Sonnet 5

Everyday agentic coding with a balance of capability and cost.

M22 / Google

Gemini 3.5 Flash

Fast agentic, coding and long-context workflows at production scale.

M07 / Anthropic

Claude Opus 4.8

Focused complex fixes that need frontier quality with faster execution.

M01 / xAI

Grok 4.5

Strong coding quality at reasonable cost when latency is less important.

M08 / Anthropic

Claude Opus 4.6

A proven general-purpose option for difficult coding work.

M05 / Anthropic

Claude Fable 5

The hardest, highest-value tasks where success matters more than cost or speed.

M15 / Google

Gemini 3.1 Pro

Complex, long-context or multimodal tasks within the Google ecosystem.

M23 / Google

Gemini 3.1 Flash Lite

Simple, high-volume tasks optimized for speed and minimal cost.

M24 / Google

Gemini 3 Flash

Selected route

M20 / Anthropic

Claude Haiku 4.5

Lightweight Anthropic workloads where speed and vendor consistency matter.

M09 / Anthropic

Claude Sonnet 5

Everyday agentic coding with a balance of capability and cost.

M22 / Google

Gemini 3.5 Flash

Fast agentic, coding and long-context workflows at production scale.

M07 / Anthropic

Claude Opus 4.8

Focused complex fixes that need frontier quality with faster execution.

M01 / xAI

Grok 4.5

Strong coding quality at reasonable cost when latency is less important.

M08 / Anthropic

Claude Opus 4.6

A proven general-purpose option for difficult coding work.

M05 / Anthropic

Claude Fable 5

The hardest, highest-value tasks where success matters more than cost or speed.

M15 / Google

Gemini 3.1 Pro

Complex, long-context or multimodal tasks within the Google ecosystem.

M23 / Google

Gemini 3.1 Flash Lite

Simple, high-volume tasks optimized for speed and minimal cost.

M24 / Google

Gemini 3 Flash

Selected route

M20 / Anthropic

Claude Haiku 4.5

Lightweight Anthropic workloads where speed and vendor consistency matter.

M09 / Anthropic

Claude Sonnet 5

Everyday agentic coding with a balance of capability and cost.

M22 / Google

Gemini 3.5 Flash

Fast agentic, coding and long-context workflows at production scale.

M07 / Anthropic

Claude Opus 4.8

Focused complex fixes that need frontier quality with faster execution.

M01 / xAI

Grok 4.5

Strong coding quality at reasonable cost when latency is less important.

M08 / Anthropic

Claude Opus 4.6

A proven general-purpose option for difficult coding work.

M05 / Anthropic

Claude Fable 5

The hardest, highest-value tasks where success matters more than cost or speed.

M15 / Google

Gemini 3.1 Pro

Complex, long-context or multimodal tasks within the Google ecosystem.

M23 / Google

Gemini 3.1 Flash Lite

Simple, high-volume tasks optimized for speed and minimal cost.

Production volume

2.75T+

Tokens routed monthly

Cost reduction

~30%

at 30 ms added latency

Routing reliability

99.999%

successful routes

Routing is just the start.

Router chooses the right model for the job, then applies 100+ optimizations to get it done for less.

Pay for what the job needs. Nothing more.

Every week, the price-intelligence-latency frontier shifts. Router tests each new model on real work, then automatically sends every request to the lowest-cost model that clears its quality bar.

Explore the full benchmark→

Implement in a few lines of code

One endpoint gives you leading closed and open models. Router handles routing, fallbacks, and provider updates so you benefit from new models without rewriting your application.

$terminal

curl

curl https://router.ramp.com/v1/responses \ -H "Authorization: Bearer rk_live_8f3a2c91e7b04d6a" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini", "input": "Hello from Router" }'

“At Ramp, Router cut our LLM costs by 30% while making our features smarter and faster.”

Rahul Sengottuvelu

CTO, Ramp

Every trick, out-of-the-box.

Router handles caching, compaction, semantic attribution and 100+ optimizations on every request to make it faster and cheaper.

Smart Routing

Flex vs Standard

Compression

Caching

Spend controls

Flex

Timing

Smart Routing

Flex vs Standard

Compression

Caching

Spend controls

Flex

Timing

Smart Routing

Flex vs Standard

Compression

Caching

Spend controls

Flex

Timing

See who spent what, where.

Attribute every request by model, product, team, and project with Ramp Token Spend Management.

Explore Token Spend Management→

Total spend

$106K →32%

Price per million tokens

$9

[truncated for AI cost control]