SaaS & Software·Jul 21, 2026

Ramp Router

Article URL: Comments URL: Points: 4 # Comments: 0

Hacker News9 min readSingle source
Ramp Router
Image · Hacker News
The gist
5-point summary · 1 min

Article URL: Comments URL: Points: 4 # Comments: 0

  • Ramp RouterThe right model forevery request.We built Router to keep 100+ AI use cases at Ramp on the right model.
  • It cut our LLM costs by 30% while making our features smarter and faster.
  • We regularly add support for new models as they become available.How much does it cost?Ramp Router is free to use at launch.
  • As a thank you to our early access users, the first 500 people invited will receive $100 in promotional credits to get started.
  • Ramp Router gives you one endpoint across eligible models and can route requests to the lowest-cost option that meets the quality bar.
$106K$9.51$5K$4K$3K$2K
In this article

Ramp RouterThe right model forevery request.We built Router to keep 100+ AI use cases at Ramp on the right model. It cut our LLM costs by 30% while making our features smarter and faster. Now we’re opening it up to everyone.01 / PROMPT“Extract line items from these invoices”Structured extraction / high volume02 / PROMPT“Classify this support ticket”Fast / lowest cost03 / PROMPT“Summarize this board deck”Long context04 / PROMPT“Review this contract for renewal risk”High accuracy05 / PROMPT“Explain this transaction anomaly”Complex reasoning06 / PROMPT“Generate SQL for this question”Technical accuracy07 / PROMPT“Debug this failed API request”Coding08 / PROMPT“Translate this customer document”Multilingual09 / PROMPT“Draft a personalized sales email”Tone and creativity10 / PROMPT“Moderate this user message”Low latency01 / PROMPT“Extract line items from these invoices”Structured extraction / high volume02 / PROMPT“Classify this support ticket”Fast / lowest cost03 / PROMPT“Summarize this board deck”Long context04 / PROMPT“Review this contract for renewal risk”High accuracy05 / PROMPT“Explain this transaction anomaly”Complex reasoning06 / PROMPT“Generate SQL for this question”Technical accuracy07 / PROMPT“Debug this failed API request”Coding08 / PROMPT“Translate this customer document”Multilingual09 / PROMPT“Draft a personalized sales email”Tone and creativity10 / PROMPT“Moderate this user message”Low latency01 / PROMPT“Extract line items from these invoices”Structured extraction / high volume02 / PROMPT“Classify this support ticket”Fast / lowest cost03 / PROMPT“Summarize this board deck”Long context04 / PROMPT“Review this contract for renewal risk”High accuracy05 / PROMPT“Explain this transaction anomaly”Complex reasoning06 / PROMPT“Generate SQL for this question”Technical accuracy07 / PROMPT“Debug this failed API request”Coding08 / PROMPT“Translate this customer document”Multilingual09 / PROMPT“Draft a personalized sales email”Tone and creativity10 / PROMPT“Moderate this user message”Low latencyThinkingGemini 3 FlashM24Preview experimentation with fast agentic and multimodal workflows.Selected routeClaude Haiku 4.5M20Lightweight Anthropic workloads where speed and vendor consistency matter.EvaluatedClaude Sonnet 5M09Everyday agentic coding with a balance of capability and cost.EvaluatedGemini 3.5 FlashM22Fast agentic, coding and long-context workflows at production scale.EvaluatedClaude Opus 4.8M07Focused complex fixes that need frontier quality with faster execution.EvaluatedGrok 4.5M01Strong coding quality at reasonable cost when latency is less important.EvaluatedClaude Opus 4.6M08A proven general-purpose option for difficult coding work.EvaluatedClaude Fable 5M05The hardest, highest-value tasks where success matters more than cost or speed.EvaluatedGemini 3.1 ProM15Complex, long-context or multimodal tasks within the Google ecosystem.EvaluatedGemini 3.1 Flash LiteM23Simple, high-volume tasks optimized for speed and minimal cost.EvaluatedGemini 3 FlashM24Preview experimentation with fast agentic and multimodal workflows.Selected routeClaude Haiku 4.5M20Lightweight Anthropic workloads where speed and vendor consistency matter.EvaluatedClaude Sonnet 5M09Everyday agentic coding with a balance of capability and cost.EvaluatedGemini 3.5 FlashM22Fast agentic, coding and long-context workflows at production scale.EvaluatedClaude Opus 4.8M07Focused complex fixes that need frontier quality with faster execution.EvaluatedGrok 4.5M01Strong coding quality at reasonable cost when latency is less important.EvaluatedClaude Opus 4.6M08A proven general-purpose option for difficult coding work.EvaluatedClaude Fable 5M05The hardest, highest-value tasks where success matters more than cost or speed.EvaluatedGemini 3.1 ProM15Complex, long-context or multimodal tasks within the Google ecosystem.EvaluatedGemini 3.1 Flash LiteM23Simple, high-volume tasks optimized for speed and minimal cost.EvaluatedGemini 3 FlashM24Preview experimentation with fast agentic and multimodal workflows.Selected routeClaude Haiku 4.5M20Lightweight Anthropic workloads where speed and vendor consistency matter.EvaluatedClaude Sonnet 5M09Everyday agentic coding with a balance of capability and cost.EvaluatedGemini 3.5 FlashM22Fast agentic, coding and long-context workflows at production scale.EvaluatedClaude Opus 4.8M07Focused complex fixes that need frontier quality with faster execution.EvaluatedGrok 4.5M01Strong coding quality at reasonable cost when latency is less important.EvaluatedClaude Opus 4.6M08A proven general-purpose option for difficult coding work.EvaluatedClaude Fable 5M05The hardest, highest-value tasks where success matters more than cost or speed.EvaluatedGemini 3.1 ProM15Complex, long-context or multimodal tasks within the Google ecosystem.EvaluatedGemini 3.1 Flash LiteM23Simple, high-volume tasks optimized for speed and minimal cost.Evaluated01 / PROMPT“Extract line items from these invoices”Structured extraction / high volume02 / PROMPT“Classify this support ticket”Fast / lowest cost03 / PROMPT“Summarize this board deck”Long context04 / PROMPT“Review this contract for renewal risk”High accuracy05 / PROMPT“Explain this transaction anomaly”Complex reasoning06 / PROMPT“Generate SQL for this question”Technical accuracy07 / PROMPT“Debug this failed API request”Coding08 / PROMPT“Translate this customer document”Multilingual09 / PROMPT“Draft a personalized sales email”Tone and creativity10 / PROMPT“Moderate this user message”Low latency01 / PROMPT“Extract line items from these invoices”Structured extraction / high volume02 / PROMPT“Classify this support ticket”Fast / lowest cost03 / PROMPT“Summarize this board deck”Long context04 / PROMPT“Review this contract for renewal risk”High accuracy05 / PROMPT“Explain this transaction anomaly”Complex reasoning06 / PROMPT“Generate SQL for this question”Technical accuracy07 / PROMPT“Debug this failed API request”Coding08 / PROMPT“Translate this customer document”Multilingual09 / PROMPT“Draft a personalized sales email”Tone and creativity10 / PROMPT“Moderate this user message”Low latency01 / PROMPT“Extract line items from these invoices”Structured extraction / high volume02 / PROMPT“Classify this support ticket”Fast / lowest cost03 / PROMPT“Summarize this board deck”Long context04 / PROMPT“Review this contract for renewal risk”High accuracy05 / PROMPT“Explain this transaction anomaly”Complex reasoning06 / PROMPT“Generate SQL for this question”Technical accuracy07 / PROMPT“Debug this failed API request”Coding08 / PROMPT“Translate this customer document”Multilingual09 / PROMPT“Draft a personalized sales email”Tone and creativity10 / PROMPT“Moderate this user message”Low latencyThinkingM24 / GoogleGemini 3 FlashSelected routeM20 / AnthropicClaude Haiku 4.5Lightweight Anthropic workloads where speed and vendor consistency matter.M09 / AnthropicClaude Sonnet 5Everyday agentic coding with a balance of capability and cost.M22 / GoogleGemini 3.5 FlashFast agentic, coding and long-context workflows at production scale.M07 / AnthropicClaude Opus 4.8Focused complex fixes that need frontier quality with faster execution.M01 / xAIGrok 4.5Strong coding quality at reasonable cost when latency is less important.M08 / AnthropicClaude Opus 4.6A proven general-purpose option for difficult coding work.M05 / AnthropicClaude Fable 5The hardest, highest-value tasks where success matters more than cost or speed.M15 / GoogleGemini 3.1 ProComplex, long-context or multimodal tasks within the Google ecosystem.M23 / GoogleGemini 3.1 Flash LiteSimple, high-volume tasks optimized for speed and minimal cost.M24 / GoogleGemini 3 FlashSelected routeM20 / AnthropicClaude Haiku 4.5Lightweight Anthropic workloads where speed and vendor consistency matter.M09 / AnthropicClaude Sonnet 5Everyday agentic coding with a balance of capability and cost.M22 / GoogleGemini 3.5 FlashFast agentic, coding and long-context workflows at production scale.M07 / AnthropicClaude Opus 4.8Focused complex fixes that need frontier quality with faster execution.M01 / xAIGrok 4.5Strong coding quality at reasonable cost when latency is less important.M08 / AnthropicClaude Opus 4.6A proven general-purpose option for difficult coding work.M05 / AnthropicClaude Fable 5The hardest, highest-value tasks where success matters more than cost or speed.M15 / GoogleGemini 3.1 ProComplex, long-context or multimodal tasks within the Google ecosystem.M23 / GoogleGemini 3.1 Flash LiteSimple, high-volume tasks optimized for speed and minimal cost.M24 / GoogleGemini 3 FlashSelected routeM20 / AnthropicClaude Haiku 4.5Lightweight Anthropic workloads where speed and vendor consistency matter.M09 / AnthropicClaude Sonnet 5Everyday agentic coding with a balance of capability and cost.M22 / GoogleGemini 3.5 FlashFast agentic, coding and long-context workflows at production scale.M07 / AnthropicClaude Opus 4.8Focused complex fixes that need frontier quality with faster execution.M01 / xAIGrok 4.5Strong coding quality at reasonable cost when latency is less important.M08 / AnthropicClaude Opus 4.6A proven general-purpose option for difficult coding work.M05 / AnthropicClaude Fable 5The hardest, highest-value tasks where success matters more than cost or speed.M15 / GoogleGemini 3.1 ProComplex, long-context or multimodal tasks within the Google ecosystem.M23 / GoogleGemini 3.1 Flash LiteSimple, high-volume tasks optimized for speed and minimal cost.Production volume2.75T+Tokens routed monthlyCost reduction~30%at 30 ms added latencyRouting reliability99.999%successful routesRouting is just the start.Router chooses the right model for the job, then applies 100+ optimizations to get it done for less.Pay for what the job needs. Nothing more.Every week, the price-intelligence-latency frontier shifts. Router tests each new model on real work, then automatically sends every request to the lowest-cost model that clears its quality bar.Explore the full benchmarkImplement in a few lines of codeOne endpoint gives you leading closed and open models. Router handles routing, fallbacks, and provider updates so you benefit from new models without rewriting your application.$terminalcurlcurl https://router.ramp.com/v1/responses \ -H "Authorization: Bearer rk_live_8f3a2c91e7b04d6a" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-4o-mini", "input": "Hello from Router" }' “At Ramp, Router cut our LLM costs by 30% while making our features smarter and faster.”Rahul SengottuveluCTO, RampEvery trick, out-of-the-box.Router handles caching, compaction, semantic attribution and 100+ optimizations on every request to make it faster and cheaper.See who spent what, where.Attribute every request by model, product, team, and project with Ramp Token Spend Management.Explore Token Spend ManagementInsightsAI token spendOverviewBriefingAPI keysPeopleTotal spend$106K →32%Price per million tokens$9.51 →17%Peak dayLast 30 daysJune 1 →82%Total spend over time30 days · USD$5K$4K$3K$2K$1K$0$1.4K6/1$3.3K$3.7K6/5$4K$3.7K6/9$3.5K$2.4K6/13$1.4K$3.7K6/17$3.9K$4.1K6/21$4K$3.8K6/25AnthropicOpenAIGeminiCursor7-day rolling averageFrequently asked questionsStraight answers about how Ramp Router works, who it’s for, and how to get started.What is an LLM router?An LLM router is one API endpoint for accessing multiple AI models. Instead of wiring your app to one provider at a time, you send requests through our router which can choose the right model for the job based on quality, cost, and availability.How does it work?Your request goes to Ramp Router first. We’ll authenticate the request and help you track the usage, model, provider, and cost. We’ll automatically route eligible requests to a more cost-efficient tier when it won't affect quality.Why is Ramp building an LLM router?Saving you time and money is the whole reason Ramp exists. AI tokens are the fastest-growing spend category, and we want every token you use to be worth it.What models do you support?Ramp Router supports the latest models from OpenAI, Anthropic, and Gemini plus select open-source models, including Kimi. We regularly add support for new models as they become available.How much does it cost?Ramp Router is free to use at launch. You’ll pay list price for the tokens you use. As a thank you to our early access users, the first 500 people invited will receive $100 in promotional credits to get started. Subject to offer terms.Do I need to be a Ramp customer?No. You don’t need a Ramp card, a company account, or even an LLC. We’re inviting an initial group of users to get early access. Join the waitlist to request your spot.What's the difference between using Ramp Router and going direct with a provider?Going direct usually means choosing one provider, one model catalog, and one pricing structure yourself. Ramp Router gives you one endpoint across eligible models and can route requests to the lowest-cost option that meets the quality bar. As new models launch, change prices, and performance change, Ramp Router will adapt automatically without requiring you to constantly rework your integration.Do I have to rewrite my code to use this?No. Ramp Router has an OpenAI-compatible API, so if you’re already using the OpenAI SDK or another OpenAI-compatible framework (which is basically all of them!), switching should be a one-line change: update your base URL to Ramp Router’s endpoint.Who should use this?Router is for anyone who wants to get more from AI without overpaying for it. It helps you get the same work done for less, whether you’re testing an idea on your own, building for a team, or managing AI across an enterprise. We’re inviting an initial group now. Request an invite.What happens if a provider has an outage or rate-limits me?If a provider goes down or rate-limits you, Ramp Router can route eligible requests to another available model, so your app has a fallback when one provider cannot serve it.How does Ramp Router handle my data and credentials?Please see the Ramp Router Privacy Policy for information on how Ramp manages personal information.Tokens are money. Save both.Model providers make more when you use more frontier intelligence. Ramp wins when you spend less.

Integrity note  ·  Xela does not rewrite or paraphrase article content. The excerpt above is the source publication's own words, sanitized for display. For the full piece — including any quotes, charts, or images — read it at Hacker News. Xela's rewritten version is off for this story, so there's no editorial angle attached — you're getting the source's reporting unfiltered. When the rewrite is on, we add a What this means block underneath with the operator/trader takeaway.

What people are saying

Discussion

Hot takes

0/280

Loading takes…

Comments

Discussion · 0

Sign in to comment, like, and save articles.

Sign in

Loading comments…

Newsletter

Track saas & software every morning.

Daily digest tuned to this beat. The 5 stories most worth your time. Unsubscribe anytime.