Eragon AI projects hundreds of thousands in inference savings with Merge Gateway

Eragon AI projects hundreds of thousands in inference savings with Merge Gateway

Eragon AI fact sheet

Product
AI operating system for work
HQ
San Francisco, CA
Industry

Artificial Intelligence

Merge category
No items found.
Merge Common Models
No items found.
Get a demo

Merge Gateway lets us route queries to the most appropriate model for the task, and this year that's gonna save us hundreds of thousands of dollars in inference cost.

Lance Mathias
Research Engineer
Problem

Model provisioning quickly became a full-time engineering function

Eragon AI is an AI operating system for work. Its agents connect to the systems a company already uses and automate tasks across email, CRM systems, document stores, software development, and operations. To serve that range of work, Eragon AI needs to choose the right model for each request while respecting customer preferences, data constraints, cost, and availability. By combining its own routing logic with Merge Gateway, the company now projects hundreds of thousands of dollars in inference savings this year and has reclaimed about one and a half engineers' full-time workload.

In Eragon AI's early days, most requests went to the most capable model available. That gave the team a straightforward starting point, but many tasks did not need frontier-level intelligence. Checking an email or retrieving information from documents could often be handled by a faster, less expensive model with the same quality.

Eragon AI began building routing logic that could classify each prompt and select a model based on the task. Customers could set their own preferences, such as using one model for coding and another for writing emails. Data residency added another requirement. A workflow involving sensitive European data might need to stay inside Eragon AI's environment or use a model hosted by the customer.

The model catalog behind that routing quickly became difficult to manage.

 "If we wanna use models across, say, 20 different companies, that's 20 accounts, 20 different invoices, 20 different API keys that we have to keep track of," Lance said. 

Some providers lacked admin APIs, so engineers had to log into a website and create keys by hand. Setting up 100 accounts could mean repeating that flow 100 times.

The burden landed across the engineering team. Lance, a Research Engineer at Eragon AI, described account provisioning, key setup, and handoff as "totally just a full-time job for the engineering team." At the time, the company had three engineers. About half of that team's capacity was going toward operational setup instead of the product.

Provider limits created another source of risk. The team hit Anthropic rate limits and ran out of credits daily during Eragon AI's early growth. A frozen account could interrupt access to the model selected for a task, and every new provider introduced another account that engineers had to watch.

Solution

One Gateway endpoint automated provisioning and expanded model choice

‍"We managed to do the Gateway integration in literally just a matter of minutes,"

The team added one Merge API key, set the Gateway endpoint, and kept its routing logic in place.

Eragon AI continues to decide which model should handle a request. Its own systems classify the prompt, apply the customer's task preferences, and enforce any deployment or data constraints. Once a model is selected, Merge Gateway provides a single endpoint for reaching it.

Through that connection, Eragon AI can access proprietary and open-weight models as they become available. It also enabled Eragon and their customers to bring their own provider keys and use spend committed to those accounts. Managed fallbacks give requests another path when an account reaches a limit or a smaller provider becomes unavailable.

Merge Gateway's Admin API removed the manual work around customer provisioning, allowing Eragon AI to create, disable, and rotate keys programmatically. The team paired the Admin API with an Eragon AI operations bot, automating the setup flow that engineers previously handled one account at a time.

The automation freed up about one and a half engineers' full-time workload almost immediately, according to Lance.

The usage dashboard gives the team one place to follow cost, adoption, traffic, and model selection. That visibility helps Eragon AI confirm that each type of work is reaching the intended model as customer usage and the catalog grow.

‍

Outcome

Gateway Routing cut evaluated inference cost by 70 to 90% while maintaining performance levels

Eragon AI's evaluations show the economic effect of matching model capability to the task. Instead of sending every request to the highest-cost model, the company can use a lower-cost model when its benchmarks show no loss in performance. Those evaluations show a 70 to 90% cost reduction at the same performance.

"Merge Gateway lets us route queries to the most appropriate model for the task, and this year that's gonna save us hundreds of thousands of dollars in inference cost," Lance said.

At Eragon AI's current scale, that difference compounds quickly. In one week, the company processed about 12 billion tokens across 200,000 requests. Lance projects that task-specific routing through Merge Gateway will save hundreds of thousands of dollars in inference costs this year.

The operating model has changed along with the economics. Provisioning and key-management time in production has fallen essentially to zero. At least 50 different models now run in production behind one Gateway connection, compared to  roughly five models across two providers in Eragon AI's early days. The Admin API keeps new-account setup from consuming engineering capacity as that usage grows.

Eragon AI's engineers can now spend that capacity on the product. As Lance put it,

"Now that we have Merge Gateway, our engineers at Eragon have been freed up to devote their time to making Eragon agents smarter and more knowledgeable."

‍

Customer stories

How Assemble uses Merge to add HRIS and ATS integrations in a matter of minutes
HRIS
HRIS
ATS
ATS
How Causal sped up their self-serve time-to-value by 20% with Merge
Accounting
Accounting
HRIS
HRIS
CRM
CRM
How Ramp uses Merge’s HRIS integrations to improve the user experience for thousands of customers
HRIS
HRIS
Ticketing
Ticketing

Make integrations your competitive advantage

Stay in touch to learn how Merge can unlock hundreds of integrations in days, not years

But Merge isn’t just a Unified 
API product. Merge is an integration platform to also manage customer integrations.  gradient text
But Merge isn’t just a Unified 
API product. Merge is an integration platform to also manage customer integrations.  gradient text
But Merge isn’t just a Unified 
API product. Merge is an integration platform to also manage customer integrations.  gradient text
But Merge isn’t just a Unified 
API product. Merge is an integration platform to also manage customer integrations.  gradient text