FastRouter.ai - Route requests to the right LLM for cost, latency & quality
by•
FastRouter is a unified AI gateway and control plane for developers and enterprise teams building with LLMs. It routes every request to the right model across 200+ LLMs through one OpenAI-compatible API, optimizing for cost, latency, quality, and reliability. With intelligent routing, failover, observability, and governance, teams can scale AI apps without vendor lock-in or code changes.

Replies
we ended up building a crude version of this internally last month wish we found you guys earlier haha.
@adams_parker Haha, totally get it. The first version is one thing; maintaining it as providers and models change is another 😅 It’s worth keeping an eye on the ongoing DevOps work and the latency the router itself adds.
With FastRouter, you get a lightning fast gateway and also useful features such as AI evals, prompt optimization, and real-time alerts without having to build those separately. Would love to hear what you built and see whether we can save your team some work! Reach us at support@fastrouter.ai.
Congrats on the launch. Can I set my own routing policies in settings on top of the default logic?
@iamanantgupta Thanks, Anant! Yes, you can configure your own routing policies using Virtual Model Aliases: https://docs.fastrouter.ai/explore-features/virtual-model-aliases
Options include Random Shuffle, Lowest Latency, Highest Throughput, Lowest Usage, Lowest Price, Priority Routing, and Category-Based Routing. For example, you can prioritize a specific provider and fall back to the next only if a request fails, or choose the lowest-cost option from a selected set of providers.
In the next couple of days, we’re also launching Optimization Settings, through which you’ll be able to assign a setting to individual model choices within an alias - e.g. try the flex tier always for a particular model within an alias; use prompt compression with one of choices; etc. Stay tuned!
Can devs prioritise which criteria they want to optimise for: cost, latency, quality, and reliability at the routing layer?
@nuseir_yassin1 Yes, you can configure routing policies using Virtual Model Aliases. These include Lowest Latency, Highest Throughput, Lowest Price, Lowest Usage, Priority Routing, Category-Based Routing, and Random Shuffle. For example, you can prioritize a preferred provider with fallback to the next ordered provider on an error, or choose the lowest-cost option from a selected set of providers.
In the next couple of days, we’re also launching Optimization Settings, which you can assign specific settings to individual model choices within an alias so you have greater control on how you access a model within a provider - for e.g. use the flex tier from OpenAI for model X.
For Quality, our AI Custom Evaluations feature and Insights help you compare models on your actual requests, so you can decide which models to opt for based on cost or speed or quality or combination thereof. Insights gives you recommendations to act on with evidence. The actual switching is a decision you can take based on reviewing the data.
The first users usually seem to come from places where the problem is already being discusses. Curious if anyone has had better results from communities than from posting on their own profiles.
Many congratulations, @ritprasad , @andrej_gamser2 and team! 😊
@dhanrajchoudhary Thank you! Your support means a lot. Really appreciate it.
Really like the failover functionality. If a provider becomes unavailable or throttles during the request process, how quickly is it done and is the streaming capability and tool calls maintained for all the models?
Btw, Congratulations @ritprasad , @andrej_gamser2 and team FastRouter 🚀✌️
@aymi_malik The retry is immediate as soon as we get an error from the provider. For more context, please see an article we wrote up recently: https://fastrouter.ai/blog/posts/anthropic-went-down-fastrouter-didnt-150-requests-one-live-outage-zero-downtime
Further, on the Activity Log, you get the full retry chain of the request including the failed attempts with other providers and the overall request level response time and latencies as this is what matters to your end-user. As regards, streaming capability and tool calls - both work across the working provider. Happy to discuss any specific use case - please do send us an email at support@fastrouter.ai.
Awesome product, congrats with a launch and good luck!
@annmast Thank you for your message. Really appreciate it and grateful :)
@nalin_rajendran Thank you for your support. Really appreciate it and grateful!
Finally someone solved the headache of writing manual fallback code every time OpenAI drops.
@ashir_murtaza1 Thanks Ashir! That was one of the main pain points. Nobody should have to rewrite retry-and-switch logic - and waste time integrating new provider SDKs. We have moved a long way since then and added a lot more value added features around insights and optimizations. Thanks for checking us out!
This solves a massive headache for us. we were literally spending hours last week figuring out how to balance claude and gpt costs.
@sansa_grey Yup, figuring out when to use Claude vs. OpenAI or even other models, tracking spend, and checking quality can quickly become a job of its own. That’s exactly why we built an easy to use AI evals product; as well as proactive cost recommendations with Insights -- to help you identify when to leverage different models. Thanks for checking us out!