So you want to use OpenRouter?
Posted by player85 3 days ago
Comments
Comment by joshstrange 1 day ago
OpenRouter sells the idea of swapping being commodity providers but it couldn’t be further from the truth. Provider A is often not swappable for B or C (again, as this author found). It can be crazy-making as you sit there thinking “OpenRouter has no clothes right?! Am I the one that’s wrong?”.
I love the _idea_ of OpenRouter and maybe Stripe can improve this situation but the only sane way I’ve found to use it is to tightly pin providers to the point I wonder if I should just use the providers directly.
Without pinning you are in for a world of hurt and unreliability (varying model capabilities, speed, etc).
Comment by embedding-shape 1 day ago
But that's the intention right? Even the name implies they just send stuff around for you, and if you want to control the routing, you'd lock down providers. I don't see how they could build what they wanted to build, and not have it end up unreliable if you freely round-robin between providers, it's bound to work exactly like this.
> I love the _idea_ of OpenRouter and maybe Stripe can improve this situation but the only sane way I’ve found to use it is to tightly pin providers to the point I wonder if I should just use the providers directly.
This is quite literally the point of OpenRouter. A unified interface, so you can easily switch providers without changing a ton of code which using providers directly would most likely mean, as there are slight differences between them. And the providers all run different weights, so of course quality/performance will differ among them.
I guess OpenRouter is a bit like Amazon, in that they're just routing stuff around for you, but to actually find the good and usable stuff, you need to focus in on what providers/manufacturers you know are good, and stick with those. Still, the unified interface helps you to shop around and try different ones when you want to.
Comment by Spacemolte 1 day ago
2. As long as the apis use the openai spec, it should be fine? And what makes you say the providers are using different weights? The blogpost says the exact opposite?
3. Great example, if amazon does not lead me to good products, I will stop using it, and this is why I dislike amazon. There are so many crap products, and the search seems to try to push crap products instead of what i'm actually looking for at a good price.
You take a bunch of providers with not great uptime, put them in a pool and now you get great uptime - but it doesn't work if it's at the cost of shitty performance or failing toolcalls.
Comment by eli 18 hours ago
Comment by embedding-shape 1 day ago
OpenRouter does reliably route to your specified model and provider, otherwise it'd pretty much be fully broken. Parent is complaining about the auto-provider chosing, not that all providers are unreliable.
> 2. As long as the apis use the openai spec, it should be fine? And what makes you say the providers are using different weights? The blogpost says the exact opposite?
In theory, yes. In practice, no, there are differences. Ollama, llama.cpp, vLLM and SGLang all say "ChatCompletionRequest" compatible, but the devil is in the details, they don't have 100% the same request/response schema across all compatible models.
> 3. Great example, if amazon does not lead me to good products, I will stop using it, and this is why I dislike amazon. There are so many crap products, and the search seems to try to push crap products instead of what i'm actually looking for at a good price.
Yup, makes sense! If you're unable to find models when you use OpenRouter, it makes zero sense to continue to use OpenRouter.
> You take a bunch of providers with not great uptime, put them in a pool and now you get great uptime
Huh? That's not how it works or does it make sense, nor have I've seen anyone use OpenRouter like that.
Comment by pessimizer 21 hours ago
Parent is claiming that the choosing is unreliable. If I rely on a provider to bring me tuna to some spec, but they get it from many different fishermen, it doesn't mean that the tuna doesn't have meet the spec. The complaint is that they're given a bunch of knobs that simply don't work with providers that they could be switched to. That's like saying that I want my tunas to be 20lbs. minimum, and I get switched to a provider that doesn't weigh their tuna at all.
The choosing is all OpenRouter provides. If it doesn't do that, then what is it good for? If I have to permanently pin the one provider who doesn't ignore what I've asked, why shouldn't I just deal with them directly?
edit: it's really supposed to reduce providers to a commodity market. If you're selling e.g. produce to a commodity market, you can't just ship whatever the hell you want. You ship something indistinguishable from others, or more likely the market itself allows you to grade what you're shipping so it's put into a bin with virtually identical stuff. The customer just buys Grade B Wheat.
Comment by ShalevYoni 1 day ago
Comment by dofm 17 hours ago
This helps forecast one possible future for OpenRouter: they begin to offer in-house provision, and people begin to migrate their uses off the "marketplace" providers and onto the "fulfilled-by-us" provision.
Comment by bbor 1 day ago
if you want to control the routing, you'd lock down providers
I don't want to control routing. I want the model to work how the giant, prominent "BENCHMARKS" section says it works, not randomly have a 100x error rate.If OpenRouter is a marketplace to pick a provider while avoiding huge problems, it is terrible at that job. It surfaces literally none of that info in the top-level list, the graphs below are mislabeled and useless at best, and doesn't notify you of this horrifying situation anywhere, even in passing. There's not even a way to compare providers, AFAICT -- you can only compare models.
This is quite literally the point of OpenRouter.
Their tagline is "better prices, better uptime, no subscriptions". The first two of these directly and inherently contradict your understanding -- neither would be possible if OpenRouter was just a fancy way to change something in their GUI rather than changing the target url of your gateway.Comment by embedding-shape 1 day ago
Why do you care about the public benchmarks at all?
The way companies and effective individual developers use OpenRouter, is that you first create your evaluation framework/benchmark, for your specific tasks and use cases, and make that real easy to run and use various models and providers with it.
Then you run this to gather data. Then you use said data to figure out what works and what the quality/cost tradeoff you want to make is. Then you lock that down in production while you keep iterating on your benchmark to make it match with real-world use cases and keep adding the new models that pop up.
I don't think anyone serious is just willy-nilly making individual requests against OpenRouter and similar platforms, get a "feel for a provider" then use only that provider. Not only would it be wildly inefficient, but also you need hard numbers to compare so you can make informed choices.
For this process and workflow, OpenRouter is great, because adding/changing providers and models is essentially changing two strings, rather than having a adapter for each platform you want to try out.
If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRouter, it doesn't make much sense, but use the provider the model maker has available, as almost all of them run their own endpoints.
Comment by bbor 1 day ago
I don't think anyone serious is just willy-nilly making individual requests against OpenRouter
Despite your confidence, that is indeed the basis of this massive corporations entire business plan. If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRoute
If OpenRouter is only for bad accuracy, they should say as much and fade into deserved obscurity.Comment by porridgeraisin 1 day ago
> If openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large budgets, it should say so and thus fade into deserved obscurity.
That is the way LLMs have to be used for highest reliability. While the term stochastic parrot has been co-opted by unreasonable LLM skeptics, that is indeed what LLMs are. You have to have grounded evals that check outcomes if you want to use them reliably - or a human in the loop works too.
The more general your family of tasks, the less likely you can make automated evals. So for "general" coding agents, you need a human in the loop that can verify and it's not that easy to write an eval.
But if you have specific tasks, then you can spend the time to make a eval, and then you can optimise the way you use the LLM and get extremely good success rates. It's not like it's black magic. Nor does it need large budgets.
> that is indeed the basis of this massive corporations entire business plan.
No. Individual developers using codex (for extremely underspecified general engineering) needs human in the loop, is not amenable to evals but is only a fraction of all LLM usecases.
Comment by embedding-shape 1 day ago
Comment by infecto 1 day ago
10month old account with 20k karma. Low value rubbish postings as a professional user. Sad.
Comment by porridgeraisin 1 day ago
Proper LLM adoption beyond fucking around with claude code or github copilot is so low and is only going to go up as people figure it out. I also think the new cloud agents thing might accelerate adoption among these companies, since it's a bit more plug and play. But they are more likely to be stingy about it, so I'm not sure about the high margin expectations of certain model families.
[1] not witch, but BFSI. Surprisingly parts of witch companies have it down to science already, they embedded openai or anthropic or startups like devrev more than a year back.
[2] again not high fly SF companies, BFSI.
Comment by lelandbatey 23 hours ago
That's not quite true. The only thing they don't show per-provider is benchmark data, cause I don't think they are doing continuous benchmarking of each model from each provider, as I assume they feel that's too expensive. You can see hugely detailed breakdowns for near-time metrics per provider for any model by visiting the page for that model on Openrouter. For example see the page for Qwen 3.8 27B: https://openrouter.ai/qwen/qwen3.8-27b
Some of the killer stats they show per provider:
- Pricing: Effective price accounting for cache hit rate, by provider
- Performance: Throughput in tok/s, latency, E2E latency, tool call error rate, structured output error rate, and more; all per provider.
- Uptime: You have to click on the provider to see their specific uptime, but doing so does show the last-7-days uptime, and you can click to see more.
Comment by Gracana 23 hours ago
Comment by numlocked 22 hours ago
Comment by bbor 9 hours ago
Comment by kelvinjps10 22 hours ago
Comment by Aurornis 1 day ago
I’ve tried using un-pinned models and the experience is exactly as you described: Some providers are so unreliable that the majority of requests fail. Some providers do weird things like abruptly end the response (which I get billed for and have to re-submit). Some providers are clearly running heavily quantized versions of the model because their eval performance is terrible. Some providers advertise features on OpenRouter but will reject those requests when submitted to their API.
So pinning is the way to go.
Comment by highfrequency 1 day ago
Comment by arjie 1 day ago
If I could pay per request without maintaining a balance or credit card out of a single wallet (using crypto or something maybe) I would happily simply write the integration myself because OR’s caching is often not as good without some hoop jumping.
Comment by Mairoce 1 day ago
Comment by fc417fc802 1 day ago
How many account credentials, balances, and tokens do you want to maintain? Even without automatic failover services such as openrouter are still incredibly useful.
Personally I pin a single vetted provider in the interest of minimizing risk.
Comment by miki123211 1 day ago
For truly production use cases, use Novita, Fireworks, Toghether or something of the sort.
Comment by nacs 1 day ago
> Fireworks scored 46% on TAU, a 30 point gap
Another surprise was DigitalOcean being bottom of barrel too.
Companies are apparently willing to risk their brand name by being deceptive about these heavily quantized/flawed model-serving.
Comment by svachalek 1 day ago
Comment by tandr 23 hours ago
Comment by FergusArgyll 19 hours ago
The providers should be benchmarking their offerings daily
Comment by random3 1 day ago
Comment by kadoban 1 day ago
Comment by artursapek 1 day ago
Comment by drakmail 18 hours ago
Comment by szundi 23 hours ago
Comment by jmward01 1 day ago
Comment by numlocked 22 hours ago
Open to feedback on how to make this better.
Comment by jmward01 14 hours ago
- Customers being able to decide their own routing with true logic is a huge feature. Open router provides the seamless switching/api, route switching decisions are available in a client.
- Similarly, providing hooks at this level gives a chance for stats/other things that are hard to plug into prod code elsewhere
- a true middle man hosting for other things like MCP may also turn into a real win once it is implemented.
Just a random thought though. My point about quality/cost being clobbered by bad providers remains. The fact that cache and quality is badly handled makes me doubt that training data choices are being respected. You need a more public trust/certification process for providers with real teeth when they cheat. I'm going to wait a bit to see how things evolve and check back later.
Comment by randomblock1 17 hours ago
Comment by ElectricalUnion 18 hours ago
Can I block providers (for a model, not in general) that set up cache write cost when the mode is free cache writes?
Comment by stavros 20 hours ago
Comment by vova_hn2 1 day ago
Providers probably serve quantized versions without disclosing it. Which is a real shame, because for certain tasks I would be perfectly willing to trade accuracy for cost. But, unfortunately, it is impossible to explicitly choose how quantized do you want your model to be, unless you are running it yourself on your own (or rented) hardware.
BTW, does anyone knows if LLM Gateway suffers from the same issues? Currently looking at trying it, but haven't got to it yet.
Comment by joelthelion 1 day ago
It should be OpenRouter's responsibility to protect you against it, by regularly benchmarking providers and giving you the control to avoid bad providers. In fact, that's a big opportunity for them, since it justifies their place as a middleman between users and inference providers.
Comment by ActivePattern 1 day ago
If OpenRouter wants to succeed as a business, they need to be auditing the providers they connect to (i.e. benchmarking) and removing fraudulent ones from their service.
Comment by ipaddr 1 day ago
Comment by pessimizer 21 hours ago
Or by fining them, or getting rid of them altogether.
Comment by joshheitzman 23 hours ago
In my coding agent harness I've included 25 open weight providers mainly because I keep having to find new ones when what was previously a great combination of model and provider becomes pretty bad. vllm has defect that causes reasoning to get dropped much of the time for the GLM family of models. sglang has a defect that causes the elements of array args to get dropped for the deepseek family of models. Some providers need some very specific additional config passed through for reasoning to make it back to the model.
I've not tried OpenRouter as adding yet another layer will just make it that much more difficult to get a model and provider combination working well.
I suspect people's bad experiences with open weight models have a lot to do with these headaches. Finding a good model and provider combination is pretty tedious and so far its been a never ending process. I'd really like to host my own models but it isn't economically feasible for one person for the open weight models that work well (i.e. the 300B+ ones).
Comment by desterothx 1 day ago
Comment by kadoban 23 hours ago
Comment by neya 1 day ago
I used to get content: "" all the time and I used to triple check my code to see if I was doing something wrong until I realized most AI providers in general have vibe coded their infrastructure as well and it is just a futile attempt to even fight it.
Comment by sandinmyjoints 20 hours ago
Comment by wren6991 21 hours ago
Another common annoyance is having a request go to a provider that dribbles out ~1 tps (even for small models like DeepSeek V4 Flash). If you cancel the request, you still get charged for the prefill and the handful of generated tokens. If you don't cancel the request, you might be waiting 10 minutes for the turn to finish.
The overall experience is pretty good, and it's the best way to try new models, but they don't appear to do any real vetting or apply any quality standards to their providers, and occasionally it bites you.
Comment by rolfus 1 day ago
Comment by aembleton 1 day ago
Comment by bko 1 day ago
Android app called "AudioRun"
https://play.google.com/store/apps/details?id=com.audiorun.a...
Comment by binarymax 1 day ago
Comment by ndr_ 7 hours ago
For my contribution to OpenAI's gpt-oss red-teaming competition (published as arXiv:2510.01259), I initially used OpenRouter with DeepInfra and Together AI. The results were much too noisy to draw reliable conclusions from, and markedly different behavior through AWS Bedrock was the last straw. I ended up renting GPUs through vast.ai and running the model myself with vLLM.
Comment by numlocked 1 day ago
Thanks everyone for the feedback here. Some of this we are aware of, some of it we aren't. Some we can fix, some of it is inherent to inference (and we in fact improve the situation dramatically).
Philosophically, at OpenRouter we are trying to do two different things, that are sometimes at odds with one another:
1. Let you use a lot of capacity across a lot of providers, in a way that "just works" and you don't need to worry about it.
2. Have a huge variety of inference available so you can pick radically different price/performance tradeoffs, data policy decisions, geographic destinations, inventive hardware, etc.
These are inherently odd bedfellows, and we are still very much improving how we can make both of them true at the same time.
Some quick thoughts on the article itself:
1. Benchmarks: YES! Providers benchmark differently. We run benchmarks on the live endpoints continuously, monitor the median performance, and kick providers out of the default routing pool if they vary by more than a standard deviation. We work hard (and continue to invest) to make sure that providers serving sub-par inference can't game the system, and that our routing actively avoids them. So the chart is accurate (it's our chart) and it actively influences our routing decisions!
2. That is bad and we will fix it. Sorry.
3. When we on-board providers we run essentially the same test as the author did to verify that the param is working as expected. If it isn't, we don't launch the provider. However this is not one we are running constantly in production. We are working on making this more robust in general and I do believe is fundamentally solvable in a way where it will "just work".
4. We 100% agree that users should not filter by quantization. It's a bit of a legacy concept in general; there is a huge amount of code between "model weights" and "inference API" and in almost all cases quality degrades in that part of the stack, NOT in the model weights themselves.
5. Hmm...we will dig in here. We monitor tool calls in real time and route around providers that are regularly mis-parsing tool calls. So you should get a very low rate of these in general. Another area we have invested a lot in: https://openrouter.ai/docs/guides/routing/auto-exacto
6. We will dig in here as well. I'm surprised this is happening frequently enough to be noticeable. We eat the cost when the finish reason is an error, but not when it is "stop". Perhaps we can expand our "insurance" program: https://openrouter.ai/docs/guides/features/zero-completion-i...
7. Will investigate.
8. We attempt to heal these, but obviously missed some. Will fix.
9. We do not rate limit by IP. Would love some more information here, as that is very surprising.
10. Ugh. That sucks. I'm sorry. We are introducing QoS tiers for production apps, which will address a lot of this.
Comment by mixu 20 hours ago
I'm glad to hear that y'all are doing this, as I was unaware that this was something OpenRouter does. I was surprised and disappointed that there are so many problematic providers that it seems like community best practice [1] is to ban somewhere in the realm of 5-6 providers. Would it be possible to provide some way to express an even stronger preference for high quality providers? E.g. "only route to first party for this model" or, "cost, but don't route to providers that more than x% worse than the first party". I'm sure something like that can be done via the API but I haven't found a UI way to do it - and having it in the UI would go a long way towards feeling like OpenRouter is looking out for me/helping solve the problem as opposed to leaving it to me to have to figure out.
[1] https://www.reddit.com/r/LocalLLaMA/comments/1mk4kt0/be_care...
Comment by pingtoven 1 day ago
(screenshot showing our internal testing of deepinfra image inputs: https://raw.githubusercontent.com/ping-Toven/images/main/ima... )
Comment by mmoustafa 22 hours ago
This was meant as more of a technical reference, sorry you had to wake up to a PR drill lol
Comment by numlocked 22 hours ago
this is amazing!!
[8:19 AM]The first obvious win is routing around providers that arent handling image inputs correctly. That should be straightforward
[8:19 AM]The effort param stuff...I thought we had addressed that, but will dig in. This is incredible feedback
[8:20 AM]We should hire this guy.
Our goal is to get better, fast!Comment by frenchtoast8 1 day ago
Comment by numlocked 1 day ago
Comment by frenchtoast8 1 day ago
Comment by numlocked 22 hours ago
Comment by noahbp 23 hours ago
Comment by noahbp 23 hours ago
To this day, your in-chat “Report an Issue” button still does not work consistently, and I am still billed for empty responses from many image providers.
Comment by numlocked 22 hours ago
Concretely, we pulled data on the last few days of image gen requests in our chatroom (50,893 requests). 5,846 got a text response (which is frustrating, I'm sure) and were billed for text appropriately. The model did not generate an image.
There were 16 requests where a customer was billed, but neither an image or text was returned. Those should not have been charged, and we'll see if we can either fix that issue or ensure that customers aren't charged.
Comment by numlocked 22 hours ago
Comment by epistasis 1 day ago
One thing to note about the first graph: nobody is doing as well as the first part on tool calling, and it's not close.
This might be the fault of the other providers, but it's probably just something slightly different that the first party does with the model inference program than anybody else, and that's not sure to weights it's due to vLLM twiddling (or whatever) and probably becuase the first part actually uses their own customized inference program rather than the standard methods that all the third party providers use. This isn't nefarious, it's just the challenge of these sorts of stochastic systems.
Having been in science for decades now, and seen benchmarking across many different fields, these results are completely expected for me. LLM serving is not mechanical, it's hard to get right and has lots of unknown footguns. Even something as extreme as scrambling a matrix will still likely get results that are nearly as good as normal, and if there's a bug deep in vLLM or the tensors metadata that results in that, then it's going to be pretty hard to find unless you're an active researcher with knowledge of the particular model you're running inference on. I kind of doubt that's happening here, but maybe!
In the scientific literature, when benchmarking methods, everybody's own method performs best in their own hands. Some attribute it to researchers gaming benchmarking for publication purposes, but I think it's just what we see here: the people who made a method are just the best at using it because they know all the quirks and use it best.
Programmers are not used to thinking with that nuance, and jump to conclusions about lying about quantizations, etc., but this is really just an unavoidable part of AI/ML methods: when things aren't perfect they're still pretty good and it's going to take the model creator to truly debug it. At least until the open weights ecosystem gets a lot better at ensuring reproducibility, and model cards are nowhere detailed enough for that to happen yet.
Comment by ndr_ 7 hours ago
Comment by Computer0 19 hours ago
Comment by roger_maddux_iv 22 hours ago
Comment by bsaul 1 day ago
Comment by john01dav 1 day ago
Comment by maeln 1 day ago
Comment by vinhnx 1 day ago
Comment by dools 1 day ago
Comment by JaceComix 1 day ago
It's been a few months since I looked around at this topic, but Fireworks and Openrouter were the two options I (briefly) tried.
Comment by dannyw 1 day ago
Comment by JaceComix 1 day ago
Comment by danvdb 1 day ago
Comment by nacs 1 day ago
Comment by james-bcn 1 day ago
Comment by maxcoding 1 day ago
Comment by lukasbm 1 day ago
Comment by bakugo 1 day ago
If you use any other meta-provider that routes your requests to third party providers, you'll likely face the same issues. If you try using any of those providers directly, you'll likely face some of the same issues as well, except you won't have the option of quickly swapping to a different one and taking your credits with you.
Extreme variance in quality and feature support per provider is probably the biggest obstacle holding back adoption of open weights models.
Comment by TZubiri 1 day ago
Not sure why people are drawn to this particular blunder. The promise of vendor neutrality maybe? I'll take working product over vendor-neutral slop anyways.
Comment by unscaled 1 day ago
They used a closed list of 3 vendors in prioritized order, and got 429ed out of two of them, while the third one stopped serving the mode.
This is less of a problem if you're running an agent locally and routing your problem to OpenRouter - you can pin to one or two models for consistency and just switch models when something goes bad. But the article is specifically about production traffic.
Comment by joshheitzman 22 hours ago
That's a serious questions that I really interested in the answer too. I have 25 providers included into my coding agent harness not because I care about vendor neutrality, but because I have to keep adding new ones as inference quality degrades at the providers I was using. Its quite tiresome.
Comment by bbor 1 day ago
Never before have I heard this sentiment, NGL. Vendor-neutrality has been an OS(/FLOSS) darling for, well, the whole time.
RE:"single vendor", if this post is to believed then you might have picked one that has 100x the tool calling errors for the next SoTA model, if your single vendor serves the next SoTA in the first place. It also completely erases the notion of competition driving down prices -- that would only hurt you in the short term, but obviously would ruin the whole ecosystem long term.
I feel like I must be missing something?
Comment by mrngld 1 day ago
The most significant lock-in to me isn't even something you mentioned, but rather it's model related; I personally put a little time into trying to optimize my prompts every time I change models, as they all have their own unique... flavor.
As for tool calling errors, it seems like first party providers are among the best, I got the feeling that's what he was suggesting, though of course that's why you test. You can also go directly to Together.ai or whoever else you please.
Like others have said I think Openrouter seems neat for testing, but even just as a hobbyist I've been drawn to go direct to particular providers due to irritating little issues that I now see just weren't me.
Comment by TZubiri 23 hours ago
Wrapping a specific implementation in a neutral function is something you learn to do in year 1 of programming.
This specific issue and argument I see in lots of different aggregator dependencies, Terraform, LiteLLM/OpenRouter.
They promise to save some hypothetical work in the future if your boss asks to change vendors, and it turns out to be very trivial work that is just a regular part of our programming job, changing a couple of lines in order to change vendor.
It's worth noting that there exists a similar set of technologies with a reasonable tradeoff, using a framework that targets different user-platforms makes sense, write-once and deploy at iOS and Android is a reasonable tradeoff, but because you are deploying to those providers simultaneously and it's a user-choice so you don't get to pick one or the other (without losing clients), there's still arguments to chosing just one and losing market share, or doubling the workload and building native for both, but this is a true engineering choice. I feel like stuff like OpenRouter and TerraForm take elements of these frontend abstraction technologies and wastefully apply them to backend tech.
A particularly egregious case is when there's an aggregation layer for aggregation layers, say, a tool that generates TerraForm or Chef configs, or a tool that generates Docker and Podman containers, or a tool that generates LiteLLM/OpenRouter configs. Sounds dumb, but it happens when there's a market share for it. Can even get to 3 layers deep.
At the foundation might be an aversion to making an irreversible choice, which is an innate emergent psychological phenomenon, but is supported by the Bezos Amazon policy of reversible and irreversible doors. But again, even if you want to be light, using some of these aggregating tools isn't necessary, you can just build on top of a tech, and switch later. The only thing you get with an aggregating layer is that the API ends up being the common denominator so you lose out on the competitive advantages of each choice, or are forced to use even more complex API logic like LLM(commonParam1, commonParam2, vendorParams= {"vendor1"=:{"vendorParam1":"blabla"}} or worse, use hard-coded aggregator provided mappings between the aggregator API and the vendor API that may be incomplete and relies on updates from the aggregator dev.
Less is more.
Comment by jeremyjh 19 hours ago
Comment by TZubiri 14 hours ago
Still the best way to ensure such policy control is to have 1 provider, tops 2 or 3.
Having a router thing that reroutes to 18 different vendors is of course no way to ensure any policy control, you can add all the internal buttons and dials on policy control and ISO and GDPR compliance, but all it will do is (incorrectly) check compliance box and increase compliance risk to the 18 different vendors.
In practice most openrouter users look for the cheapest vendor, and they tend to go for chinese vendors, who love to price dump and don't have the same views on contracts and IP as the west.
Comment by Macha 1 day ago
The models I've mainly been using recently are GLM 5.3 Flash and GLM 5.3. While obviously all these models have some variability, GLM 5.3 Flash feels like it oscillates between "I can't believe it's not Sonnet", but it costs a fraction of that and "This feels like I'm back using GPT-4, why am I even bothering with an LLM?".
Comment by copperx 1 day ago
Comment by SomeonesAccount 1 day ago
Comment by bambax 1 day ago
Comment by bradfa 1 day ago
It reads to me an argument for self hosting, maybe a less capable model to deal with smaller compute resources, but when starting to use that model and inference software to build a benchmark you can easily rerun when you change things to observe the impact of the change.
Maybe that’s a similar level of diligence but they feel different to me.
Comment by mesmertech 1 day ago
"200 OK, no answer" - insane that openrouter's main feature is literally a fallback and streaming doesn't support 200 no content to fallback to another provider or smth.
"rate-limit by IP"... now it kinda makes sense why deepseek v4.1 flash rate limits me on prod but never seems to happen on local. Makes you have to basically pin Deepseek as provider, since I've never had 429 error on them
Comment by aftbit 19 hours ago
Comment by memoryleakgame 1 day ago
Its a google model. The edge cases are crazy to deal with and have taken a long time to find. I am also finding that since it has no fallbacks but no published rate limit I am single handedly taking the model down on what I thought were a reasonable amount of request. There is no other fallback that isn't google. I'm worried that with sustained usage from my users on launch in a few days. Clearly google can't be making that much money on it it im like 10% of the usage and the top 2-3 user of it spending several hundred a month in pre launch testing.
So why would they care to help a small time start up? I will have to jump to a cost effective different model and find the footguns all over again at somepoint.
Not sure where im going with this but just wanted to let others know </rant>
Comment by voakbasda 1 day ago
Comment by memoryleakgame 1 day ago
Yay google, love living in fear of a vendor that they will rug pull me at any point because they can.
Comment by jwxz 1 day ago
I think the rate limit issue happens because of OpenRouter sending so many requests to Google.
Comment by danielmarkbruce 1 day ago
Comment by sinuhe69 1 day ago
--
Update: oh, that is the AutoExacto Benchmarks! Now I see it.
Comment by FranklinMaillot 1 day ago
The worst case of hallucination I had, was DS v4 flash switching to Italian mid-conversation and impersonating a podcast host for no reason.
Comment by habosa 1 day ago
Comment by plandis 20 hours ago
This is surprising to me. Does anyone have a definitive answer that accounts for this difference between providers?
Do the benchmarks that Open Router runs not account for the stochastic nature of the models? Do the providers lie about quantization or context window sizing? Does Open Router not take into account variations for a given model?
I could understand latency/cost benchmarks varying but not the actual generated token responses of the models given the exact same model parameters.
Comment by sheo 19 hours ago
There are multiple types of quantisation: weights and KV Cache. Quantizing kv cache can drastically hurt performance
Comment by SXX 1 day ago
If you for some reason had unusee balance and forget about it; its just gone.
Comment by radicality 1 day ago
Comment by numlocked 22 hours ago
Comment by wren6991 21 hours ago
Hmm? You have the revenue already. I know it's awkward from an accounting point of view, but you already took my money. "Letting" me keep the balance in the account is not generous.
Edit: on re-reading this came out more combative than I intended, sorry. I think what you're doing is reasonable.
Comment by SXX 19 hours ago
At the same time majority of money we pay them usually just gonna be an expense paid to actual inference providers. Then they pay taxes on their fee aka actual profits.
Its understandable, but it dont make 1 year expiration any good.
Comment by SXX 19 hours ago
Comment by epistasis 1 day ago
So I guess I shouldn't be surprised at all to see these benchmarks, but still I am!
There are such huge economies of scale with batched inference that it's clear this sort of service will continue, but it has a lot of growing up to do. Even AWS Bedrock has a Claude that feels different to me, but I haven't had a chance to do actual benchmarks that would show that.
Comment by dannyw 1 day ago
They’re close enough to not matter though.
Comment by ozborn 17 hours ago
Comment by cesarvarela 1 day ago
Comment by celrod 21 hours ago
``` OK.
Let me write.
Let me go.
OK.
Let me write the script.
Let me go. ```
I'm not sure to what extant this is a model problem, vs some providers being fairly broken. If I chose a single provider, I could know how to blame and to avoid them. With OpenRouter, I don't know which provider I was on when this happened.
Comment by fzysingularity 1 day ago
If you’re building vision-native apps, there are so many footguns in vLLM/SGLang serving configurations, let alone the routing/orchestration in providers like OR, that leave the user more confused about the model’s capabilities.
Comment by msp26 22 hours ago
Comment by fzysingularity 19 hours ago
For example, most vision models don't need 256K context length for modes like Qwen3.8-27B when all you care about is single-image captioning, so you can technically save on KV cache. Video reasoning does require that context length, so it's a different set of deployment parameters that need to be enabled.
All of this to say that the providers that offer these models, are simply using vLLM / SGLang, and mostly cater to the text inference use-case (coding, etc). Vision always seems to be a bit of an afterthought.
Comment by msp26 18 hours ago
Do you have any tips for Gemma 4 31B in particular? I quite like the model but I feel like I'm underutilizing my rented GPU hard due to skill issues. Throughput should be way higher than this, right?
---
Gemma 4 31B NVFP4, vLLM 0.29, single B200 (modal), FlashInfer, fp8 KV, prefix caching, 32k ctx
Workload: ~6k-token shared prefix (~98% cache hit) + short input, ~250 tokens out, ~250 seqs running. Get ~5-6k output tok/s, flat from 128 to 384 concurrency.
Spec decode (DFlash, n-gram) didn't give a good boost which is annoying because I feel like this task should be easy for a draft model to predict.
Individual request latency doesn't matter I just need as much throughput as possible. Its only active for a few hours when I need it.
Comment by kristianp 19 hours ago
As a regular user of openrouter I didn't know this info was available. Will have to check it out. I've definitely noticed that a high level of deepseek flash responses were looping endlessly before the 0731 release.
Edit: looks like the diagrams data is from the "Auto Exacto Benchmarks" section of the performance. Looks like they haven't run the benchmarks of deepseek flash 4.1 on the deepseek provider yet: https://openrouter.ai/deepseek/deepseek-v4.1-flash#performan...
Comment by johnsmith1840 22 hours ago
My basics are I have a test suite that: 1. Finds newest models of my versions 2. Inferences every single model and a few providers for each with a short problem 3. Analyze latency and if a model failed the stupid simple questions drop it and the provider 4. Run larger context haystack kinds of problems.
It's cheap and fast less than 5$ so I can do this daily, hourly, whatever depending on how much I care. If it's mission critical I would say you need a two day study running once an hour to know the STD of model variance.
Then lock a top 3 contenders via latency dropping routing.
Is this easy? No. Is it cheap? Also no. Is it better than just using a trusted labs api? Also not really.
But it does give you exponentially more flexibility. Being able to run 10 unique models at the flick of a switch on a problem for pareto front analysis is amazing. And giving a dropdown for customers for multiple model options is powerful.
Comment by mmoustafa 22 hours ago
Example I forgot to mention: `:nitro` ranking is not the fastest, I do a round robin sampling with representative payloads to find out the fastest providers and reorder my list on the fly.
Comment by johnsmith1840 22 hours ago
I considered an automatic promotion path but decided against it I want to actually review the data myself first.
What's your model churn rate like? I was worried about customer experience by same day maybe same work getting a totally different model response (also caching is worse)
Comment by ltononro 1 day ago
Comment by dangoodmanUT 1 day ago
This is actually expected behavior. No content and no tool call is the same as content-only: the agent decided it's done. Anthropic has done this for a while.
Comment by eitally 1 day ago
https://newsletter.semianalysis.com/p/clustermax-20-the-indu...
Comment by nojs 1 day ago
I wonder what tricks one could use to ensure the model is actually performing on par with the reference api, like matching seeds or running exact benchmarks.
Comment by _ink_ 1 day ago
Comment by quietsegfault 1 day ago
Comment by EDEdDNEdDYFaN 1 day ago
Comment by SomeonesAccount 1 day ago
Comment by dnugget 21 hours ago
Comment by alexcz 1 day ago
Comment by polytely 18 hours ago
Comment by ketzu 1 day ago
Yesterday I wanted to do a bit of benchmarking a prompt across multiple models. Small requests. Outside the big providers, the experience became awful. This explains that experience.
Comment by benjbrooks 1 day ago
Source: Enterprise customer doing $XXM annual run rate of inference spend on their platform
Comment by Semaphor 1 day ago
Comment by ghm2199 1 day ago
Noob question: do good harnesses automatically optimize for bad tool calling behavior automatically?
Comment by jwxz 1 day ago
Comment by neilmovva 1 day ago
Comment by jwxz 22 hours ago
I had quite a few Kimi K3 requests served by Sail. Most of these were fine and had the reasoning traces (I like reading them), however I noticed that some requests were being generated unusually quick and had no reasoning trace, which led me to believe another fallback model was being used, since Kimi K3 always emits reasoning tokens.
Comment by rpjt 1 day ago
Comment by anguishe 22 hours ago
Comment by Krisso 1 day ago
Comment by SpyCoder77 19 hours ago
Comment by jwrallie 1 day ago
Comment by jeremyjh 1 day ago
Comment by kinard 1 day ago
Comment by philipp-gayret 1 day ago
Comment by bakugo 1 day ago
https://app.answerhq.co/openrouter-ai/articles/credits/credi...
Comment by irthomasthomas 1 day ago
Comment by bluepeter 1 day ago
Comment by ofisboy 1 day ago
Comment by try-working 17 hours ago
Comment by CubsFan1060 1 day ago
On top of that, near as I can tell, there are no protections for your API key. No restrictions by country, IP, etc...
Comment by nhecker 1 day ago
Comment by CubsFan1060 1 day ago
I get it, the API key was my responsibility, but, setting the dollar limit is exactly the guard they suggest against that.
Comment by numlocked 1 day ago
Comment by CubsFan1060 1 day ago
Comment by srcreigh 1 day ago
Comment by hqm_ 1 day ago
Comment by probabletrain 21 hours ago
Comment by seanieb 1 day ago
Comment by user- 1 day ago
This aligns with what I've experienced using openrouter.
Are there competitors that handle these same issues better?
Comment by ggdG 1 day ago
Comment by Havoc 23 hours ago
Comment by claudeIsDown 1 day ago
Comment by aranaur 1 day ago
Is it, though?
Comment by anonzzzies 1 day ago
Comment by joelthelion 1 day ago
I feel there is still a lot of progress to be made before we can really trust LLM providers.
Comment by xienze 1 day ago
* Hardware availability
* Competency
* Scruples
Comment by MattyRad 1 day ago
Comment by owenshen24 17 hours ago
+1, stuff like this is a big turnoff for reading
Comment by fl0id 1 day ago
Comment by system2 23 hours ago
Comment by dvdkon 23 hours ago
Comment by system2 23 hours ago
Comment by MallocVoidstar 1 day ago
This isn't necessarily an OpenRouter issue, Google's Gemini will sometimes do it on their own API. Since they summarize reasoning this means you pay the whole cost and can't get anything out of it.
Comment by teaearlgraycold 1 day ago
Comment by grim_io 1 day ago
I was so disappointed that I won't consider any of them for at least a few years.
Comment by agcat 1 day ago
Comment by polski-g 1 day ago
Comment by bbor 1 day ago
I feel like it's absolutely insane that some providers serve the same model with much less efficacy. That doesn't make sense to me. What's going on?! I'm suddenly feeling intense shame for having routed all my non-subscription usage through them so far, and honestly some white hot anger that they would blatantly lie about something so important.
What am I missing? Is this really true?
Comment by numlocked 1 day ago
And errrr...yeah...that auto-exacto performance chart is both 100% useless, and totally unclear. We will get that fixed. But under the covers it is doing a lot of valuable work! https://openrouter.ai/docs/guides/routing/auto-exacto
Comment by bbor 9 hours ago
Submitting an application to your product role now. "All of the models are at least okay" is heartening, but there's a whole bunch of fun places to take this.
¡Viva La OpenRouter! (again)
Comment by rima_667 8 hours ago
Comment by joelsol 1 day ago
Comment by lluisantoni 1 day ago
Comment by lellow 1 day ago
Comment by abuds 3 days ago
Comment by anik200 3 days ago
Comment by npn 1 day ago
It is pretty trivial to pin a single provider for a model. Better yet, instead of calling the model directly, use presets instead. You can easily change the setting on openrouter without having to update your app every time.
Comment by ptsneves 1 day ago
Comment by hadeer626 3 days ago