Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
Posted by adam_rida 1 day ago
I’ve been building Echo (https://echo.tracerml.ai/), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task.
It started with a simple experiment. I took a group of models, including GLM-5.2, Kimi K2.7 and others, and ran them on the same evaluations. Then I measured what would happen if, for each problem, you somehow knew in advance which models would be useful and how their outputs should be combined.
That hypothetical system performed substantially better than any individual model in the pool. Of course, it is not something you can actually deploy because it relies on knowing which decisions were good after seeing the result. Echo is my attempt to recover some of that advantage without having that information in advance.
For each request, Echo decides how much computation to allocate, which models should participate, and how their work should be combined. Some prompts may only need a relatively small amount of inference, while others benefit from multiple models working on different parts of the problem.
One thing that surprised me while building it was how complementary the models are. A model that is clearly weaker overall can still be extremely useful on particular problems or as part of a combination.
On my first evaluation mix, Echo consistently performed better than the best individual model in its pool. It also reached roughly the same aggregate result as Fable, which I used as one of the stronger comparison systems, at around one third of the inference cost.
There are still some cases where Echo makes the wrong allocation or combination decision. I’m currently spending a lot of time understanding those failures, as well as testing whether the same approach holds up on coding and agentic tasks where measuring the quality of each decision becomes much harder.
I built a chat interface (echo.tracerml.ai) and an OpenAI-compatible API (https://echo.tracerml.ai/docs/api) so the system can be tested outside the evaluation setup.
Here is a short/high level video on how it works: https://www.youtube.com/watch?v=lJFJSvOdXhg
I wrote up the evaluation methodology, individual model results, costs and current limitations here: https://echo.tracerml.ai/eval
I would love for you to try it! Especially if you hit any weird failure cases or places where the allocation looks unintuitive.
Comments
Comment by fmx 17 hours ago
I've literally taken one step on your website - the one your site design invited me to take - and immediately got tripped up. I'm not coming back.
Comment by ciefa 13 hours ago
Comment by adam_rida 12 hours ago
Comment by rufasterisco 11 hours ago
Since you are fixing dark patterns, can you also put data privacy links somewhere other than on the gdpr banner?
Once clicked is impossible to see them, or at least I could not find them.
Comment by adam_rida 7 hours ago
Comment by barapa 15 hours ago
Comment by pizzao 15 hours ago
I dont know, it's something I can understand as someone also building AI stuff
Comment by idonotknowwhy 14 hours ago
This dark pattern is reminiscent of those online test sites in the 2000's where you spend 10 minutes filling out some quiz, then get prompted for an email address to see the results.
https://chat.mistral.ai/chat <- let me chat and actually responded without signing up.
MoonshotAI had this fake chat box dark pattern.
So I signed up with Mistral instead.
Comment by lukan 15 hours ago
Comment by boesboes 15 hours ago
Comment by adam_rida 1 day ago
i am going to try to address a couple of topics that came up often:
- i'll keep publishing stronger evals, including more difficult coding and agentic benchmarks, to map out more precisely the differences with sota
- the public eval dashboard will keep expanding and be updated (very open to more benchmark suggestions as well!)
- some people found issues in the eval dashboard ui and the sign up flow, should be now all fixed in prod
some important precisions as well:
- NO credit card is required to try Echo
- each acount includes 10$ of free credits to try on both the API and the chat
on the approach itself: the idea i'm exploring is more broader than model routing, i'm looking at how to allocate inference efficiently across open-weight models, deciding not only which models to use, but also how much computation a request deserves and how intermediate work should be combined.
ensembling by itself is not new. since random forests and probably even before in statistics/classic ml we knew that bringing multiple models together can outperform individual ones. the interesting problem for Echo is how to model and leverage this without paying the full ensemble cost at each request.
while there are conceptual similarities with systems like Fusion or Fugu, the architecture and optimization objective are different.
thanks again for all the thoughtful feedback.
Comment by user_7832 21 hours ago
Comment by adam_rida 7 hours ago
Comment by troupo 19 hours ago
This isn't required either. Of course there's an xlcd for that: https://xkcd.com/936/
Besides,
--- start quote ---
Using complexity requirements (that is, where staff can only use passwords that are suitably complex) is a poor defence against guessing attacks. It places an extra burden on users, many of whom will use predictable patterns (such as replacing the letter ‘o’ with a zero) to meet the required 'complexity' criteria.
https://www.ncsc.gov.uk/collection/passwords/updating-your-a...
--- end quote ---
Comment by user_7832 13 hours ago
Oh, I'm well aware (with the concept and comic both). My passwords that I set for myself almost always are like that. It's just that I wasn't sure if OP was aware of it, due to what their site was asking.
Comment by RugnirViking 14 hours ago
Comment by bitexploder 12 hours ago
Seriously though, the best password boxes are one that have an entropy meter and check that your password has never been in a breach.
Comment by jorisw 14 hours ago
Comment by rahulroy 17 hours ago
Comment by ciefa 13 hours ago
You had my interest, now not anymore.
Comment by afzalive 13 hours ago
Comment by cheema33 1 day ago
I am guessing this is not targeting those of us on the heavily subsidized $200/mo plans. Sure, these plans may be temporary, but none of us really know how temporary they are. Until then, 1/3rd of the published API pricing is not very appealing.
Comment by stilesja 1 day ago
Comment by jambalaya8 21 hours ago
Comment by Foobar8568 19 hours ago
I have in mind 50k-100k ish for 3d studio max or was it softimage? (Well seems softimage https://www.awn.com/animationworld/siggraph-news-announcing-... ).
So... Basically we are back to this era.
Comment by stkdump 18 hours ago
Comment by cube00 17 hours ago
The value was the multiple CDs of MSDN documentation and code samples which where very handy considering the slim pickings on the internet in 1995.
These weren't included with an IDE perpetual license retail box.
Comment by jambalaya8 8 hours ago
Comment by mattjoyce 13 hours ago
The other weird thing about the subs is that if the agents aren't grinding if feels like I'm losing money.
Comment by DougN7 21 hours ago
Comment by good8675309 22 hours ago
Comment by sunaookami 15 hours ago
Comment by Bombthecat 15 hours ago
And it was damn cheap too! The main page cost was like 5 dollar.
Comment by hmottestad 21 hours ago
Comment by sheepscreek 19 hours ago
Dude that topped Meta's tokenmaxxxing board before it was shut down used 265 billion tokens in a month. I kid you not.
Comment by killingtime74 1 day ago
Comment by rodrodrod 1 day ago
Comment by lnrd 18 hours ago
Comment by nickthegreek 1 day ago
Comment by monk_grilla 1 day ago
Comment by byzantinegene 23 hours ago
Comment by nl 20 hours ago
Comment by lelanthran 16 hours ago
I can not imagine some shelling out $200/month and then using that product lightly.
Comment by nl 15 hours ago
The people paying for the plan are not the same people using it.
Of 6 people I have data on the $200/plan only 2 regularly use more than $400 value.
> I'll bet the other way: the plan is not cost effective unless you are coding
The person I've personally seen use the most tokens isn't a coder. They do the "second brain" thing and wow it uses a lot of tokens.
They believe in the value, and TBH I've seen them do some pretty interesting and impressive things with it.
> the most junior developer, so green they almost need mowing, are going to throw the agents into a loop
I think this is also true.
But loops actually hit the cache a lot and most people who are calculating the value they are getting from a subscription aren't taking this into account.
SemiAnalysis published a snippet of their analysis, and they believe their tokens are an effective price of $0.99/million, rather than $20/million the naive pricing calculation would give you.
Comment by bensyverson 22 hours ago
Comment by byzantinegene 23 hours ago
Comment by suby 22 hours ago
It's not a 98% margin loss if your users are unwilling to pay 50 times the cost that they were previously paying, and if they have other options like open source providers. The calculus isn't so simple because some portion of users would switch to API, and so it's about how many would continue using the service rather than leaving for a competitor.
I'm aware they need to recoup the enormous cost of training and data centers, but on a purely inference cost level I'm not convinced that the 200 dollar plans are unprofitable.
Comment by nl 20 hours ago
Especially considering not everyone is tokenmaxxing, and in most parts of the world people take leave and companies do not cut their subscriptions.
I suspect they are priced to have a lifetime average price/token amount that is roughly break-even, or maybe a slight loss leader.
> have seen Dario say in multiple interviews that they are profitable on inference, which maybe he was only meaning to refer to API usage, but that's not the impression I got.
I think he does mean API usage. Don't forget they can (and do) adjust the number of tokens you get on each plan at any time to adjust their margins on those.
That means he knows that is controllable, and it only the underlaying inference that defines the succes or otherwise of the company.
Comment by kelnos 18 hours ago
Exactly. I have the Claude $100/mo plan, and use it moderately for open source hobby stuff. I still haven't dipped my toes into the Fable pool, but I always use Opus 4.8 on xhigh, and I never hit my limits.
On the other hand, though, there have been times when I've looked at /usage for a long-running session (e.g., 7-10 days, after it's compacted a few times), and it showed I'd used ~$450 worth of tokens just for that session. So I'm clearly getting value for the money here when it comes to the subscription cost. But I still don't hit limits, so...
Comment by victorbjorklund 22 hours ago
Comment by orsorna 22 hours ago
Are your thought patterns worth 9800 dollars a month?
What's the RoR on analyzing those thought patterns?
Comment by aroman 21 hours ago
Comment by somenameforme 21 hours ago
Comment by davedx 19 hours ago
Comment by lnrd 18 hours ago
Comment by hahahaa 20 hours ago
Comment by neonstatic 21 hours ago
Anthropic emailed me today:
Fable 5 moved to usage credits on July 20. It is still available to you, but it requires pay-as-you-go usage credits and is not included in your subscription rate limits.Comment by MertsA 21 hours ago
Comment by dd8601fn 19 hours ago
The half usage limit still applies.
Comment by agar 21 hours ago
Comment by nl 20 hours ago
For Premium seats it is.
https://support.claude.com/en/articles/15424964-claude-fable...
Comment by runtime_lens 20 hours ago
Comment by eunos 17 hours ago
Comment by yieldcrv 19 hours ago
the supercycle is on device models, and one of those evolutions is models baked into chip die, and you just upgrade chipsets every few years instead
so it's the hyperscalers that will take the L in that environment
Comment by kamranjon 1 day ago
Anyhow, this kinda reminds me of that quote about architecture: "We replaced our monolith with micro services so that every outage could be more like a murder mystery."
Comment by adam_rida 1 day ago
It currently exposes 907 stored rows across seven benchmark families, with prompts, outputs, grades, and cost records. More benchmarks are coming soon.
Echo does not disclose its per-request routing decision because that policy is the product. We can, however, publish some of the eligible open-weight model pool, version dates, aggregate allocation mix, and evaluation settings without exposing the request-level recipe.
New video is also being made.
Comment by dannyw 23 hours ago
My honest advice: that's going to pull away a decent amount of potential customers, even though I think your idea/concept is fantastic.
For example, if we were to consider it for Canva, observerability and full transparency is critical requirement; we can't accept not knowing which model serves a request. Both for legal/contract reasons, co-ordinated capacity planning with API providers, or even just evaluating our prompts and harnesses; and debugging/tracing results that went wrong. So that renders it out of consideration; and also suggests some kind of adversarial relationship where customers aren't trusted with critical information.
I definitely understand you need to keep business value, but I don't think hiding which model a request is routed to, is the right one, or at least if you want to expand to bigger potential customers / more advanced LLM deployments.
Comment by lionkor 15 hours ago
Comment by monk_grilla 1 day ago
I have not had need for a router product thus far so excuse my ignorance if this is standard, but how could I possibly use and improve a product built on a router like this if I am not permitted to see which model served my request? If I got a bad answer back in my LLM-powered app, do I really have no way of knowing which model was responsible?
Comment by seizethecheese 1 day ago
Comment by guessmyname 1 day ago
The benchmarks are here → https://echo.tracerml.ai/eval/
They are not good benchmarks but at least they exist.
Comment by seizethecheese 1 day ago
In my project, I wasted a huge amount of time trying to improve GPQA Diamond results above ~93% range. I realized my mistake when Fable dropped and made no improvement on this benchmark vs. Opus.
Comment by yorwba 1 day ago
Comment by seizethecheese 1 day ago
Comment by codekansas 1 day ago
I just wish this were solving an actual problem rather than being a fairly transparent attempt to say something approximating, "Hey VCs, OpenRouter just became a unicorn but I can basically vibe code it"
Calling it "Fable-level" feels intellectually lazy / dishonest, but then again, what do you expect when there's so much money on the table.
Comment by vunderba 1 day ago
Comment by throw10920 1 day ago
Comment by vunderba 1 day ago
This space is so crowded it feels like I see a new "model router" pop up every few weeks.
Comment by seizethecheese 1 day ago
Comment by dimitrios1 1 day ago
"microservices turn function calls into distributed computing problems"
Comment by WD-42 1 day ago
seem very confusing to grug
Comment by XCSme 21 hours ago
Comment by j45 1 day ago
Eval tests while giving general indicators might not be similar for each use case.
Comment by guesswho01 18 hours ago
Comment by subygan 1 day ago
Else, you break the cache by doing a round robin of the same conversation across different models. Likely you'll end up paying more than what it would've cost with a cache aware system
Comment by lbriner 17 hours ago
Comment by hahahaa 20 hours ago
Comment by tj800x 1 day ago
Comment by adam_rida 1 day ago
Echo also starts users with free credit and does not require a credit card to try it. The current signup flow did not make that clear enough, so we are fixing that presentation too.
Thanks for calling both out.
Comment by seizethecheese 1 day ago
Comment by DatCodeMania 23 hours ago
Comment by cynicalsecurity 1 day ago
Comment by dluan 1 day ago
Comment by glasss 1 day ago
Comment by onlyrealcuzzo 1 day ago
You needed to search all of them to find something decent.
That's roughly analogous to today. Ignoring cost, you'd be way better off asking all the LLMs to solve a problem (like coding) where you can verify the answer.
So the question is, for things like that -> can a group of models perform better than frontier models, especially at a reasonable cost?
Fable is not a great value, so unless you're trying to find answers to Erdos questions, you can probably do better on cost.
You can probably typically ask 3 or 4 of the top Chinese models for an answer and get a response for the same Fable question... Given that Fable isn't that much better, it's not surprising you can do better for a large subset of problems.
Comment by Terretta 1 day ago
Precisely.
Comment by user_7832 21 hours ago
The more vague and non committal and hand-wavey and subjective the field for AI to answer, the better the results (imo).
Comment by rablackburn 21 hours ago
Because the correct response to that query is "I have no idea -- you will need to provide more information"
and LLM Agents suck at that.
Comment by ai-x 23 hours ago
Comment by joe_the_user 21 hours ago
Well, search engines are trash again. Perhaps it should come back
Comment by onlyrealcuzzo 13 hours ago
That's mostly because the ability to make money on the web made the web trash.
When you couldn't monetize your websites, everyone's websites were passion projects.
Now everyone is trying to figure out how they can make $8M a year off yet another recipe website.
Comment by ljlolel 1 day ago
we did the same: https://trustedrouter.com/blog/prometheus-2-new-draco-state-...
Comment by ignoramous 21 hours ago
Our whole stack is radically open source — frontend and backend alike, Apache-2.0 licensed — and so is everything behind this benchmark. That is how a benchmark number earns trust: verifiability, not hype.
The repos have since been moved to BUSL-1.1: https://github.com/Lore-Hex/quill-router/commit/8155ac666ae0...Comment by hahahaa 20 hours ago
Comment by gchamonlive 1 day ago
Comment by ignoramous 1 day ago
Comment by fmajid 1 day ago
Comment by johnvanommen 1 day ago
Did OP invent an “intelligence router?”
Comment by alizaki 1 day ago
Comment by moralestapia 1 day ago
Comment by cdelsolar 1 day ago
Comment by smokeeaasd 10 hours ago
The industry has largely focused on building larger models, but your results suggest that intelligently routing requests to the right combination of specialized models can deliver greater gains at a much lower cost.
It also reinforces the idea that weaker models are not necessarily obsolete. They may simply excel in different areas and become much more valuable when combined with others. I'm curious to see whether this still holds for coding and agentic tasks, where choosing the right models is likely much more challenging.
Comment by adam_rida 10 hours ago
on agentic and coding what's make the problem even deeper is the granularity. how and when to use each model and at which layer of abstraction (session, goal, task, turn/tool calling). this is also something we are working on actively!
Comment by slashdave 1 day ago
Comment by meander_water 1 day ago
Comment by seizethecheese 1 day ago
Fusion generates many replies then synthesizes. This adds a ton of latency and cost, so it's going to be better only for cases where you're willing to wait a lot and pay a lot more.
Routers (like this project) are a different thing, they can theoretically improve performance and cost at the same time without increasing latency much. I'm a bit skeptical though, since knowing which LLM is going to be better on a cost adjusted basis is hard (see https://artificialanalysis.ai/models/capabilities/coding?cos..., where the cost per task vs. performance is not what you expect, for example comparing Qwen 3.7 Max to GPT Sol.
A project I'm working on is aimed at improving performance without added latency but from a different angle. Instead of waiting for all replies for synthesis (like OpenRouter Fusion), it streams the "best" reply immediately (using a router to pick the best model) then synthesizes with emoji reactions and optional replies from the background models. It's free to use here with no login: http://pellmell.ai
Comment by Arshad-Talpur 19 hours ago
Comment by huflungdung 19 hours ago
Comment by springtimesun 1 day ago
Comment by XCSme 21 hours ago
Comment by jmspring 1 day ago
Comment by motbus3 11 hours ago
in short, all of them are token diarrhea. Annoyingly logorrheic. they output so much useless stuff that it comes to the point that it makes me thing if that's not on purpose.
it was not because of price, but quality, I started testing other models and honestly, glm 5.2 is FAR SUPERIOR than fable 5 in every aspect. It comes to a surprise when glm 5.2 does not finish the task successfully. Kimi k2.7 does require a bit more of guidance but still better experience than opus 4.8. I have not yet the chance of trying k3.
openai latest model are.... ridiculously bad at software design and implementation.
(note i am only speaking to the domain of my work which has lots of data analysis, machine learning and software engineering.)
Comment by ljlolel 1 day ago
Comment by adam_rida 1 day ago
You also found a real UI bug: the inspector should show both stored answers and currently does not in some rows. We are fixing that.
On the row you reran: the page records a frozen matched run. It does not claim that Fable is incapable of solving that prompt on another run. We are adding repeated matched trials and making run count and variance explicit. Your rerun is exactly the kind of external check the row-level page is intended to make possible.
The broader result remains: Echo is competitive with Fable across the evaluated task mix at materially lower measured inference cost. We are filling in the harder agentic-code evidence now rather than asking anyone to infer it from HumanEval+.
Comment by jdthedisciple 19 hours ago
I find that very off-turning!
Comment by elnatro 19 hours ago
Comment by jdthedisciple 19 hours ago
And still Fable beats it hands down 8-0 in one of them, and is at worst even in some others.
Also it doesn't make logical sense: A router can save costs, yes, but not magically be "smarter" somehow.
That's like selling "free energy".
Comment by saberience 11 hours ago
Literally promising frontier-equal results but at 1/3 price, i.e. cheaper than Kimi K3?
Doesn't offer any real benchmarks or explanation of how this magic trick is accomplished. No credible team or notable scientists behind it...
Comment by Alifatisk 1 day ago
Comment by seizethecheese 1 day ago
Comment by NetOpWibby 1 day ago
Comment by hmokiguess 1 day ago
Comment by tintor 1 day ago
Comment by jmaw 1 day ago
I have often wondered how tools like GHCP choose the best model for the job when set to "auto".
Comment by sudo_cowsay 20 hours ago
Comment by stevefan1999 18 hours ago
Comment by datadrivenangel 11 hours ago
Comment by NoNameAditya 10 hours ago
Comment by blobbers 1 day ago
For example, compute X tokens with model A, then feed those into model B, etc. to get chain of thought through a diverse set of mdoels rather than chain of thought through a heterogeneous chain.
Humans seem to strongly believe echo chambers are bad. Are LLMs the same?
Comment by lelanthran 16 hours ago
Comment by adam_rida 1 day ago
Comment by blobbers 23 hours ago
I'm curious though if these training methods are convergent or are models actually different; just like how in the stock market people think they're "diversified" but the truth is their exposure is likely much more risk correlated than one might think.
In certain situations, one right answer is better than a committee discussing the problem, but in others its sometimes nice to have some alternative methods of solving something. Fun project nonetheless.
My approach to using multiple models has been less about CoT but more about time to first token, and how you can use a small model to start interacting with the user while in parallel the more complex model is building a larger more complex thought. My work on this was primarily for voice backed interfaces before the voice models became quite a lot faster.
Comment by yonatan8070 1 day ago
As I understand it, in an MoE model, you essentially have hundreds of smaller sub-models ("experts") that are good at different tasks, and for every generated token, a single "master" model chooses which ones are most relevant to participate, and you only activate them.
Comment by janalsncm 1 day ago
Even more confusingly, there are older pre-LLM MoE systems which ensemble and pool the predictions from multiple sub-components. For example in a random forest you could take the majority vote of the decision trees or the average of their numerical predictions.
After that, we developed neural net architectures for predicting a single thing like whether the user will click on your ad. An MMoE is in the same family.
And so now we are at massive MoE networks for LLMs which have similarities with MMoE in that the “decision” is about the very next token to predict.
Comment by lukan 1 day ago
Have there been experiments with doing it per task? Like, "oh this is python project, use this model" "oh this is about writing fantasy, use this"?
Comment by janalsncm 1 day ago
I think OpenAI already has (had?) a feature like this called “auto” mode for thinking.
Comment by alightsoul 1 day ago
Comment by thatxliner 20 hours ago
Comment by hahahaa 20 hours ago
Comment by janalsncm 1 day ago
So “1/3 the cost” really depends.
Comment by seizethecheese 1 day ago
Comment by janalsncm 1 day ago
Comment by seizethecheese 1 day ago
Comment by zhonglin 1 day ago
Comment by SubiculumCode 18 hours ago
Comment by alex-moon 19 hours ago
Don't want to derail what you're trying to do with Echo in case I'm wide of the mark, but yeah even in that case, if you hadn't considered that use case for it, I reckon there will, probably inside six months, be a substantial market for non-technical users who are sick of seeing ads in a service they already pay a subscription for, and who don't care what the underlying model is - or, indeed, don't even understand the concept of an "underlying model" because they interface with AI as a product.
You have already taken the HFG idea way further than I had even thought of yet, and I feel vindicated in seeing someone else do it. I wish you the very best!
Comment by islambaraka 1 day ago
Comment by spidercob 17 hours ago
Comment by haris599 19 hours ago
Comment by indiantinker 1 day ago
Comment by bbstats 1 day ago
Comment by jacobgold 1 day ago
But we get ~$2500/mo worth of Fable credits for $200/mo on Anthropic pan? I'm still confused why people (who don't have to use API billing) are chasing open weight models based on cost.
Comment by teruakohatu 1 day ago
Comment by sscaryterry 1 day ago
The $200 odd plans are already out of reach of many, many people.
The attrition of customers if they were to get rid of these subscriptions plans would be untenable.
Comment by recursivegirth 1 day ago
The $200 plans are priced so that the power-users use them and then advocate about how great the product is. If you're buying a $200 plan, you're not doing it because of the price point but rather because of the amount of work it is doing for you.
Comment by trollbridge 1 day ago
Comment by anonzzzies 1 day ago
Comment by combyn8tor 1 day ago
Comment by gruez 1 day ago
https://artificialanalysis.ai/agents/coding-agents#artificia...
Comment by combyn8tor 1 day ago
Comment by dbbk 1 day ago
Comment by eikenberry 1 day ago
Comment by dberg 1 day ago
Comment by recursivegirth 1 day ago
Comment by dang 21 hours ago
Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
Comment by drnick1 1 day ago
This would be like trying to outlaw Linux or peer-to-peer file sharing. It's technically possible to write and pass a law, but it's basically impossible to enforce it.
Comment by switchbak 1 day ago
Comment by trollbridge 1 day ago
Comment by recursivegirth 1 day ago
Comment by Art9681 1 day ago
The frontier providers aren't dumb. They charge what they charge because they know this. If you think Fable is too expensive then the type of problems you are solving don't demand that level of capability.
If you are working on something cutting edge, something truly novel, the cost of frontier AI is well worth its price.
With all that being said. No one is going to complain if we can get the same capability at a lower cost. And I mean true parity. Not trading space for time.
Comment by qainsights 1 day ago
Comment by adam_rida 1 day ago
Comment by raver1975 1 day ago
Comment by adam_rida 1 day ago
Comment by fneddy 1 day ago
Comment by seizethecheese 1 day ago
Comment by bnjemian 1 day ago
While an LLM isn’t what you’d traditionally consider a weak learner, the theorems on learning systems clearly point to them being so in this context. The feigned surprise at combining them to yield better results seems disingenuous.
Even so, the work to predict which models are best suited for which task, how to delegate, and how to combine their outputs is interesting, especially if you’re placing a cost minimization objective on it. That said, this isn’t too far off from what many AI labs are already doing.
Comment by abernard1 1 day ago
It's the natural move that happened after people realized you couldn't throw away half a century of AI research.
Most of these focus on costs. But it is simply the case that the one-shot output did not scale for harder problems on workflows.
Comment by codekansas 1 day ago
Comment by ninjahawk1 1 day ago
I go to the website…and it’s a sign up. I expected a repo. Otherwise how do I use it? As a SaaS? Yeah right.
Oh well I guess at least the benchmarks are good…I find the benchmarks and many are either not present or are not what the title claims.
My main question is how this has so many updoots from HN, probably the passerby not looking closer for sure.
I mean no offense and I really do wish you best on this, but it seems like what we used to call back in the day, vaporware.
Comment by fgoose180 1 day ago
Comment by zuzululu 18 hours ago
Comment by cantalopes 1 day ago
Comment by meowface 1 day ago
Plus Fable is way less annoying to talk to than Opus 4.8. Opus 4.8's writing style is absolutely insufferable. Fable has some of the same quirks but it's way less bad.
Comment by anonzzzies 1 day ago
Comment by jambalaya8 1 day ago
Comment by purplecats 1 day ago
Comment by seizethecheese 1 day ago
Comment by retinaros 1 day ago
Comment by maxdo 1 day ago
Comment by seizethecheese 1 day ago
Comment by j45 1 day ago
Comment by ototot 1 day ago
Comment by code0igx 12 hours ago
Comment by kachnuv_ocasek 1 day ago
Comment by antrichards 48 minutes ago
Comment by hmokiguess 1 day ago
https://www.ycombinator.com/companies?query=tracerml
I don't see it?
Comment by zachdotai 1 day ago
Comment by dang 1 day ago
But in the present case, they're just a startup in the current batch.
Comment by hmokiguess 12 hours ago
Comment by wayknow 20 hours ago
Comment by matchartier 16 hours ago
Comment by alexzhangai 1 day ago
Comment by IrfanD 16 hours ago
Comment by moriwo-dev-ai 1 day ago
Comment by mohammedmsgm 16 hours ago
Comment by mandarinclips 22 hours ago
Comment by theneocorner 1 day ago