Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

Posted by cdnsteve 2 days ago

Counter67Comment27OpenOriginal

Comments

Comment by forgot-my-pw 2 days ago

Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.

Still quite impressive though.

Comment by ChrisArchitect 2 days ago

Related:

Cognition launches new SWE-2 model

https://news.ycombinator.com/item?id=49645443

Comment by varispeed 2 days ago

These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.

Comment by kzrdude 2 days ago

Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?

Comment by varispeed 2 days ago

It is my own anectodal. Few days something I worked on usually got one shotted or got quality result. Today it is very much going nowhere and is stuck in reasoning loops.

Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc.

Comment by EPWN3D 2 days ago

So it's your own anecdotal experience, but someone should build a service to track it?

Comment by singingtoday 1 day ago

There is a service to track anthropic models. Not sure how accurate it is, but my vibe says somewhat.

Comment by cbg0 2 days ago

Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/

Comment by varispeed 2 days ago

That is not conclusive, because they can detect such tracker and route it through proper not quantised model.

Comment by nba456_ 2 days ago

not really

Comment by d_tr 2 days ago

How and why do they get nerfed? To save money?

Comment by johnfn 2 days ago

Models do not get nerfed. There has never been evidence of this. This would be trivial to prove if it were true, and such a proof would be a huge story and scandal to a news market hungry for a shred of a signal on AI's downfall.

This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s.

Comment by 4chandaily 2 days ago

> This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s

Of course, it turned out that this wasn't actually completely BS. We just were accusing the wrong vendor. Not disagreeing with you on models.

Comment by johnfn 2 days ago

I'd be happy to read a source as I am fairly confident that audio transcription -> facebook (or other) ads has never been true.

Comment by 4chandaily 2 days ago

I was referring to LG specifically. Here is a link from https://news.ycombinator.com/item?id=49592375 three days ago.

Comment by johnfn 1 day ago

Wow totally missed this, thanks for sharing

Comment by esafak 2 days ago

Yes, they can, through quantization. Many providers of open source models openly serve quantized versions; check openrouter.

Comment by varispeed 2 days ago

Yes. They save on compute and customer has to use more tokens to achieve their goal which means more profit.

Comment by cliche 2 days ago

Profit? I thought these companies were making a massive loss

Comment by ricardobeat 2 days ago

One theory is that they start serving at full precision, then quantize to save on costs as adoption grows. It's kind of a conspiracy theory atm, but I have definitely felt it - I had a large project done on Opus 4.8 release day, a week later it was struggling to complete partial tasks in the same area.

Comment by samusiam 2 days ago

Which is pretty much a useless (i.e., saturated, contaminated) benchmark now.

Comment by captainregex 2 days ago

my personal experience with swe has been…suboptimal. I am not sure how much I buy these benchmarks and it has a very “just blurt it out even if it’s probably not right” style but hey it’s free.

Comment by walrus01 2 days ago

Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system.

Comment by ricardobeat 2 days ago

"Not really news" that a model from a smaller lab can beat weeks-old Fable 5.1 at 70% lower cost? What a time to be living in.

Comment by p1esk 2 days ago

It’s very far from Fable on benchmarks that matter, like TB4.

Comment by walrus01 2 days ago

no, I mean not really news specifically on the number for terminal bench 2.

Comment by llm_nerd 2 days ago

They did reinforcement learning on Kimi K3, probably specifically targeting the benchmarks.

Eh, it isn't news. I mean, it's just an echo bit of news to the great K3 release.