GPT-6 Astra on OpenRouter
Posted by Topfi 4 days ago
Comments
Comment by simonw 4 days ago
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Comment by dector 3 days ago
Comment by theturtletalks 3 days ago
Comment by theturtletalks 3 days ago
Comment by SkiFire13 3 days ago
Comment by theturtletalks 2 days ago
Comment by alexgoodhart 3 days ago
Comment by anthonyrstevens 10 hours ago
Comment by kyorochan 3 days ago
I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.
Comment by stymaar 3 days ago
Comment by maleldil 3 days ago
Comment by stymaar 3 days ago
Comment by brookst 3 days ago
Comment by WarmWash 3 days ago
Comment by stymaar 3 days ago
Comment by maleldil 2 days ago
Comment by threatripper 3 days ago
Comment by epihelix 3 days ago
Comment by jayd16 3 days ago
Comment by steve-atx-7600 4 days ago
Comment by simonw 3 days ago
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
Comment by pilaf 3 days ago
Comment by stymaar 3 days ago
Comment by Petersipoi 3 days ago
Did even better from the front. What's surprising is that it used the exact same colors as the simonw example, despite my prompt only being
> Generate an SVG of a ring-tailed lemur riding an electric scooter. Front view
GPT-6 Astra Extra High
Comment by Fuzzwah 3 days ago
Comment by benatkin 3 days ago
Comment by CamperBob2 3 days ago
Comment by ulrikrasmussen 3 days ago
Comment by swingboy 3 days ago
Comment by y1n0 4 days ago
Comment by Kranar 4 days ago
Comment by mudkipdev 4 days ago
Comment by awakeasleep 4 days ago
Comment by WarmWash 3 days ago
If that was the case, the models would have been producing near perfect outputs for it a year ago.
Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".
Comment by dgellow 3 days ago
Yes
Comment by jenniferhooley 3 days ago
I'd be shocked if they didn't myself.
Comment by Kranar 2 days ago
You really have to go out of your way to completely misrepresent what's being claimed here in order to make such a wildly off-topic reply.
Comment by kibae 3 days ago
Comment by stymaar 3 days ago
Comment by mi_lk 4 days ago
Treat it like a bit as is
Comment by benatkin 3 days ago
Comment by epiccoleman 3 days ago
Comment by benatkin 2 days ago
Comment by justinbaker84 3 days ago
It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.
Comment by jsdalton 4 days ago
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
Comment by simonw 4 days ago
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
Comment by bahmboo 3 days ago
Comment by WarmWash 3 days ago
Comment by bahmboo 3 days ago
Comment by Gander5739 3 days ago
Comment by bahmboo 3 days ago
Comment by spockz 3 days ago
Comment by Gander5739 3 days ago
Comment by xfax 3 days ago
I used Light mode and it used up all my limits for the day and had to continue the following day.
Comment by CamperBob2 3 days ago
Missing GLM-5.3, though.
Comment by carljungslabtek 2 days ago
Comment by jostylr 3 days ago
Loving the pelican silliness. My ChatGPT is over the moon about it. Gonna miss it when it really is done. Though maybe a pelican riding a bike game could become the benchmark in a year.
Comment by vb-8448 4 days ago
Comment by thimabi 4 days ago
Comment by pizza234 4 days ago
Comment by vessenes 4 days ago
Comment by TiredOfLife 3 days ago
Comment by petilon 4 days ago
Comment by andai 4 days ago
Reminded me of https://clocks.brianmoore.com/
Comment by CamperBob2 3 days ago
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
Comment by BrokenCogs 4 days ago
Comment by jeffybefffy519 3 days ago
Comment by mkagenius 3 days ago
Coz who knows if astra low will produce max like output if tried once more.
Comment by alastairr 3 days ago
Comment by mkl 3 days ago
Comment by simonw 3 days ago
I heard a rumor that Luna is a slightly different architecture from Sol and Terra, which makes me wonder if Luna and Astra might be more related to each other than to Sol and Terra.
Bit of a big leap to make from a token count though!
I just checked the tiktoken library and couldn't see any changes relating to Luna: https://github.com/openai/tiktoken
Comment by az226 2 days ago
Comment by Xunjin 4 days ago
I'm wondering if this is being trained on by the models today.
Comment by r_lee 4 days ago
Comment by scotty79 3 days ago
Comment by PacificSpecific 3 days ago
Comment by yreg 3 days ago
Apparently the crowd agrees because they keep upvoting these.
Comment by dolebirchwood 3 days ago
Comment by lspears 4 days ago
Comment by droidjj 4 days ago
Comment by simonw 4 days ago
Comment by simonw 4 days ago
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
Luna — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 7.83 | 1.57 |
| xhigh | 4.24 | 0.85 |
| high | 2.46 | 0.49 |
| medium | 1.26 | 0.25 |
| low | 0.76 | 0.15 |
| none | 0.71 | 0.14 |
+--------+--------+---------+
Per million tokens:
Before: $1 input / $6 output
After: $0.20 input / $1.20 output
Sol — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 48.55 | 32.37 |
| xhigh | 24.11 | 16.08 |
| high | 10.38 | 6.92 |
| medium | 10.55 | 7.03 |
| low | 8.33 | 5.55 |
| none | 5.90 | 3.93 |
+--------+--------+---------+
Per million tokens:
Before: $5 input / $30 output
After: $4 input / $20 output
Terra — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 32.09 | 25.67 |
| xhigh | 14.67 | 11.74 |
| high | 3.74 | 2.99 |
| medium | 3.46 | 2.77 |
| low | 3.47 | 2.78 |
| none | 2.60 | 2.08 |
+--------+--------+---------+
Per million tokens:
Before: $2.50 input / $15 output
After: $2 input / $12 outputComment by fHr 4 days ago
Comment by EugeneOZ 3 days ago
Comment by kart23 4 days ago
Comment by gpugreg 3 days ago
Comment by poopiokaka 3 days ago
Comment by man4 4 days ago
Comment by jedjjfjf 4 days ago
Comment by leoqa 4 days ago
Comment by samuelknight 4 days ago
Comment by satvikpendem 4 days ago
Comment by ComplexSystems 4 days ago
Comment by StopTheCringe 4 days ago
Comment by bnorton 4 days ago
Comment by maxlapdev 4 days ago
Comment by sumedh 4 days ago
Isnt Netherlands the leader in bike riders and they dont wear helmets.
Comment by neutronicus 4 days ago
Comment by jjcm 3 days ago
Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...
And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
Comment by copperx 3 days ago
I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.
Comment by jermaustin1 3 days ago
Everything costs more now than it did a year ago... except for THIS, and we are still complaining that a 90-99% reduction in cost is STILL too expensive. And a 50-75% reduction in time is STILL too long.
We used to have to wait for weeks for a design like that when I worked at a consultancy, and that is a week of salary. For the design, then it got handed off to a front end developer to slice it and get built so the back end developer can hook it up to a CRM. We are talking a month turn around with design, revisions, development, testing, and bug fixing.
It can now be done in a couple of hours for less than a single hour's cost. If it were 10x slower and 10x more expensive, it would STILL be "good deal".
Comment by kolinko 3 days ago
Comment by paxys 3 days ago
Comment by JumpCrisscross 3 days ago
Compared to what?
Comment by copperx 3 days ago
Comment by risyachka 3 days ago
Yeah if you are solo developer without budget.
For any business this is nothing, the ROI is massive.
Comment by mydreamof 3 days ago
Comment by CapsAdmin 3 days ago
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
Comment by jjcm 2 days ago
Another way to think about it is you could prompt Astra to put the building back in. You couldn't prompt Opus 5 to get the correct curvature of the line / cutout. That part has always been a huge struggle for models.
Comment by frenchtoast8 3 days ago
I would recommend staying away from OpenRouter. No matter how good the service is, if anything does go wrong, you have no recourse and you lose every credit in your account. Ironically some of the few responses I actually saw in the Discord were doubling down on their “no refunds no matter what” policy.
Comment by Gecko4072 3 days ago
Comment by bellowsgulch 3 days ago
Comment by miyuru 3 days ago
Comment by enraged_camel 3 days ago
Comment by tocariimaa 3 days ago
Comment by weberer 3 days ago
Comment by predkambrij 3 days ago
Comment by frenchtoast8 3 days ago
Comment by dominick-cc 3 days ago
Comment by bellowsgulch 3 days ago
Comment by XCSme 4 days ago
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
Comment by satvikpendem 4 days ago
Comment by dprkh 4 days ago
Comment by mceachen 3 days ago
Comment by embedding-shape 3 days ago
Comment by dprkh 3 days ago
Comment by kulahan 4 days ago
Comment by appplication 4 days ago
Comment by redox99 3 days ago
Comment by XCSme 3 days ago
Comment by embedding-shape 4 days ago
Comment by XCSme 4 days ago
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
Comment by kulahan 4 days ago
Comment by XCSme 4 days ago
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
Comment by kulahan 3 days ago
Comment by XCSme 4 days ago
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
Comment by kzrdude 3 days ago
Comment by sehugg 3 days ago
Comment by embedding-shape 3 days ago
Comment by embedding-shape 3 days ago
Comment by XCSme 4 days ago
Comment by embedding-shape 3 days ago
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
Comment by XCSme 3 days ago
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
Comment by readams 4 days ago
Comment by PUSH_AX 2 days ago
Comment by viccis 3 days ago
>Playing ping pong from the side of the table
I think marketing might be getting a bit absurd at this point
Comment by mawadev 3 days ago
I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"
Comment by jubilanti 3 days ago
The entire mainstream media and political establishment, and every normie I meet is convinced Skynet is already upon us
Comment by ryanschaefer 3 days ago
Comment by MisterMunchkin 3 days ago
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
Comment by kolinko 3 days ago
Comment by wilkystyle 3 days ago
Comment by kolinko 2 days ago
Comment by bigyabai 3 days ago
Comment by sschueller 3 days ago
Comment by kolinko 3 days ago
It must be super interesting working there rn :D
Comment by wookmaster 3 days ago
Comment by gentlewater 3 days ago
Comment by ghosty141 3 days ago
Comment by yurishimo 3 days ago
Comment by ghosty141 3 days ago
Comment by KptMarchewa 3 days ago
Comment by simianwords 3 days ago
Comment by kingstnap 4 days ago
Comment by InsideOutSanta 4 days ago
Comment by embedding-shape 4 days ago
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
Comment by wincy 3 days ago
Comment by wincy 4 days ago
Edit: nevermind it JUST gave me a notification to use it!
Comment by paxys 4 days ago
Comment by cromka 1 day ago
Comment by saidnooneever 3 days ago
almost had the feelin it was watching its little brother fail and had to 'step in' for a moment :').
time to go play outside...
Comment by sumedh 4 days ago
Comment by killerstorm 3 days ago
Comment by cmrdporcupine 3 days ago
It's terse, like all GPT models by default, but the sentences feel less obscurantist.
It's also more pro-active about problem solving.
Comment by theagenticleade 3 days ago
Where do we go from here??
Comment by sejje 3 days ago
Comment by kzrdude 3 days ago
Comment by kdnvk 3 days ago
Comment by 1saadcodes 3 days ago
Comment by vb-8448 4 days ago
Comment by WASDx 3 days ago
Comment by algoth1 4 days ago
Comment by friendlypenguin 3 days ago
Comment by redox99 3 days ago
Comment by simianwords 3 days ago
Edit:
GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens
GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens
So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive.
Why is the burden of proof on me tho!?
Comment by slopinthebag 3 days ago
astra high is also 3x cheaper than opus max at basically the same intelligence.
astra high is also about as expensive as sol max while being more intelligent.
astra medium is cheaper than sol max while also being cheaper and roughly same intelligence.
im going to replace my sol usage with astra high/medium i think
caveat: benchmarks are really fuzzy with llms
Comment by haaz 3 days ago
Comment by wickedsight 3 days ago
Comment by dgellow 3 days ago
Comment by ygty 2 days ago
Comment by d2p 3 days ago
Comment by forrestthewoods 3 days ago
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
Sure seems like Astra is expensive AF.
Comment by ebiester 3 days ago
Comment by starik36 4 days ago
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?
Comment by hhh 4 days ago
Comment by claiir 4 days ago
Comment by nibbleyou 3 days ago
Comment by itsjustkev 4 days ago
Comment by Peanuts99 3 days ago
Comment by r_lee 4 days ago
Comment by jiggawatts 3 days ago
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
Comment by r_lee 3 days ago
but again, seems like there's no word from Azure if this applies to them.
very confusing.
Comment by jaesonaras 4 days ago
Comment by NSUserDefaults 3 days ago
Comment by gavinray 4 days ago
I'm a Business plan user with Cyber verification enabled, FWIW.
Comment by embedding-shape 4 days ago
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
Comment by thenthenthen 2 days ago
Comment by upcoming-sesame 3 days ago
Comment by logged4upvoting 3 days ago
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
Comment by DrProtic 3 days ago
Comment by swe_dima 3 days ago
Comment by christophilus 3 days ago
Comment by ellessarr 3 days ago
Comment by cute_boi 4 days ago
Comment by azinman2 3 days ago
Comment by vatsachak 4 days ago
Comment by hermesrouter 2 days ago
Comment by gertlabs 3 days ago
Comment by marsven_422 4 days ago
Comment by StopTheCringe 4 days ago
Comment by Osama0456 4 days ago
Comment by r_lee 4 days ago
Comment by OpenGayEye_ 4 days ago
Comment by pointitkememe 3 days ago
Comment by boxed 3 days ago
Comment by noob0053 4 days ago
Comment by taywrobel 4 days ago
Comment by thejassbrain 3 days ago
Comment by 6thbit 4 days ago
Comment by drivers99 4 days ago