Portal by Spotify cut my Claude Code token usage by 90%
Posted by cebert 4 days ago
Comments
Comment by schainks 3 days ago
Comment by larodi 3 days ago
not X, but Y... so this is very likely Opus 5 text, given timing and how it reads. has all the marks of it, one being very weird wording which makes it so hard to read the text.
Comment by ssivark 3 days ago
> modes are the load-bearing piece:
Okay, this write-up is filled with Claudisms.
Comment by larodi 2 days ago
- listing modes - this gate, other gate - canonical stuff etc.
would've been nice if Claude produces such output now and then. with 4 agents doing my stuff on a daily basis, I feel like vomiting at some point, not mere nausea, but disgust. damn Codex seems to fare better at this imho.
curiously I've tried many times to instruct it to not produce this nonsense, but _deus ex machina_ always finds a way around it. wonder if these "instructions" are actually treated as restrictions, rather than... mere obstacles. like - you can try to plumb a river, it always finds its way around plumbing.
Comment by andrew_mason1 3 days ago
Comment by schainks 3 days ago
Comment by larodi 2 days ago
Comment by hotelsacher 3 days ago
Comment by schainks 3 days ago
Comment by solenoid0937 4 days ago
I've never had an issue with Codex or Claude reading massive files, they're really good at precise greps.
Comment by jampa 3 days ago
Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).
I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".
Comment by johnnythujone 3 days ago
There’s an added benefit that the manager’s focus on strategy and task decomposition before actually handling the user’s prompted task directly seems to be a very good way to interact with Claude’s Fable safeguards, and I haven’t had any refusals doing this.
And while I haven’t ran any numbers, I can get orders of magnitude more out of my claude subscription doing this, especially with deepseek-v4-flash being as good as it is for as cheap as it is.
Comment by vintermann 2 days ago
Comment by johnnythujone 12 hours ago
Comment by jwillmer 3 days ago
Comment by astrange 3 days ago
This is a Claudism, right? I feel like I never saw "gated" used this way before it.
Comment by gopher_space 2 days ago
Comment by johnnythujone 12 hours ago
Just a guy trying to make his subscription last longer than the single Fable prompt anthropic includes for 100 bucks a month, lol.
Comment by bloomfieldj 3 days ago
The community edition was open sourced when the creator got hired by OpenAI a few months ago.
Comment by jimmySixDOF 3 days ago
Comment by hankbond 3 days ago
very good way to put it.
Comment by lxgr 3 days ago
That property would be very useful here, but I don't see how it would be achievable using LLMs.
Comment by ramraj07 3 days ago
Comment by Artimus 3 days ago
So fable and opus use opus to explore. Sonnet uses sonnet.
I replaced my built in explore agent with one hardcoded to sonnet low effort.
Comment by phoghed 3 days ago
Comment by shubhamjain 3 days ago
Why not, though? I started using OpenCode + GitHub Copilot, but I burned through my Claude Sonnet quota in just three days. I switched to GPT-5.4-mini, which uses far fewer tokens, and it’s often just as good as Sonnet. I think optimizing token usage is a good exercise. We often assume a model will be terrible, when it really isn’t.
Comment by solenoid0937 3 days ago
Comment by CuriouslyC 3 days ago
Comment by jurgenburgen 3 days ago
“Often” doesn’t sound great. If the smaller model fails then I just wasted a lot of time and tokens.
Comment by bensyverson 4 days ago
And why stop at 90%? I have this one weird trick to reduce Claude Code token use by 100%: use a different harness and model!
Comment by lxgr 3 days ago
Comment by jrm4 3 days ago
Comment by solenoid0937 3 days ago
Comment by jrm4 2 days ago
You kids don't get it, it's not about the tokens, it's about the principle of the thing. No self-respecting real programmer would accept the loss of even a few tokens over programmatic efficiency and cleverness.
Comment by fy20 3 days ago
The app would start using it for exploration tasks, and then as it improved it became the default for writing code and tests too. You can change it of course, but I find it does a pretty decent job if you have a large model directing it.
The parent model of course checks the work, but most of the time the handoff is good enough that no edits are needed.
It's also pretty fast and cheap, firing off a bunch of sub-agents to explore different parts of the codebase is a regular occurrence for the way I work.
Comment by zxspectrum1982 3 days ago
Comment by lamine-yaml 2 days ago
Comment by 14u2c 4 days ago
Comment by phreack 3 days ago
Comment by vintermann 2 days ago
Comment by ricardobeat 3 days ago
On top of that, saving 90% of input tokens != saving 90% "of tokens", output is wildly more expensive.
[1] especially if it's a really old model like Gemini 2.5!
Comment by Syntaf 2 days ago
I'm personally skeptical of optimizing for minimal token consumption, the closer to a vanilla setup I am the more confident I feel I'm always getting the best performance out of my models.
Just look at how JetBrains measured rtk and found that while an individual call saves tokens, agents on average perform *more* turns and use *more* tokens to accomplish one task[1]
[1] https://blog.jetbrains.com/ai/2026/07/rtk-claude-code-token-...
Comment by steveBK123 3 days ago
Comment by Jtarii 3 days ago
Comment by steveBK123 3 days ago
15 seconds.
Incredible stuff.
Comment by bspammer 2 days ago
The app was done in 2011. They could make a massive improvement across the board with a single git reset --hard command.
Comment by SebaSeba 3 days ago
And no, the problems you described are not there.
On my Windows 11 laptop it takes about 2sec to start and basically any song starts in less than a second.
Maybe it's Apple making Spotify crappy on Mac to promote iTunes?
Comment by bigyabai 3 days ago
I wonder if the macOS build has a bunch of static libraries it has to load first.
Comment by hotelsacher 3 days ago
Comment by shimman 3 days ago
Never been happier listening to music ATM.
Comment by subscribed 3 days ago
Related: https://bedrocknews.com/article/technology/2026/apr/11/ai-im... and https://www.headphonesty.com/2025/01/spotify-ghost-artists-c...
Just skip Spotify at all.
Comment by steveBK123 3 days ago
I’ve noticed sometimes my wife puts on a playlist and it’s almost like weird covers or remixes I can’t quite put my finger on. Like it has a melody of a real song but the style / tempo / vocals are changed.
Comment by Tistron 3 days ago
Comment by LiamPowell 3 days ago
Comment by steveBK123 3 days ago
Comment by hotelsacher 3 days ago
And I'm sure their prompt whisperers are proud of their work.
Comment by steveBK123 3 days ago
Comment by wging 3 days ago
More specifically, I'm talking about tapping the '...' on a track for the options menu. That should be latency-free, and used to be.
Comment by Tistron 3 days ago
Though I agree it should be.
Comment by hasteg 3 days ago
Comment by SaltyBackendGuy 3 days ago
Comment by hasteg 2 days ago
Comment by stavros 3 days ago
Comment by autotune 3 days ago
Comment by remus 3 days ago
Comment by samtheprogram 3 days ago
Comment by bitwizeshift 3 days ago
Comment by ssalka 3 days ago
Comment by stavros 3 days ago
Comment by asimovDev 2 days ago
as someone on the internet once said, it's good that no-one at Spotify realised that people would buy the subscription much faster if it was a crying baby soundbite playing every time instead of an advertisement
Comment by fuckinpuppers 3 days ago
Comment by bigtex 3 days ago
Comment by steveBK123 3 days ago
Comment by hotelsacher 3 days ago
Comment by surcap526 3 days ago
Comment by pmdr 3 days ago
Comment by landr0id 3 days ago
"But it adds motion" was the classic reply.
Comment by bluGill 3 days ago
There is nothing wrong with art. It is a great thing, I hope to see more art in the world. However, if the goal is sharing information, which is supposedly the goal of a fairly large number of websites, art needs to be secondary to sharing that information. And then those things that add motion, whatever: you're adding art and harming the real purpose.
Comment by PinkSheep 3 days ago
The comment above shows this type of people don't understand it. Yet they are the ones getting hired due to formal qualification. Those who do are at the intersection of design, engineering and computer science. The latter give you enough experience to understand the culture to recreate familiar _look and feel_.
Comment by ssl-3 3 days ago
Comment by dotancohen 3 days ago
I found no problem scrolling. Either Firefox on Android doesn't support whatever trickery they are doing, or they reverted it.
Comment by retsibsi 3 days ago
Comment by PinkSheep 3 days ago
Comment by amsterdorn 3 days ago
Comment by nosioptar 3 days ago
Ublock on Firefox mobile seems to keep scrolling unfucked on this page for me. (Designer/"dev" of this page still sucks wet hobo socks.)
Comment by _rwo 2 days ago
Comment by jnwatson 4 days ago
You can also just delegate this to subagents with Claude Code (though you have a more limited choice of models unless you swap the cheaper models via OpenRouter).
I'm OK using a dumb model as a smart grep, but the whole point of using the frontier models is using their intelligence for the hard stuff like coding.
Comment by CaveTech 3 days ago
Comment by spockz 3 days ago
Basically I run in luna high or extra high continuously with a terra subworker dedicated to planning and difficult research questions. Then I end with a final review in Terra or Sol depending how big the feature is.
Comment by donavanm 3 days ago
- Write ~1 paragraph of developer instructions (AGENTS.md): Use subagents for tasks that can be decomposed, worked on in parallel, or delegated. Describe common examples. I put a reference to a "how to use subagents" skill for more details. The "skill" isnt' always read (as subagents arent always useful) which saves some tokens. But you pay the once-per-session read-skill cost when its relevant.
- Describe how to use subagents in ~1 page or less (SKILLS.md): use them for sub tasks. select model size/quality based on task ambiguity, scope, unbounded work, or conflicting requirements. Use reasoning effort for complexity, interdependence, or ambiguous success criteria. How to evaluate complexity & common subtask examples across the spectrum. give tasks a relevant name like "model-family_version_reasoning-effort_task-description" so you can actually understand what theyre doing by name.
- in dev instructions (SKILLS.md) provide a table of agent names (low complexity summarizer, bounded implementation, complex implementation), model+effort (gpt-5.6-luna medium, gpt-5.6-luna high, gpt-5.6-sol medium), and short description of 2-3 task "types" for each.
- Explain they can use "default" or specify their own custom model settings if needed.
- Define your list of subagent profiles in ~/.codex/agents/ which matches names (low_complexity_summarizer.toml) from previous. In each you'll need to set model, reasoning, and `developer_instructions` that describe *how* to do a task, *not what* to do.
Details to know: - IMO subgent profiles are "task centric" because `developer_instructions` are required. You can't just specify model & reasoning, you also have to give valid developer_instructions that will be merged in to every session/prompt. I address this by defining a few different agents for tasks that are commonly encounted like summarization, synthesis, planning, implementation, etc. The different agent profiles (~2-5 per category) will "scale" the model + reasoning based on the complexity and ambiguity. This work pretty well in practice. And you don't need to over due it, the harness/agent can still launch a "custom" profile that uses the parent sessions developer instructions.
- You need to use agent profiles with codex because "v2" models (terra & sol) can't launch "v1" models (luna). There are a couple of code paths to avoid this, the agent profile is the simplest.
Anyways, write you skill & subagent profiles and it basically "just works".Comment by alastairr 3 days ago
Comment by shikck200 3 days ago
Comment by faangguyindia 3 days ago
Try it yourself, use a big model like Opus or Sol to implement everything by first making a plan using plan mode.
Then try distributing the task to a cheaper models like Luna Max or Gemini Flash 3.8.
During planning, the big model already reads the relevant files in context, while giving a smaller model a slice of work itself requires the big model to reason about the task distribution, review, etc.
So do you really save on tokens?
Comment by majormajor 3 days ago
Comment by klodolph 3 days ago
When I do this, I can have it use cheap subagents with models like Luna to read the relevant files.
Comment by MPSimmons 3 days ago
Comment by skybrian 3 days ago
Comment by sognetic 3 days ago
Comment by Banditoz 3 days ago
Comment by orliesaurus 3 days ago
Comment by tobinfekkes 3 days ago
Comment by jokethrowaway 3 days ago
If you think a cheap model is smart enough to filter information to give to your expensive model, you can save some money. If you think your cheap model is smart enough to format your expensive output, you can save some money.
In practice, this didn't work well until Qwen 3.8.
Qwen 3.6 and (abliterated) Gemma 4 were almost there but still making mistakes.
Comment by khaki54 3 days ago
Comment by gruez 4 days ago
>Tested against a Java monorepo across four scenarios, measuring tokens Claude would consume reading files directly vs. consuming the bulk-reader's summary or writing code via the code-writer. Mean bulk-read savings were around a whopping 90%.
>The code-write scenario is harder to measure in tokens because without shunt, Claude both reads the reference files and generates the output as expensive output tokens. With shunt, the code goes straight to disk, Claude never sees it.
So nothing about accuracy or actual performance? At least run against DeepSWE bench or something.
Comment by gilmtz 3 days ago
So the actual performance was bad.
It might be an acceptable trade off tho. If token costs become prohibitive, then using a meat engineer to actually debug could be cheaper.
Comment by nerdyadventurer 1 day ago
Comment by tolugenius 4 days ago
Comment by CharlesW 3 days ago
Comment by spockz 3 days ago
In codex I don’t see this behaviour despite having added the instructions to do so to my agents file. I also let that agents file be reviewed by Sol to come up with the right phrasing but no luck so far.
Comment by florians 3 days ago
Comment by CharlesW 3 days ago
Comment by Dmitry-kov 2 days ago
Comment by ellessarr 3 days ago
Comment by dominicl 3 days ago
(fully vibe coded)
Comment by andai 3 days ago
> If Claude needs to make edits based on the analysis, it still has to read the specific section directly.
Yeah iirc the sysprompt tells it that it must always (re-)read a file before editing it. I noticed this because I customized Claude Code to just read all source at startup (if the project was only a few thousand lines of code). And it would still read the stuff it had already read! Because the system prompt explicitly told it to...
Comment by amelius 3 days ago
Comment by ryuuseijin 3 days ago
There is a tool that uses ripgrep and treesitter that does this [1], adapted from the maki coding agent.
Comment by avazhi 3 days ago
We’re fucked.
Comment by kristianp 3 days ago
Comment by stephbook 3 days ago
Next sentence was also an AI juxtaposition. Done.
Comment by lowbloodsugar 3 days ago
Comment by pmontra 3 days ago
> The modes are the load-bearing piece:
Why do people write like LLMs? Maybe they delegate all the work to a LLM and don't have the time or the will to edit the copy. How about telling another LLMs to replace at least the most common LLM patterns with something human looking?
Comment by eterm 3 days ago
I'm fairly confident this is just LLM writing the majority, possibly tweaked by a human.
Opening line is a form of, "It's not X, it's Y": ".. isn't thinking. It's I/O".
Then the start of the second paragraph is that weird breathless kind of writing:
> Reading five files to answer a question about one method. Generating a test file that follows the exact same pattern as the twenty test files next to it.
More "It's not X, it's Y": The seat license isn't what hurts, it's the tokens.
The softly pressed insistence that AI is worth it, really: "The tooling pays for itself but only if..."
Comment by blehn 3 days ago
Comment by vagabund 3 days ago
Comment by florians 3 days ago
Comment by prmoustache 3 days ago
Comment by florians 3 days ago
Comment by fbn79 3 days ago
Comment by wejick 3 days ago
Eg. On opencode there's explorer subgagent that we can set to use lower level model. Many people even use haiku level model for this.
Comment by shevy-java 3 days ago
Wasn't the AI promise to save costs? So it was a lie.
Comment by bluGill 3 days ago
Measuring productivity is, of course, a hard problem and I'm going to leave this completely out of this comment.
Comment by FelineStateMach 3 days ago
Comment by bakugo 3 days ago
Comment by dotancohen 3 days ago
Comment by Zambyte 3 days ago
Comment by hexo 3 days ago
Comment by cute_boi 3 days ago
And, I can't believe this is from official spotify.... What a joke.
Comment by throooooo 3 days ago
Comment by pmontra 3 days ago
Comment by 1saadcodes 3 days ago
Comment by simianwords 3 days ago
Comment by sandos 3 days ago
Comment by doubleorseven 3 days ago
fable already calls subagents (hoping not opus), so i don't really see the gain here
Comment by thewhitetulip 3 days ago
Comment by guluarte 3 days ago
Comment by tetrisgm 4 days ago
Comment by kotaKat 3 days ago
Right?
Rules for thee...
Comment by m3kw9 3 days ago
Comment by claiir 3 days ago
> The modes are the load-bearing piece:
lol
Comment by m3kw9 3 days ago
Comment by ig0r0 3 days ago
Comment by blopusai 3 days ago
Comment by andrethegiant 3 days ago
Comment by krzys 3 days ago
And it's just the execution, I'm not even going to comment on the idea itself, as it's even worse.
How this slop-post landed on the front page...?
Comment by nektro 2 days ago
Comment by nyxtom 2 days ago
Comment by addedlovely 3 days ago
Comment by 3371 3 days ago
Comment by dominotw 3 days ago
imagine working at these shitty companies with their shitty wrapper tooling.
Comment by acceptallmag 1 day ago
Comment by omid-io 2 days ago
Comment by scraplabs 2 days ago
Comment by fif7y 4 days ago
Comment by bizcalclab 3 days ago
Comment by Buoylog 4 days ago
Comment by manganate06 3 days ago
Comment by BottieZimmie 4 days ago
Comment by shledery 3 days ago
Comment by nirmeet011011 3 days ago
Comment by jing09928 3 days ago
Comment by malinono 3 days ago
Comment by jens_tlb 3 days ago
Comment by ahmedelsama 3 days ago
Comment by lxgr 3 days ago
One is strictly a performance optimization, the other is a speed/quality tradeoff. It might well be a very good one, but it’s a tradeoff nevertheless. The framing is misleading.