Show HN: Claude-thermos keeps your Claude session warm for you
Posted by s0ck_r4w 1 day ago
Comments
Comment by SwellJoe 1 day ago
How Claude handles its sessions is none of my business. I'm going to let them do the best they can to provide good service for everyone, and if they can't/won't, I'll switch to a provider that can.
Using these massive models is already pretty danged extravagant, I'm not going to demand to be at the front of the queue at all times, too.
Comment by brookst 21 hours ago
Cached input tokens cost 10% of uncached. So if you’re model runs for 45 minutes, generates 300k output tokens and asks you a question, it costs 10x more if you wait 5.01 minutes to answer.
Sure, you may be willing to pay 10x more (or get 10x less for your subscription). But the time limit is arbitrary and has nothing to do with other peoples’ workloads. So I think your point is a non sequitur.
Comment by Jensson 20 hours ago
No it has to do with others workloads, now you keep their cache for longer so others will get less. And no its not arbitrary, they run out of memory, if more people do this they will have the dial it down further or run out of capacity.
Comment by brookst 11 hours ago
Are you imagining this a fixed MRU where duration scales with usage? Becasue that is not at all what Anthropic documents: https://platform.claude.com/docs/en/build-with-claude/prompt...
You would not get more than 5 minutes if you were the only user in the world. You would not get less at their peak hours.
Comment by toasty228 18 hours ago
Because ram/memory is free and not in demande at all these days?
Comment by searealist 19 hours ago
2) Now imagine Anthropic or OpenAI now charge your per minute of reserved VRAM time. It would be more fair if they did. Would you still want to run a tool like this?
Comment by brookst 11 hours ago
Comment by himata4113 22 hours ago
Comment by dannyw 23 hours ago
Comment by gruez 22 hours ago
Comment by spacemanspiff01 21 hours ago
Comment by GMoromisato 1 day ago
In this specific case, Anthropic can avoid keeping the cache if it detects this kind of prompt (i.e., if max tokens < some number).
Comment by s0ck_r4w 1 day ago
Comment by SwellJoe 1 day ago
Comment by dannyw 23 hours ago
If you keep this running for hours without doing anything, it will drain your limits and API. The use case of keeping the main thread cache warm while subagents work is very genuine and legitimate.
Comment by s0ck_r4w 1 day ago
How is this comparable to going to lunch or taking a walk?
Comment by SwellJoe 1 day ago
Comment by devnonymous 1 day ago
Comment by idonotknowwhy 1 day ago
Comment by dannyw 22 hours ago
Comment by kakugawa 1 day ago
Comment by dannyw 22 hours ago
The cache is discounted for a reason. They WANT you to use it.
Comment by searealist 19 hours ago
1) Start charging for VRAM reservations.
2) Charge _other_ customers more.
3) Eat the cost themselves.
Comment by dannyw 18 hours ago
Anthropic (and now OpenAI too for 5.6) prompt caching is not free.
Comment by brookst 21 hours ago
I don’t think your understanding works.
Comment by fearmerchant 1 day ago
Comment by dannyw 22 hours ago
Keeping your cache warm is a good thing, caching saves compute and electricity.
Cached input is cheap for a reason, it is in everyone’s mutual interests to maximise cache hit rates.
Comment by OccamsMirror 22 hours ago
Comment by brookst 21 hours ago
Comment by sznio 1 day ago
Comment by pkulak 1 day ago
Comment by engineer_22 23 hours ago
Comment by devnonymous 1 day ago
> Detect the danger window. When the main lineage goes idle and a subagent is actively running, the main prefix is at risk of expiring.
So, this is not demanding to be at the front of the queue, it's just paying someone to take the place you already had in the queue, when you want to take a leak.
Comment by SwellJoe 1 day ago
Comment by devnonymous 1 day ago
Comment by unholiness 1 day ago
If you're paying API rates, you can choose 5m or 1hr yourself (and pay different rates).
Keeping a 1hr cache warm could still be useful, sure, but outside that, I don't see much use of this today.
Comment by s0ck_r4w 1 day ago
Comment by Wowfunhappy 23 hours ago
https://code.claude.com/docs/en/prompt-caching#on-a-claude-s...
(Thank you to EliasWatson for giving me this link just a few days ago, as I was previously confused too.)
Comment by alukin 1 day ago
Comment by gogobio 1 day ago
Comment by supern0va 1 day ago
Comment by Lalabadie 1 day ago
Comment by supern0va 1 day ago
Comment by l1n 1 day ago
Comment by agluszak 1 day ago
Comment by munk-a 1 day ago
Fixed that for you.
Comment by cadamsdotcom 1 day ago
Cache duration is arbitrary. What it actually does (if used en masse) is decrease the amount of oversubscription their infra can handle..
Comment by janderson215 1 day ago
Lately, I’ve been thinking about how this related to fractional banking. If you were to eliminate fractional banking introduced in the US by Hamilton, you would destroy a lot of current prosperity.
Comment by sillysaurusx 22 hours ago
Comment by Cyberdog 1 day ago
Comment by boc 1 day ago
You would have avoided that cache hit if the LLM session was kept "alive" for those few hours. Why not automate the part where you keep the large main thread alive until you're ready to analyze the results?
Comment by PcChip 1 day ago
Comment by leemoore 23 hours ago
Comment by purpleidea 1 day ago
Hearing one byte refreshes the whole thing is huge! 5min is wayy too slow, because sometimes I want to spent more than 5 min looking at a diff before choosing where to go next.
Kind of outrageous, I hope this kind of feature gets built into claude code =D
Comment by jonas21 1 day ago
Comment by s0ck_r4w 1 day ago
*UPD:* actually it appears the default is authentication-dependent. API key gets 5 minutes, subscriptions - 1 hour.
Comment by foota 1 day ago
Imo it's their fault for not having pricing that aligns incentives.
Comment by cosmotic 1 day ago
Comment by davesque 1 day ago
Comment by Wowfunhappy 23 hours ago
Comment by jaimehrubiks 1 day ago
Comment by gabigrin 1 day ago
Comment by skeledrew 23 hours ago
Comment by ATMLOTTOBEER 1 day ago
Comment by SwellJoe 1 day ago
Comment by skeledrew 23 hours ago
Comment by SwellJoe 18 hours ago
Comment by cortesoft 1 day ago
Comment by s0ck_r4w 1 day ago
Comment by addaon 1 day ago
Comment by randomblock1 1 day ago
Comment by 2001zhaozhao 1 day ago
Comment by leemoore 23 hours ago
Comment by razodactyl 1 day ago
Comment by broodbucket 1 day ago
Comment by leemoore 23 hours ago
Comment by broodbucket 23 hours ago
Comment by cortesoft 1 day ago
Comment by edot 1 day ago
Comment by sublinear 18 hours ago
Comment by j45 1 day ago
For example, there might be something I intended to complete in one sitting, but took two sittings in the same day unexpectedly. Maybe it could just be a few cache delays per day or something, tagged in advance somehow.
Comment by smokeeaasd 10 hours ago
Comment by IrfanD 15 hours ago
Comment by szin 1 day ago