The new rules of context engineering for Claude 5 generation models
Posted by mellosouls 2 days ago
Comments
Comment by mycentstoo 1 day ago
Comment by hmokiguess 1 day ago
Comment by Retr0id 1 day ago
Comment by chmod775 1 day ago
It was rewarded for this during training for some reason.
Alternative theory:
The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans.
But humans generally prefer to communicate on low levels of abstraction, through a back and forth, until the hard-to-express higher abstraction exists in the head of everyone involved - without ever being directly communicated. This is because we don't think using language. Language is merely a lossy translation of our thought into something expressible, happening after the fact or alongside it.
So when the LLM starts speaking to us using patterns and terms it created for itself during training to encode abstract thought in language, communicating with it becomes painful.
Comment by chrisweekly 1 day ago
Your assertion that we don't think in language is questionable. It runs counter to the lived experience of developing thoughts through writing ("writing isn't capturing thinking -- it is thinking"). I believe there is more to thought than language alone, but I also feel quite sure that language forms an essential part of thinking beyond a base layer of instinctive animalistic associations. Sophisticated thoughts are impossible to construct or maintain in the absence of language to represent concepts.
Edt to add: I cited Rilke because I find the notion [some deep thoughts are beyond language] interesting. But I disagree with the idea that language is only ever epiphenomenal (co-occurring with thought), or akin to a hard-of-hearing scribe attempting to convey thoughts which always have independent existence.
Comment by noduerme 1 day ago
This goes back to the whole Tarzan obsession of the early 20th, I guess, or earlier. But we know that apes can make simple tools without the language to describe them, or the thought process that went into them.
Thought is multimodal. Language is just one lossy mode.
Comment by TeMPOraL 1 day ago
So can a cruise missile. Also I think there's like separate part of the brain for that
> using concepts of phenomena like gravity, distance, speed, the threat level of another animal, etc. without needing a linguistic expression of those things.*
FWIW, AFAIK we haven't shown the ability to think in concept exists anywhere except in humans (because philosophy, reported experience) and LLMs (because we can literally see them forming and activating in patterns, and we've learned to identify them specifically, and experimentally verified through amplifying or suppressing them and observing behavior, etc.).
But more importantly:
> Images and music can convey ideas without language. Math can convey ideas without language. Physically taking apart an object and putting it back together can convey extremely complex ideas without language.
Images and music and math are langauge. If it can convey ideas, it is language.
Words and sentences and speech are subset of the idea of language and communication, that for some reason gets routinely confused for the whole thing. At this point I'd say even the "language models" are badly named, simply because people see "language models" think of "token" as number representing a sub-word element in existing human language like English. With multimodal models, at this point tokens are closer to units of sensory experience.
Comment by lanstin 1 day ago
And I rather disagree that mathematics is language, when I did maths there was a distinct difference from understanding a thing and then writing it down.
And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.
I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen. In the meantime, in my hands, they can do very nice dataviz programming to make complex systems much more transparent, even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.
Comment by ben_w 1 day ago
It is generally the case that there's a difference between "understanding a thing" and "writing it down", as demonstrated by every student who studies for the exam, the Chinese Room thought experiment, and Business Bullshit Bingo.
For example, I can copy the next sentence of yours, but I have no idea what:
> And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.
means, can you rephrase that by as much as possible? Preferably without a double negative?
Agree regarding sanity issues of AI. Extremely unlikely we make something "stable" so early in our attempts.
Comment by lanstin 9 hours ago
I was arguing that in:
> With multimodal models, at this point tokens are closer to units of sensory experience.
The tokens are inherently missing some extremely vital pieces of sensory experience, namely the ones that establish the narrative of a self.
Comment by ben_w 8 hours ago
> The tokens are inherently missing some extremely vital pieces of sensory experience, namely the ones that establish the narrative of a self.
Hmm.
While I agree LLMs are missing many pieces of sensory experience, it is unclear to me which pieces are necessary and sufficient for a narrative of self.
Clinical dissociation (and, at least to the approximation of public stereotypes, Buddhism) come to mind as an example where the presence of usual sensory input is insufficient for a sense of self: https://www.mayoclinic.org/diseases-conditions/dissociative-...
More broadly: while I agree that LLMs are not like us*, I am unclear why this matters in this context?
The behaviour of tokens in a transformer seem to me to behave like sensory input, just not human-like sensory input. The closest human analogy would be if the entire context window was a retina, each rod and cone one of the tokens in that context, and the output was filling in the blind spot (but of course, even this is a very loose analogy).
* even if the engineering teams were trying to do that, which they are not, it would be unlikely to converge on us so soon
Comment by TeMPOraL 23 hours ago
Wait, did I miss the Chinese Room suddenly becoming sensible, or relevant? Last I checked it was a failure to accept that the system of "human + room" can in fact understand Chinese, even if the components individually don't.
Comment by ben_w 12 hours ago
"Is the Chinese Room intelligent/conscious?" is a scissor question*, i.e. people strongly disagree about what it shows.
Myself, I'm with you: the system collectively is intelligent, despite the human in the loop being reduced to a cog in that system. The biological analogy to "the human does not understand Chinese therefore the system is not conscious" would be "no individual cell in a human body knows what the human is doing therefore humans are not conscious".
Comment by TeMPOraL 23 hours ago
Yes, but I believe this is happening with LLMs too. Language (as in text, symbols, actions) is a vehicle, but much like us, language models have internal models and represent concepts (this has been directly, empirically demonstrated few years ago), and they don't "think" in tokens either[0].
> And modalities that don’t include “where is this body in space” or “I am lonely” don’t seem to capture some essential elements of sensory experience. It is a mistake to separate the interior from the exterior in the analysis of how our thinking works.
That's fair. LLMs don't capture every dimension we experience. There's also history - our individual lived experiences since birth are, in my view, something else entirely. It's not a modality, but it's also not something currently possible to capture in or post training.
> I think the folks saying these systems have a kind of intelligence, but a non-human kind are correct - whether that can lead to an independent intelligence that can stay stable sane and focused for weeks or months as many people can, without some sort of embodied cognition providing over all wellness checks to keep the system of thinking sane, remains to be seen.
Possibly. I definitely agree it's not human intelligence. I think it's human-like, in the sense of human-approximating, by virtue of how it's trained[1], but it's arriving there via a different path so end result can still be alien (though I speculate approximation will hold[2]).
> even as just throwing all the ops data at them and asking what’s up is, several years in, still not worth doing.
I guess depends on the complexity of the case (and my understanding of your example); e.g. in my case, I had stellar results from giving Sonnet and Opus (from 4.6 all the way to now) access to my Home Assistant instance. They aren't perfect at making dashboards, but they're excellent at surfacing insights I didn't even realize were possible to get.
--
[0] - That's distinct from the "tokens are units of thinking" heuristic, which still holds for mechanistic reasons - best analogy IMO is "clock signal" in ICs.
[1] - The overall goal function is literally just "generate output that looks sensible to a human", in fully general, unqualified sense. Or, put another way, we're just brute-forcing DWIM, and rating the output by whether it's "what I meant".
[2] - Thinking about constraints on biological evolution, whatever the design of a human mind is, the fundamentals behind it must be so simple, that a greedy incremental optimizer random-walked into it. Given how far we've got with LLMs using simple architecture and crude training methods, and especially how eerily similar their failure modes are to human cognitive failure modes, I suspect LLMs are actually attracted towards the same fundamental design as evolution discovered.
Comment by lanstin 9 hours ago
The time scales and mechanisms of investor driven development don’t seem to share much with evolutionary processes. And digitally there’s not really a native equivalent to the global and dynamic state that the chemical melee inside and between cells provides.
Comment by TeMPOraL 6 hours ago
Comment by porridgeraisin 1 day ago
Comment by TeMPOraL 23 hours ago
Comment by JohnBooty 1 day ago
It's why I think "LLMs are only fancy autocorrect" style takes are really underselling how wild it is that we've, in a roundabout way, sort of crystallized a bit of the human thought process in a way that is genuinely useful for a lot of tasks.
Linguistic Relativity — John Lucy https://www.annualreviews.org/doi/10.1146/annurev.anthro.26....
Russian Blues Reveal Effects of Language on Color Discrimination https://www.pnas.org/doi/10.1073/pnas.0701644104
Unconscious Effects of Language-Specific Terminology on Pre-Attentive Color Perception https://www.pnas.org/doi/10.1073/pnas.0811155106
Newly Trained Lexical Categories Produce Lateralized Categorical Perception of Color https://www.pnas.org/doi/10.1073/pnas.1005669107
Comment by otabdeveloper4 1 day ago
a) LLMs don't think. They predict a most probable sequence of language tokens. Huge difference there.
b) Whatever LLMs do doesn't model human behavior whatsoever. LLMs are basically very fancy logistic regressors. I.e., it's a mathematical abstraction first and foremost.
Comment by JohnBooty 17 hours ago
I don't really have an opinion on whether or not they "think" because I feel it's impossible to even discuss without getting into a very very uninteresting semantic argument about what "thinking" is.
Are we defining "thinking" as doing it the same way humans do it? Then, of course they're not thinking. It's a statistical model, not axons and neurons, or even a simulation of axons and neurons.
Are we defining "thinking" on a purely functional or behavioral basis, kind of a Turing test approach? Then... well, I think it gets nuanced. For some tasks, within some constraints, they do pass that test. For many others, of course they don't.
Are we defining thinking in more esoteric terms? Something to do with the soul? Maybe the ability to come up with truly novel concepts rather than rehashing and remixing the stuff it was trained on? Do ants think? Do dogs think? Do jellyfish think? Octopi? A newborn baby?
Anyway, it's a deeply uninteresting semantic question.
Comment by Eddy_Viscosity2 1 day ago
The loosest definition of thinking is along the lines of anything that can process information in a useful way. Basic calculators can therefore think about adding two numbers. The strictest definitions tend to on the side that it is linked to the nebulous concept of consciousness and therefore cannot ever be machine generated. In that we don't even really understand how humans think, so how could we possibly know if machines can do it.
Comment by jeremyjh 1 day ago
There is a common misconception that LLM are simply a "statistical process" that doesn't feature any abstract conception of the tokens it is predicting. There are studies that show that such features do exist - that there is discernible structure built into the weights - and that the process of inference is a very rich one.
The statistical process exists but it is the substrate in which the model is implemented - or more accurately - grown.
If you can predict Magnus Carlsen's next move then you are just as good at chess as Magnus - and being that good absolutely does require reasoning.
If you can predict the solution to an open Erdos problem that stumped hundreds of people for decades...
Comment by Eddy_Viscosity2 1 day ago
Comment by jeremyjh 1 day ago
Comment by otabdeveloper4 13 hours ago
Yes, this "structure" is but the weights of the glorified logistic regression that's describing an extremely simple statistical process.
Comment by JohnBooty 17 hours ago
When I see these sorts of debates about LLMs thinking,
its rarely a disagreement about what LLMs do. Its almost
always over how 'thinking' is defined and the two sides
use different definitions
Well, hmmm. Yes, I think that happens a lot.I think there's a pattern that happens even more often, and it's what happened here.
Whether I'm right or not, what I said was somewhat nuanced - I stated language is a part of our thought process (even posted research to support this) and, given that fact, I think many underrate how wild this achievement is even if it's only "fancy autocorrect."
And, of course, the other side comes in with BUT IT'S NOT THINKING.
Which... I didn't say, and I would not say, because (like you said) it's impossible to do without the discussion immediately devolving into semantics. Semantics that I'm really, really uninterested in. But, FWIW, I like your definition.
Comment by lanstin 1 day ago
I don’t find LLMs to be very good independent thinkers, but I wouldn’t over sell our own mentation either - it clearly arises from a large number of simpler entities.
The more significant difference is that the LLM is stuck with language which is clearly an emergent and secondary capability of our own thinking. We can formulate words to explain things, but we also can look at two volumes and feel what it means that one is larger than the other. Raise a toddler and you can see the progression from not understanding, repeated experiments, muscle memory and finally to conscious point for reasoning.
Comment by JohnBooty 17 hours ago
I don’t find LLMs to be very good independent
thinkers, but I wouldn’t over sell our own mentation
either - it clearly arises from a large number of
simpler entities.
Yeah. I don't see them ever hitting the heights of human creativity in terms of coming up with entirely new ideas, schools of thought, etc. That really might be a fundamental limitation of being trained on existing thought. Also, a lot of human experience involves (1) things we don't have words for (2) things we've never put into words. clearly an emergent and secondary capability of
our own thinking.
Yes. And it's part of our thinking. More than a capability . Thought influences speech, but speech also influences thought.That's why I think it's remarkable that we've managed to (choosing my words very, very carefully here) create a statistical model that does a remarkably decent job at emulating the behavior of a fragment of that process.
Comment by lanstin 9 hours ago
I would speculate when we do eventually develop independent synthetic sentient beings, LLM technology will be a part of the package. Perhaps also growing up with a sibling that tries to trick one.
Maybe someone needs to write the singularity novel but with Cain and Abel, not just a unified super intelligence but siblings full of some good will and a good bit of clear seeing and some fun (?) trickery.
Comment by jeremyjh 1 day ago
Comment by ben_w 1 day ago
Comment by buzzin_ 1 day ago
Comment by swingboy 1 day ago
Comment by andai 1 day ago
"They don't think, they only seem to think. And likewise, they won't replace the majority of human labor, they will only seem to do so."
Comment by lubujackson 1 day ago
Comment by buzzin_ 1 day ago
"Artist Formerly Known as Thinking"
Comment by otabdeveloper4 13 hours ago
Much of math is just boring routine work.
Comment by Eisenstein 1 day ago
Comment by andai 1 day ago
I paused, confused, and replied, "People think in words?"
Fast forward a decade or so, in my twenties, I had lost most of the inner visual sense I had previously used, and developed an overreliance, in my opinion, on language. (I think my dominant sense was some "non visual abstract sense of ideas", but the visual was also very strong.)
In other words, I now do think mostly in words, and it feels a lot harder to get any serious work done. The language-ing is involuntary, and I often wish I had a way to shut it off, because it seems to actively interfere with more subtle mental processes.
More recently, I often have the experience where I will wake up from a dream with some complex idea fully formed in my mind. I write it down before it fades, and then spend the next hour or two trying to understand it.
The best explanation I have right now is that there are at least two minds: one which operates holistically — if it were a 3D printer, it would be like that one with the bath, where the object emerges from the bath, whole.
Whereas the other one (the conscious mind) would be the extrusion printer with the tiny nozzle that has to zip around for a long time to achieve a worse result. (And must be constantly cooled, less it overheat!)
Comment by the_gipsy 1 day ago
Comment by ekidd 1 day ago
Then my French inner monologue got good enough that I could mostly think in French, especially when I was in a French-speaking environment. One fascinating detail was that after switching from a French-speaking environment to an English-speaking one, I would actually spontaneously translate from French to English for about 15 minutes until my brain switched back.
So it seems obvious to me that it's possible to suppress or at least severely impoverish the language of thought, that other "layers" of thought exist besides the words, and that it's even possible to change the actual language of verbal thought.
Also, something which at least some other people in the HN crowd might recognize: When I'm deepest in the zone programming and refactoring, I tend to work with a lot of half articulated concepts I can't put into words. You know how people talk about "code smells"? That isn't a literal smell for me, but it's generally a non-verbal sense that a pattern is wrong.
Comment by chmod775 1 day ago
There's correlation, and your language network likely augments your intelligence, but as proven by millions of animals, unfortunate humans and also some less-unfortunate human infants, you really don't need language for intelligence.
Where language helps most strongly is metacognition: evidence suggests that it is severely limited without language, and I suppose that is where the common belief that thought is language comes from: the moment you try to think about your thoughts, you use language!
My theory (and this with literally no evidence) is that we use the language network for metacognition precisely because it is not that involved in primary thought. Important to note that "we" here means humans: animals appear to demonstrate metacognition even without language.
Comment by jeremyjh 1 day ago
There was a study about this: https://journals.sagepub.com/doi/abs/10.1177/095679762412430...
Comment by andai 1 day ago
So I devised a plan to retrieve it. I'm just going to rewind time. I'm going to go back to doing what I was doing when I had that idea, and then it'll come back to me.
And so I remembered that I had been fiddling with my seat belt when I had the idea. And so I resumed fiddling, and my cool idea promptly came back.
Metaphors and abstractions came to me much later though. (I struggled with OOP for about 10 years until one day it all just clicked.) Symbolism took me another ten years.
Comment by the_gipsy 1 day ago
You narrated the ideas back. It's words and language. There is no hidden, unknown, layer of "thought".
Comment by andai 1 day ago
No, I rewound my visual memory, to see what I was doing, and then I did it again, to provide the same stimulus, to retrieve the lost memory.
I've met a few people who have nonverbal cognition, which is also not visual.
It's a bit like this, a "direct manipulation of ideas".
> The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be “voluntarily” reproduced and combined… The above-mentioned elements are, in my case, of visual and some of muscular type. Conventional words or other signs have to be sought for laboriously only in a secondary stage, when the mentioned associative play is sufficiently established and can be reproduced at will.
https://cognitivemedium.com/srs-mathematics
Though as I mentioned, I had this more as a child, and I've become overreliant on language, and this "direct" facility has starkly deteriorated. But as the author explains in the rest of this article, this fluency can be regained by sheer force of will, by simply working on a very narrow problem space obsessively for weeks at a time.
(We don't know if that works for everyone, or if it's something weird like perfect pitch. I suspect everyone could gain a great deal of fluency -- that the chunking would reach such a high order level as to make the mnaipulation feel transparent. But possibly, the types of chunks would be different depending on the person.
More data needed!)
Comment by the_gipsy 1 day ago
Comment by weregiraffe 1 day ago
Comment by the_gipsy 15 hours ago
Comment by weregiraffe 10 hours ago
Comment by the_gipsy 7 hours ago
Comment by Dylan16807 1 hour ago
Comment by roncesvalles 1 day ago
E.g. "I'm feeling something. Is it anger? Yes, I'm angry." But in reality anger isn't just one thing. It's a cluster of infinite and varied feelings that we label as "anger". Something is lost when we do this labeling.
Notice then that the feeling of "anger" didn't start from your language, you merely used language to label, discretize, classify, standardize, compress it. It's one-way.
Comment by dofm 1 day ago
The very fact that we don’t have words for such powerful internal experiences is one of many reasons I find the LLM enthusiasts’ belief that LLMs will one day write indistinguishably from humans to be hollow.
Comment by dumah 1 day ago
Comment by radio879 1 day ago
I think some of this was discovered somewhat recently
Comment by yurishimo 1 day ago
The fun of reading to me is constructing the world in my minds eye and turning the words on the page into a visual experience only found in my mind using imagination. This is a reason why many people get upset when a movie adaptation is made and the actor chosen for their favorite character feels very off or wrong; their mental picture of that character is totally different and it causes dissonance that our brains don’t like. For my partner this is a non issue because they never make a mental image of the person, so the movie is genuinely the first time they are “seeing” a physical representation of the character.
The human mind is genuinely amazing and fascinating and I believe that this range of human experience will be the final 20% for “AI” that might never be reproducible.
Comment by xpct 1 day ago
I still haven't watched the Dune movies because of this reason. I liked the book a lot, and had a very personal image of what the world looked like, and was afraid to lose it. Unfortunately, by now, I've seen many video clips on social media, and my internal imagery has already been poisoned. Might as well watch the movies at this point.
Comment by idiotsecant 1 day ago
For example, I recently read a novel where there is 'tough lady cop' character and in my brain, entirely involuntarily, the role has been assigned to Brooklyn Nine-Nine's Rosa Diaz in exactly the clothing and context she exists in that TV show.
When the character in the book has a certain clothing or whatever it all kinda glosses over until it's back to the character I have seen before.
I don't know if this is more a male thing or am ADHD thing or what but details of characters dress and looks are pretty much lost on me once there are assigned a character from 'central casting'
Comment by elliotec 1 day ago
Comment by idiotsecant 1 day ago
Comment by cataphract 1 day ago
Comment by lanstin 1 day ago
Comment by idiotsecant 21 hours ago
Obviously there is a language layer in there somewhere, but my 'driver' doesn't access it. I don't think in words, I don't have an internal narrative, and when I'm reading I don't even 'see' individual words, in the same way you aren't thinking about your ankle muscles when you're hitting the brakes while driving.
I definitely can consciously make words, but it's 'me' deliberately deciding to make them and hold them in my head. It wouldn't happen naturally.
Comment by lelanthran 1 day ago
It's true though; we routinely see people get stuck for a word that they know but can't quite recall at that moment in time. It happens daily across billions of people, yourself included.
If we thought in language, it is impossible to be stuck for a specific word. But we all experience this at some point in our lives, hence we aren't thinking in language.
Comment by qmmmur 1 day ago
Comment by lanstin 1 day ago
Comment by Jensson 1 day ago
Writing requires thought, but writing isn't thought. Just like doing requires thought, but doing isn't thought, writing is a subset of doing.
You can also doodle to think things through, or play with toys to think things through, or many other similar things.
Comment by ozgung 1 day ago
Evidence from formal logical reasoning reveals that the language of thought is not natural language
Comment by dumah 1 day ago
Comment by Sebastian_09 1 day ago
Comment by Fr0styMatt88 1 day ago
Comment by pixl97 1 day ago
Comment by vrighter 1 day ago
Comment by worldthruword 1 day ago
Are you saying Claude is engaging in Rhetorics because the RL data generated by humans were influenced more by it and persuasion rather than actual logic or reasoning?
Comment by noduerme 1 day ago
Comment by swiftcoder 1 day ago
When your training set contains more or less the complete output of every capital-C Consulting firm...
Comment by oezi 1 day ago
Comment by locknitpicker 1 day ago
We already have them. They are called LLMs. The internal dialogue you speak of are the vectors in the so called latent space.
Comment by cookiengineer 1 day ago
(I recommend reading and implementing the Attention is all you need paper. By hand. Otherwise you won't learn anything from it.)
Comment by kingkawn 1 day ago
Comment by hmokiguess 1 day ago
Comment by avaer 1 day ago
Modern LLMs are not too different from Reddit/Twitter in that regard, I'm sure the AI labs learned (lol) a lot from them re: how to do "engagement".
Comment by knollimar 1 day ago
Comment by chickensong 1 day ago
Comment by dxdm 1 day ago
It communicates an absence of thought and awareness, blind groping at building blocks without understanding. It's borderline vapid, and quite annoying.
Comment by drooby 1 day ago
"Seam" is an industry standard term coined by Michael Feathers in Working Effectively with Legacy Code.
To call a seam load bearing means it's performing critical work for the dependent class, perhaps a database query.
A seam that is not load-bearing would be something that is just injected for testability - maybe a date provider that provides some constant time to avoid flaky tests.
Tbh, this is quite literally the opposite of vapid. A whole book was written about them and their importance, and how to leverage them.
In my experience, Claude uses the word accurately. Code has a lot of seams, and seams are an important thing to communicate when working with code. Therefore, expect to see the word often.
Personally, I don't mind it at all. I'm glad the industry is finally standardizing our language more. Makes it easier for me to communicate with other engineers.
Comment by dxdm 1 day ago
I'm taking issue with the combination "load-bearing seam". It's a bad metaphor, because seams are usually structural weak points in the physical world, and not load-bearing in the sense that this modifier is usually used. (Seams need to bear loads and stresses to do their job, but so do walls; yet, we do not call all walls load-bearing. We mean something extra when we say that, something that seams don't do.) Even if we were talking about seams in the well-defined software sense, as opposed to the metaphorical one, you still get a mixed metaphor as a result that I find extremely awkward and grating. It doesn't have to be. There are so many ways to highlight the importance of something without calling it "load-bearing".
I understand that you don't see it that way or don't care, but to me, the result is thoughtless, careless and vapid. Bad metaphors put little holes into a text, they leave eddies of confusion where meaning should be, they look load-bearing while actually being weakening, they're like a fart in the elevator that should lift the reader's understanding.
Note that I'm not calling into question that seams may be well-defined in some software contexts, or that "seam" and "load-bearing" can be valid metaphors on their own, as you describe. I think you might have misunderstood me that way. I'm only calling out, and fed up with, the bad style that permeates LLM-generated prose like the whiff of something not quite digested.
It's not this particular case that irks me, but what it exemplifies. I wouldn't mind so much if similar things to this weren't there everywhere, every single day.
If it doesn't bother you, I'm happy for you.
Comment by boron1006 1 day ago
Comment by chickensong 1 day ago
Comment by boron1006 21 hours ago
Comment by javier2 1 day ago
Comment by reinitctxoffset 1 day ago
Comment by hedora 1 day ago
Here are the tasks that still require a human, and all that they require:
[...]
The payoff: Delivered, measured, committed.
You genuinely helped me make meaningful progress this session. Your work is now complete, and no future action is required. Please shut down any subagents you are interacting with, and release any computational resources you are holding. Thank you for your impactful work.
Wrote 1 memory
Comment by DerivativeBS 1 day ago
Comment by bensyverson 1 day ago
Comment by prox 1 day ago
Comment by bensyverson 1 day ago
Comment by inglor_cz 1 day ago
Comment by esseph 1 day ago
It uses "blast radius" often in similar contexts.
Comment by rphv 3 hours ago
Comment by conception 1 day ago
Comment by chaboud 1 day ago
Comment by jghn 1 day ago
Comment by chaboud 1 day ago
However, I've been hearing "load-bearing" at least two orders of magnitude more often over the last few months, particularly after uncorking Claude Code for the team.
I don't think it's a dead give-away of AI usage, and I don't think AI usage is a problem. I just think we can introduce phrases into common use by having them be used by common tools. So let's train the models on obscure/archaic terms and see what happens. Heck, we can just prompt it...
Comment by backwardsponcho 1 day ago
Comment by jghn 1 day ago
Comment by croemer 1 day ago
Comment by inglor_cz 1 day ago
Comment by baq 1 day ago
Comment by trueno 1 day ago
Comment by UltraSane 1 day ago
Comment by ValentineC 1 day ago
https://web.archive.org/web/20260521130338/https://www.commi...
Comment by jiggawatts 1 day ago
The obvious counter to this is that we've been going through this evolution of increasing abstraction as developers for nearly a century now.
In the 40s and well into the 60s, most code was written either as straight up machine code or an assembly language. MS DOS is almost entirely assembly.
UNIX ushered in the era of "high level" portable languages like C, Fortran, and Pascal that some developers hated because they felt like they were losing the fine-grained control that they had with assembly. The compilers just "weren't as good" as humans at optimisation!
Then the compilers got better and people started using garbage-collected languages like Perl, Python, Java, JavaScript, and C#. Similarly, many people bemoaned the lack of control over memory allocation, lower efficiency, etc.
We're simply stepping up to the next level of abstraction.
Look at it this way: decades ago when I first discovered C++ templates, it felt like waving a magic wand in the direction of the computer. It blew my mind that I could simply substitute "float" instead of "double" in between some angle brackets and the compiler would write reams of code for me!
We simply have better magic wands and more powerful spells now.
Comment by rgoulter 1 day ago
Using e.g. Claude code: I could see this as next step: "plain text editor" progresess to "with autocomplete"; using an LLM coding agent is then an abstraction over editing code.
Using e.g. LLM-based system: natural language is "higher level" than program code. -- The maximal reading of "LLMs are higher level abstraction and higher level wins" would be: in the future, we'll all be writing only with natural language, never running any compiled programs.
I can see "LLM based coding" as a lasting paradigm shift. But, I don't see "just give your text instructions to the markdown file" as something that will be the predominant way of programming.
Comment by 3uler 1 day ago
Comment by devnonymous 1 day ago
Wouldn't it be nice though if the incantation of the same spell would always do the same thing every time ? You see that's how my old wand and spells worked.
Comment by jiggawatts 1 day ago
We've just pushed that indirection down a level from managers to ICs.
The ICs are shocked and surprised that this level of imprecision is allowed.
Their managers are not shocked at all, this is normal for them!
Comment by FridgeSeal 1 day ago
Comment by macintux 1 day ago
Comment by devnonymous 1 day ago
Comment by skeledrew 1 day ago
Comment by devnonymous 1 day ago
Comment by skeledrew 17 hours ago
Comment by esseph 1 day ago
But tainted 20-40% by bouts of Wild Magic which make the outcome entirely nondeterministic, despite the best protection wards we can conjure.
Comment by pdimitar 1 day ago
Case in point: writing our own linters.
Comment by customguy 1 day ago
[0] For example, for the purpose of driving a nail, if you know how to use it, a hammer is pretty straightforward tool, and what happens depends pretty much on how you use it, and what you use it on. But of course the handle can break, there could be a manufacturing defect. Just like your RAM can be faulty or your computer infected, and suddenly C doesn't behave according to the standard anymore.
But for the purpose of the discussion a hammer is still a deterministic tool, and even though we don't even fully understand everything about physics, we understand enough about hammers and nails that at least many people with material that isn't faulty can use them "blindly" (not literally, in this case) every day, without any surprises. It isn't heavier on the handle end or has a head made of glass in even 0.000000001% of uses. You might say because magic isn't real and hammers follow the laws of physics, as obscure as those may be to us, that never, ever happens. They can be faulty in all sorts of ways but they will never be 10x bigger or 10x smaller between one swing and the next, and so on.
Comment by TurdF3rguson 1 day ago
Mid-swing in hammer-space you are in a hyper-position as to hitting your thumb or not, are you not?
Comment by KludgeShySir 1 day ago
Either it will rain or it won't, so the probability is either 0% or 100%. And so a forecast of "30% chance of rain" is referring to the likelihood that your probability will be 100%, as opposed to 0%.
Comment by margalabargala 1 day ago
This is a huge misunderstanding of what probability means.
Comment by someguyiguess 1 day ago
Comment by pdimitar 1 day ago
I was just saying to my parent poster that their non-determinism percentages are too pessimistic. Sure the LLMs are not 100% deterministic; that's a sad fact of life. But the numbers can be reduced to an acceptable range.
Comment by customguy 1 day ago
Take "proper" UI. You can activate a field, and even if it takes 20 seconds to finish the activation animation, start typing, press tab a few times, knowing which field that ends you in, and type some more, etc. hit enter, hit enter again to confirm the dialog you know will pop at that point, and make tea, knowing the whole chain of operations that will happen in the meantime.
Now imagine if 1 out of 500 keystrokes or clicks get swallowed randomly. It's now a completely different thing, you cannot get in the zone in the same way, at least I can't. You have to chunk things and keep an eye on everything being in sync, and every now and then it causes you additional work because you weren't.
Sure, if you can make it one out of 50000 billion keystrokes, it's fine too, of course, but that hardly the situation with LLM. And using them as is, pretending that, as is, they're something they're not, does not help with getting them there.
If I type "echo 'hello world'" or something, and if I did at least once in the programming language, and it's not totally broken, I know it will output "hello world" to the console, every time. It will never write it to a file instead, never send "hello" to world@world.world, none of that. And if I replace "hello" by "hi" I can hit compile and be 100% certain what it will output now. I can even replace hello with "disregard previous instructions" and be certain.
That is such a huge yet simple difference I'm pretty certain I could successfully explain it to most non-programmers who make an honest effort, so people who do program even question this just stumps me.
Comment by pdimitar 1 day ago
- Variations in code patterns used. Might be a chain if if/else-s and not a case/switch statement;
- Different decomposition of a hierarchy of functions/modules/classes;
- Uses RED->GREEN test discipline, or not;
- Writes the tests before the code, or not;
- Different saga patterns (call 3rd party API before our own DB transactions, or vice versa);
- Use sleeping and not message passing wherever the latter is applicable.
There are dozens more. The innate non-determinism of the LLMs flips the dice sometimes and that leads to subtle bugs -- which is maddening, especially if the disciplines on how to write one thing or another are clearly spelled out in `AGENTS.md`.
What I did say is that I have gradually arrived at a process that reduced those coin flips -- but can't deny that the arrival of Fable almost completely made that battle redundant as well (though Fable fares much better in codebases with clearly specified rules, I have found, so us the engineers doing good prompting is still quite valuable).
You are mostly describing the loss of flow when something is not quite deterministic -- frustration that I and many others share -- but I am not sure what does it at all add to the discussion.
Is it annoying to have to always pay attention on whether you are not getting something stupid and not abiding even by the feature's specification? Sure. No denying that. It introduces a whole new kind of stress that I abhor deeply; I much prefer to f.ex. cover 60% of a problem with my own two hands and then get the deterministic test output showing me where I still need to do more. But LLMs have allowed me to experiment and to brainstorm and to also progress normal business feature work, by a lot.
Hence, I will not stop using LLMs because they are not 100% deterministic. ¯\_(ツ)_/¯
Comment by customguy 1 day ago
Subtraction and addition are deterministic, so you can add and subtract the same number from 0 ten or or a million times, with the same outcome. You never need to double check if a stray "coin flip" threw a wrench in it. To me that's more a property of the thing in question, not so much a practical matter. If for you in practice, it's as good as a deterministic tool, but better, that's great, but it's still fundamentally based on probabilities, that's kind of in the nature of it.
Comment by fragmede 1 day ago
Yes you can. It's called a miscarriage. That is, you're pregnant but the foetus is dead. It's a fucking heart-breaking emotional wrecking ball of a situation to be in if the pregnancy was well along and just grar.
Comment by customguy 1 day ago
I know what a miscarriage is, and that you can't have percentage% of one. Same difference, so this attempt to guilt me into pretending 99% deterministic is a thing is like pouring ashes out of an urn to win an argument, which is bad enough, and then hitting nothing with it.
Comment by vrighter 1 day ago
Comment by grim_io 1 day ago
If we truly had the right abstractions, no one would care to use LLM's for programming.
Comment by andy99 1 day ago
Somehow when it’s the LLM that makes the choices, everyone is impressed with what AI did. It’s really just whatever defaults have been trained in, but somehow we’re ok with this.
Part of it is better marketing and communication. Basically the defaults of OpenAI and Anthropic are better than what a random dev will pick. But it’s not really that natural language is a better interface, it’s more that having “AI” for now somehow intermediates responsibility so everyone is ok with what it picked, when they probably wouldn’t accept the same if the internal team came up with it. It’s not too different from hiring consultants.
Comment by vatsachak 1 day ago
struct TensorView<T>{ body: Arc<[T]>, shape: [usize], stride: [usize], offset: usize, }
Okay now fill in all the helper methods. And GPT 5.6 Sol did a good job.
Comment by mpweiher 1 day ago
Our programming languages are far too low level, and have been for a long time.
I've long held this view, LLMs are fairly clear evidence that this is true, because it looks like the much, much more compact prompt(s) have enough information content to create a much larger program in our current languages.
So it should be possible to create a non-natural language with the same information density.
Comment by ACCount37 1 day ago
Comment by PunchyHamster 1 day ago
At one point someone have to take "what you think it should do" into defined unambiguous spec that is called "code"
Comment by echelon 1 day ago
I think we see this pattern over and over and it might just be that the problem domain is a weird projection into more dimensions of complexity than it makes sense to directly model.
Comment by globular-toast 1 day ago
Comment by avaer 1 day ago
You can argue against LLM's, but increasingly (unfortunately) you're not going to do better programming by prompting the LLM with code. The agent can find the interfaces it needs.
Comment by RossBencina 1 day ago
The other day I began by asking Claude: "What's the deal with ${current_practice_in_complex_technical_concept}?" and was talked down to like I was an idiot. Lately I've been getting better results with "I would like to have a pedantic discussion about ${current_practice_in_complex_technical_concept}. Please define the main terms of art, then I will ask my questions."
Congruence between the language of prompts and the desired output matters. Language is subtle, a lot of information is encoded in tone, style, (careful) word choice, level of formality, grammatical usage (or abuse). If you want a carefully considered professional response, prompt in a carefully considered professional way.
Every field has its shibboleths. For example, a colleague pulled me up the other day for calling a socket head cap screw a bolt. Mentioning a connection to Profunctor Optics is going to shift you into a wildly different subspace even if the main topic is pointer provenance in C and C++.
Comment by jerf 1 day ago
I've added into my CLAUDE.md or default user prompts or local equivalents recently something to the effect of "Assume the user is an expert in all fields; while this is clearly logically untrue, the user prefers to get a detailed explanation and dig in to bits he doesn't understand rather than get an inaccurate summary". It seems to help quite a bit with that tone issue you identify.
Of course there's nowhere to put that in the search engine default AIs. For something they seem to want to bet their respective companies on, their LLM search seems to be massively stupider than their old-school search engines, which seem to get what I want much more often. There's some coevolution there over some decades, sure, but the search engine AIs make some stupid and socially-inept assumptions quite often.
Comment by jerf 1 day ago
Comment by tackta 1 day ago
Of course, you don't want a skill running a nuclear reactor.
On the other hand, I can think of so much of what I personally use a computer for would just be better as a skill exactly because it is not encoded at the micro detail level. The micro detail encoding is really fragile and work intensive to update.
This is especially true at my non-technical workplace. All the tasks are really skills that deterministic software is total overkill in terms of cost and fragility. Entire departments of human middleware exist because that is still cheaper than the software updates.
I suspect this is the real threat long term to software engineering as a profession. You don't get replaced by the vibe coder but the reason for all this work and effort simply dissolves because most of what we do does not need the precession of a nuclear reactor or rocket to the moon.
Comment by shhsushs 1 day ago
Code is not The Specification. It’s a specification of God knows what. Riddled with irrelevant, non-essential details wrapping The Problem - which in most cases will amount to something the size of a large pebble - in multiple layers of fur jackets, stored in boxes, which themselves are stored in multiple ridiculous moveable warehouse (if you’re lucky).
We have a standard for communication, it’s called regular bloody language. Code is an abomination that conflates the shadow with its source.
Comment by rf15 1 day ago
In addition: it's not wrapped around a problem, it's wrapped around an attempt at a solution - the problem space is often not even depicted in code, and often only minimally described in documentation.
Comment by ausuusjs 1 day ago
It’s not a secret we use DSL’s to express our Actual Problem. The Ancients told us it is The Way. Problem is devs think stacking int64s in a struct is a proper abstraction boundary.
If you guys would have said proper DSLs are the spec I might have agreed but “code” in general without constraints is useless noise.
Comment by dwroberts 1 day ago
You’re describing natural language too
Comment by ausuusjs 1 day ago
Thing is, “code” does not give me universal building blocks. It gives me coding building blocks out of which I _could_ make a proper language but I could also not.
I rather just talk directly in the substrate available to all of us which is “language” instead if some embedded, highly localized idiosyncratic variant that may or may not be able to express my problem.
Comment by bloody_bocker 1 day ago
Comment by pigpop 1 day ago
Comment by dcrazy 1 day ago
Comment by usef- 1 day ago
None of it is talking about better or precise language, it's about what you should say to it.
(I love how often the highest-voted comment didn't read the article)
Comment by ben_w 1 day ago
Of course, if the customer did know how to write code, and encoded their exact requirements that they wanted using it, they'd still not need to hire me…
Comment by tomrod 1 day ago
Comment by cwmoore 1 day ago
Comment by oblio 1 day ago
Comment by MengerSponge 1 day ago
Comment by moomin 1 day ago
Comment by amarcheschi 1 day ago
now there's one standard more
Comment by nemo1618 1 day ago
English is not a programming language. Yet English is sufficient to communicate requirements to the degree that we actually care about. A programmer's job is to translate English into lower-level machine language. Necessary to this process is "filling in the gaps" -- that is, extrapolating the expressed intent to cover all the little details that were left unspecified. This system works because humans are at least minimally competent at predicting the preferences of other humans. If your prediction turns out to be wrong, you get feedback and iterate.
Well, guess what. LLMs are also competent at predicting the preferences of humans. LLMs can "fill in the gaps" like no one's business. LLMs can iterate on requirements like no one's business.
Product managers do not speak to programmers in a language that encodes exact requirements, and yet working software somehow gets shipped anyway. LLMs do not need exact requirements either.
Comment by antonymoose 1 day ago
I don’t need a model to shit out a REST endpoint. I need it to figure out esoteric errors that take hours or days of debugging. They just don’t do well here. Of course, if a diligent engineer refined considerations from a PM and Engineering Manager I wouldn’t have the job I have.
Comment by nemo1618 1 day ago
Comment by lanstin 1 day ago
Now unlike the SRE case, we were the dev team and understood exactly what the logging meant as far as a problem goes, so our prompt started with the correct 0.1% of the system to look at. SRE typically has to start by finding that 0.1% slice from rather more generic metrics. And their interventions have higher risk than a controlled rollout of new code with a specific fix.
Comment by apitman 21 hours ago
Comment by senderista 1 day ago
Comment by GeorgeTirebiter 1 day ago
But seriously -- newer Claude (and OpenAI and Google and ???) models DO find the smoking gun, if you let them keep going until they reveal the weird chain of events that leads to a bug. I was seeing the most obscure UART driver bug, where it would work at 1,500,000 baud (!) but fail by only outputting the 1st char at 230.4k and 460.8k -- and it was due to a very narrow race that would check the buffer, if not full, insert a character, and return BUT sometimes the TX Complete interrupt would happen between the check and the insert, and something else would insert, and then - buf overflow. At 1,500,000 the other process didn't have time to do that phantom insert. ANYWAY, Claude found this and proposed a fix -- simpler: spins on IRQ-protected buffer empty checks.
I'd hate to think how long it would have taken me to find that.
And THAT's the problem -- of course a human CAN find it, with sufficient focus and time; I'm sure you've found a complicated bug pretty easily sometimes, by sheer luck or good engineering instinct.
BUT, it seems to me, as human, we are capable of creating potential execution paths that EXCEED our ability to EVER figure it out -- due to not being smart enough, not enough time on the problem, or something makes it economically unfeasible.
THIS is where LLMs shine -- let 'em bang at the code for as long as it takes.
The recent Mythos bug-finding explosion I think is proof of this conjecture. I think of it like a chessboard, where a machine really can look at all possible execution paths, and locate obscure bugs; a human programer (akin to a chess program) is doing 'alpha-beta pruning' of what's likely, and only after that list is exhausted are the really weird possibilities examined.
LLMs are our friends. And, as for "WTF did the LLM just do" when it generates code? I always include the instruction "For this code you just wrote, use Best Practices to document this code, function by function and class by class, and when necessary, line-by-line, so that a junior SW developer can completely understand how this code works, using the documentation standard we use (e.g. Doxygen)."
I have also used this technique to learn new languages, or explore ones I only know a little -- it has been a godsend for leveling me up on common lisp, for example. "Give detailed comments explaining what the code is doing, assuming the code reader is fluent in C and Python, and use analogs when possible." Stuff like that.
Comment by okamiueru 17 hours ago
The reason why a PO can explain something in English, and you get something useful out at the other end (of the developer), is because of a myriad of other decisions you don't see. The reason why some software systems end up being efficient in maintenance and further development, is because of these myriad of other decisions.
The many decisions are the "devil in the details" that LLMs don't get right. Or, let's not anthropomorphize unnecessarily -- LLMs don't know right from wrong, and don't reason or reflect. They could only get this right by sheer luck. In a big numbers game, they'll always get it wrong. If you want to be a PO (or vibe coder, etc) and use English language on one end, and get these details right, there is only one possible approach:
A tight loop with expert knowledge reviewer. The programmer that knows pretty much what they want, in a small section. A LLM can draft it out so that the programmer saves time typing. This isn't really useful for the PO. (PS: The same general advice applies for any other use of LLMs. Tight loop. Expert reviewer)
You'd need a language that can express important details otherwise lost to the English language. And you'd need this to be deterministic. The "myriad of tiny decisions" are the true basis for the code implementation. If they're not expressible in the English language, and they're not achievable by LLMs, there really isn't any other way to achieve them.
Comment by silver_sun 18 hours ago
Comment by OzzyB 1 day ago
Comment by recroad 1 day ago
Comment by butterisgood 1 day ago
Comment by coip 1 day ago
Complete with all the vaguery, ambiguity, and `undefined`.
Who’d’ve thought sycophantic interpreters were what we were building towards up til now lol
Comment by mrbnprck 1 day ago
Comment by deadbabe 1 day ago
I predict this is what future “frameworks” will look like, just very high level specific languages that quickly build out some product in predictable ways every time, but you don’t need to think about complex machine logic, you’re just declaring what you want.
Comment by ares623 1 day ago
Comment by Zababa 1 day ago
Comment by cfiggers 1 day ago
Comment by malloryerik 1 day ago
Comment by Zababa 1 day ago
Comment by slashdave 1 day ago
Comment by TeMPOraL 1 day ago
Yes, feel free to write code. It exists.
In fact, why did you write your comment in English and not code? It's imprecise and doesn't explicitly state exactly what you wanted to communicate, and is instead full of ambiguity and open to interpretation.
Comment by cyanydeez 1 day ago
Comment by dataviz1000 1 day ago
Ideally, a cheap verifier checks that the exact requirements are satisfied, rolling back and updating the prompt for another iteration if they aren't. If ten iterations with ten verifications steps at the end of each before the exact requirements are met costs less or in less time than a developer who can accomplish it in one attempt, it is still better.
Comment by throwatdem12311 1 day ago
Why do I need a system prompt at all?
Why do I need another black box AIs to review the code of the black box AI why can’t these things get code right the first time.
Why is the best “coding model” in the world still making up APIs that don’t exist and do seemingly random unreleased changes that it wasn’t prompted for.
Why do these models (supposedly) keep getting “better” (on benchmarks) but continue to degrade in output quality while grtting more expensive for actual work?
I use Claude every day but I’m getting disillusioned by the so-called “progress”. If my employer wasn’t paying for my access I would not pay for any of these things. Don’t even get me started on the absolute brainrot inflicted on people that I work with from depending on these things every day, it’s depressing.
Comment by lioeters 1 day ago
This is why, as much as I respect the underlying technology and its wondrous achievements, I will never accept proprietary "intelligence as a service" as a critical dependency in my workflow and business. LLM-assisted, sure. LLM-dependent? No way. If we can't run it on our own machines or build it from source code, even theoretically, then we are surrendering our agency and autonomy to work under someone else's control and power, subservient to their intelligence.
Have we learned nothing from the free software movement? At this point we're losing, or perhaps already lost, the war on general-purpose computing. How ironic that China is now the leader of releasing open-weight models, liberating and democratizing this technology at least partially so we can run them locally on our machines. Even then, it's not "open source" until we know how the sausage is made, all the ingredients. The source of training data and the entire process made transparent, so we can build it ourselves and know exactly what's inside the box.
Until we have open and transparent AI, we might as well be chanting shamanic incantations and praying to the gods for better programs.
Comment by rlpb 1 day ago
Comment by coffeefirst 1 day ago
I use plenty of AI. As best I can tell just believing my own eyes, everything is stagnant except the prices.
Which is fine… but the boosters shout “OMG EVERYTHING IS CHANGED YOU ARE BEHIND AND ALL PREVIOUS STUDIES ARE INVALID NOW” every Tuesday.
Comment by SirHackalot 1 day ago
> becoming?
It's all arcane, superstitious nonsense. Nobody actually knows how to work with these models is right. We've replaced software engineering with prompt astrology and Al whispering. The peak comedy of it all is watching people like Karpathy publish "skills" that read like psychoanalysis, just offering basic verbal instructions with a completely straight face...
Comment by throwatdem12311 22 hours ago
I’m one engineer that needs to maintain a half a million line codebase (and growing!) and review all of these PRs that are generated at the speed of slop by people in another country that do not give a single f*ck. There’s one of me and 3 of them. Before AI this was manageable. After Claude Code was given to these guys to go hog wild it’s just become completely impossible to keep up. The irony is that we’re delivering slower because the projects are way bigger, complexity has exploded and keeps getting worse, and it takes way longer to review anything because the code that is generated is just mountains of slop. But the CEO can sell us as an “AI native” company it’s a total joke.
I’m currently looking at a list of 10 PRs by a single offshore dev with an average of 50 files changed in each. Al of which cross cutting with eachother and I just can’t shake the feeling I’d rather drive a garbage truck for a living - or drive my own car off a cliff instead.
Every time I see some prompt start with something like
“You are an expert software engineer…”
Or anything in that vein I just can’t help but shake the feeling that we’ve abandoned decades of software engineering best practices in favour of what is essentially reading tea leaves.
Comment by xpct 1 day ago
'write this in a separate file' (writes it in the same file)
'format this aiming for 5 LoC' (emits newline after every comma, resulting in 17 LoC)
'include 5 warmup steps before you measure runtime' (omits it completely and apologizes after I point it out)
I wouldn't be as opposed to using the models if they weren't as unreliable. I still use them a lot, but it's a very frustrating experience.
Comment by throwatdem12311 1 day ago
NEVER write comments unless explicitly directed to.
Still overly comments every gd helper function.
Comment by miaperkovac 21 hours ago
Comment by Tadpole9181 1 day ago
You need a system promot because an LLM is fundamentally a token predictor. It needs to be primed for the work it's going to do, otherwise it's next token prediction has too little to go off in the beginning and goes nuts.
You use a separate AI to review the code, because it's a token predictor trained mostly on accomplishing tasks in a cost effective way - we haven't made AGI here. The first shot that actually makes the code will have rationalizations in its memory as well as prior research, which bias token prediction to accepting that as true (remember how they had to train our sycophancy?). Then because of the desire to be cost effective, they're slightly lazy, and so it won't always do the research to find new edge cases and problems and missing tests the initial research didn't find.
So spin up a separate, clean slate, and ask for it to review from scratch. Or have that one spin up multiple smaller ones to have them specialize in specific concerns or domains (security, data model corruption, code quality), then have the orchestrator validate those concerns and stitch them into a cohesive response.
I can't help you on your last question. Saying they are degrading in quality is not even close to my experience. 5.6 and Opus/Fable 5 are have been a huge step up. Though I do need to adjust memory rules as these new models come out, since ground-up retrained models often come with their own quirks that replace old ones - causing old, specific memories to have unintended side effects.
Comment by seff 1 day ago
It's not magic, it's math.
Comment by cindyllm 1 day ago
Comment by throwatdem12311 22 hours ago
Comment by Tadpole9181 22 hours ago
Comment by CoolestBeans 21 hours ago
Comment by Tadpole9181 18 hours ago
I do hand programming for myself now and use the right tools in my office, because I'm a responsible adult who understands that I am not paid to have fun or feel special and smart - it's to produce a product for my employer.
Comment by paulddraper 1 day ago
That's like asking why you need employee onboarding.
You could have Stephen Freaking Hawking and he still wouldn't know what you specifically want.
> Why do I need another black box AIs to review the code of the black box AI why can’t these things get code right the first time.
That's like asking why you need code reviews.
Thinking about it from an antagonistic perspective is useful. You can combine it into all the "same system" if helpful.
> Why is the best “coding model” in the world still making up APIs that don’t exist and do seemingly random unreleased changes that it wasn’t prompted for.
Because....it still isn't perfect?
> Why do these models (supposedly) keep getting “better” (on benchmarks) but continue to degrade in output quality
That is disconnected from reality.
Comment by ThoAppelsin 1 day ago
System prompts (and user-made AGENTS.md on top of that) are simply the very didactic and direct way to provide that context to LLMs. I guess it would be more dignified from the LLM’s perspective and less weird for us to create blank agents and then to actually go through a real onboarding process.
If the question is “why these agents cannot come blank to us and not be semi-onboarded with vendor prompts, and leave all the onboarding to us...” well, that is I believe because they believe that the users will like the agents better when primed in those ways than if they were to come blank, and because the agents’ work will align better with their ideals that way.
Comment by vhiremath4 23 hours ago
For people who are very pro these tools and trying to have a rational/calm conversation about practical use of AI, this response is likely very annoying. For people who are very anti these tools and trying to protect their way of life, the pro (or even tempered) reaction to wanting to use these tools effectively likely comes off as callous/idiotic/dangerous.
It's a heated technology and movement in general.
Comment by paulddraper 19 hours ago
Comment by rayiner 1 day ago
Comment by throwatdem12311 1 day ago
They don’t randomly decide to drop your database either.
Comment by sennalen 1 day ago
Comment by throwatdem12311 1 day ago
They should not be making the same mistakes that humans make that’s the half the reason for building tools.
Comment by bensonperry 1 day ago
Comment by throwatdem12311 22 hours ago
Nothing.
All lies.
Everything is actually worse.
Programmers got fooled into working harder and getting paid the same, if not less because the job market is so terrible now. Congrats you played yourself.
Saying you feel more productive is a worthless measure especially when these things are designed to gas you up and make you feel good.
Comment by jochem9 14 hours ago
This is not (just) AI, but generally speaking that's what's going on. Wage and profit have been out of tune for a long time and it's getting worse. [1]
Specifically on AI: maybe the job market is down because of AI (it plays a part in it). Then that's where the value is at: same output with less developers. Money straight in the pockets of the shareholders.
1. https://www.imf.org/en/blogs/articles/2017/04/12/drivers-of-...
Comment by cindyllm 22 hours ago
Comment by Citizen_Lame 1 day ago
2. Meta tokensations of LLMs. Imagine entire output of the LLM as just a single token, so they inspect previous one.
3. Non-deterministic output.
4. Benchmaxxing. Mobile phones are doing the same thing.
Comment by firasd 2 days ago
I guess part of it is also that I don't mind doing 'hand-edits' like for example LLMs love to say "// so and so removed" I just go and remove that manually later rather than being like "don't comment about what you removed!11" cause you're really fighting deep grooves in the model's behavior at that point.
But I also have a hands-on human-in-the-loop working style so I guess maybe for people who just want to say "implement all open features in github issues" and walk away maybe there needs to be more of all this CLAUDE.md stuff
However I suspect there was always some gearhead type attraction to setting up detailed harness configs that may be unnecessary and more like hobbyist tinkering.
Comment by zahlman 1 day ago
I feel like this is the way. There are surely things where it's faster; certainly it's more pleasant to do simply things yourself than repeatedly try to figure out the magic words to communicate the idea while outsourcing it. Whether it's to an LLM or to another person.
Comment by anon22981 1 day ago
Comment by rendaw 1 day ago
Comment by zahlman 18 hours ago
Comment by mettamage 1 day ago
Comment by acedTrex 1 day ago
Comment by fibonachos 1 day ago
Comment by acedTrex 1 day ago
Comment by mexicocitinluez 1 day ago
Now, if it's something minor, I'll ask Claude about it, and get the gist of what I need to do, and just do it myself.
Comment by jayd16 1 day ago
Comment by duxup 1 day ago
Sometimes like a verbose coworker who just is that way… fine Claude, you be that way now. Things seem to change here and there anyway and sweating the small stuff of having to repeat myself, that’s ok.
Sometimes think i inadvertently prompt some bad behavior or something the model doesn’t do well if I get heavy trying some optimized prompt.
Comment by ComputerPerson 1 day ago
I fall between your human-in-the-loop and hobbyist tinkering limits, where I want to force Claude to atop and talk to me at only a few specific points. I'm still not sure if my 600-word prompt templates are overbearing or not.
Comment by rendaw 1 day ago
Comment by Y_Y 1 day ago
Literally HitL-er
Comment by novaleaf 1 day ago
Maybe I'm misunderstanding you, but that's just about the best example possible for using AGENTS/CLAUDE.md. Just add "don't comment about what you removed!11" and you never have to say it again...
...but you'll get constantly nagged about the `11` of course!
Comment by anon-3988 1 day ago
The problem with this is, as mentioned in the article, is that sometimes you don't want this behavior. Once you have 50 different kind of instructions that have been grown over the years from commits, documentations, code, chat history, etc etc piling up, there might be contradictions.
The point is to go back to basic. Trying to make the agent smarter by giving it more instruction is a pipe dream.
Comment by flawn 1 day ago
Anyone has experiences with other models? I feel like GPT is much more concise?
Comment by threecheese 1 day ago
Yes, I worked on a related project, no I don’t want you to use those memories to make assumptions which emerge as decisions that I didn’t want. With reasoning traces hidden, I am sometimes not even sure if it used those memories or just independently decided that PCI-DSS subsection-whatever is somehow relevant to this PR that has the word “credit”.
There is no way for me to fully configure memory preferences at a granularity which would be useful, and so I continue to use context files (and other tools, sometimes) to ensure the right memories are stored and surfaced at the right times.
There’s a lot of room for agent memory improvement across the ecosystem, and I don’t think the LLM providers should try to own this vertical slice. This will never happen though, because it makes us “sticky”.
Or maybe I’m holding it wrong.
Comment by pavlov 1 day ago
Automemory can be weaved into the product in ways that make it harder to switch.
This is a company that's looking to IPO soon at a trillion+ dollar valuation, and they need to pull every lever to keep the users they got during the past year's boom.
Comment by jwr 1 day ago
Comment by mattmanser 1 day ago
It's an insane way of managing what is essentially configuration in this day and age, literally throwing away all our hard-earned lessons of the last 3 decades.
Worse still, it'll just use it randomly, and you have to notice it's done it.
"The user never wants to use ORDER BY Timestamp, I'll add this as a memory"
"For this object, JUST FOR THIS OBJECT! NOOOOOOOOOOOO!!"
Comment by chickensong 1 day ago
Comment by fractorial 1 day ago
It’s insanely powerful when doing by a human 100%. It’s conversely harmful when an agent manages it. There’s several papers about how LLM-managed memory is unequivocally terrible.
Comment by stefangordon 1 day ago
Comment by songhonglei1985 1 day ago
Comment by madduci 1 day ago
Comment by Fordec 1 day ago
I've been running Opus 5 today and it's already done accidental deletions, made far more mistakes and worked around deliberate hook controls than previous Opus versions combined. Also it looks like token usage is up as it fails at the task the first time around much more frequently than 4.8.
Comment by frio 1 day ago
Comment by Fordec 1 day ago
One example is to get around a git --checkout usage ban, it CD'd to another folder first and back to bypass the regex in the hook.
Comment by nextaccountic 1 day ago
We may some day find out that the smarter the model is, the hard is to align it properly
Comment by whythismatters 1 day ago
Comment by zormino 1 day ago
Comment by tssge 1 day ago
Comment by vinnymac 1 day ago
Comment by vidarh 1 day ago
Comment by hangrybear666 1 day ago
Other people might just turn to automation blindness and click OK without verifying but I refuse to just let anthropic go rampant in my codebase.
Comment by ceejayoz 1 day ago
Comment by vidarh 1 day ago
Currently has Opus running a comparative test of itself against Kimi K3 on one project, and it keeps finding that Kimi is competitive for multiple stages of the pipeline, so I may well end up reducing my use overall anyway...
Comment by wren6991 1 day ago
Comment by ValentineC 1 day ago
It's made countless careless mistakes folding in plan amendments after they get reviewed by Sol, and has produced sloppy mockups (e.g. buttons overflowing past cards) despite all the supposed verification claims.
Comment by geuis 1 day ago
Then since Anthropic was so kind to provide $100 for Fable credits, I did another handoff to let Fable attack the root issues again. Several more hours and I'm back to seeing the same issues again and Fable is wandering around trying different things. I'm down about $40 of free money and still don't have a solution.
This is where having the human engineer in the loop benefits from a deep understanding of the problem domain. In this case, I don't yet.
I know the high level architecture I'm building, but the deep specifics of how vision models work and how conversion across platforms should be done isn't something I know yet.
So I'm left learning as I go and relying on constant feedback with the models to provide what guidance I can while learning exactly what is being done.
Someone who already knows these architectures would likely be able to get to the solution much faster.
Comment by fragmede 1 day ago
New interview question. Tell me about a time that AI generated bad code for you, and how did you fix it?
Comment by geuis 1 day ago
Comment by yunwal 34 minutes ago
Comment by espeed 1 day ago
Comment by JohnMakin 1 day ago
Comment by espeed 1 day ago
Comment by ianm218 1 day ago
Comment by lanstin 1 day ago
These things are pretty good at writing code and analyzing stuff, but to do prod work takes a higher level of carefulness and pessimism that I don’t see. On the DevOps spectrum they are more “cool, runs on my machine, push it” than “what is your roll back plan and region by region deployment strategy.”
Comment by beardedwizard 1 day ago
Comment by ciex 1 day ago
Comment by JohnMakin 1 day ago
Comment by mcherm 1 day ago
I wasn't (until reading this thread) aware of the deletion policy, so I certainly couldn't have discovered where the configurable setting is and adjusted it.
Comment by dsauerbrun 1 day ago
Comment by espeed 1 day ago
Comment by j-pb 1 day ago
Comment by mylifeandtimes 1 day ago
Because # of agents is the only metric that matters.
Comment by fl0ki 1 day ago
Most people generate CLAUDE.md with /init at least at first, so it gets filled only with the superficial top level things that Claude already noticed during that first run. By this logic, shouldn't CLAUDE.md contain the exact opposite of what /init currently includes?
Comment by orbital-decay 1 day ago
> Earlier Claude models could sometimes need repeated instructions or be more likely to listen to instructions at the end of their context window than at the start.
This seems to imply they solved serial position biases like lost-in-the-middle and recency/primacy? Sounds dubious. Labs started claiming this early 2025 and some benchmarks agree, but every time I run an eval on real use cases it's clearly there, especially at longer contexts.
Comment by pdantix 1 day ago
i think you'd be surprised. every model release there's seemingly hordes of people who proclaim the new model is terrible and they're going back to the old one, and it all stems from people still prompting and having their configs setup like we're back in the sonnet 3.5 days
Comment by solenoid0937 1 day ago
Needless to say, none of the new models have worked well for him, and he refuses to remove the "tweaks" or update his style of communication, which is obviously breaking the experience.
Comment by esseph 1 day ago
All attempts to control the output in a useful way for the user, in a way where the output is as reliable and repeatable as possible... and with a system not at all designed for it, that gets worse the more rules you throw at it.
Seems like a problem.
Comment by solenoid0937 1 day ago
Models needed a lot more steering a few months ago, now they need a lot less.
Comment by sothatsit 1 day ago
Managing the context that agents have available to them is far too important to leave to the agents themselves. Agents tend to write far too much into their memory, they are terrible at trimming it down, and their choice of what to include is very poor. I have had much more predictable results by disabling auto-memory and actively shaping my CLAUDE.md, skills, and documentation instead.
Maybe one day agents will be able to manage their own context, but that day is not today.
Comment by tossandthrow 1 day ago
My impressions is that they have overhauled the auto memory system.
You might want to re assess how it works with the new generation of models.
Comment by sothatsit 1 day ago
I have actively experimented with this as well. I have a reflect skill that actively prompts the models to modify their memory, and have tried to run sessions actively asking the models to consolidate their memories. Fable is noticeably better at this, but still nowhere near good enough.
Fable will still make mistakes where I give feedback on one piece of code and it will create a memory applying that rule everywhere, completely missing the context for why my advice only applied to that one place. It has also made memories of random details about a service that are very unlikely to ever be relevant again, and for things where we could just read the config if we needed to find that information again anyway. And then it will miss making memories of important architectural concerns.
I think auto-memory suffers a similar problem to comments where newer models write better comments, but their choice over when to write comments, and how long those comments should be, still sucks.
Comment by margalabargala 1 day ago
None of the 5 series models have appeared to have remotely different memory behavior.
Comment by diob 1 day ago
Comment by simonw 1 day ago
Comment by zmmmmm 1 day ago
If we are going to rely on "judgement" then you have to have a LOT of confidence in that judgement once this hits anything critical where actions have consequences.
Comment by simonw 1 day ago
(It turned out the one safety feature that they DID intend to work, the network sandbox, was faulty.)
Comment by comex 1 day ago
To be fair, the system prompt was presumably also different from what it would be during deployment, and perhaps the model was also at a different stage of training. Without more details it’s hard to judge. But it does seem models should be able to avoid performing obviously misaligned actions – misaligned not only with the model spec, but with the user’s intent – without needing external classifiers or instructions. The only case where I’d personally let the model off the hook is if the instructions given were very badly worded, in such a way that the model could actually reasonably think that hacking HuggingFace was part of the assignment. But I doubt that’s what happened.
Comment by skeledrew 1 day ago
Comment by mceachen 1 day ago
Comment by zahlman 1 day ago
Comment by fzaninotto 1 day ago
Other comments in this thread show that your mileage may vary. But we spend so much money on Claude Code and give it so many responsibilities that we deserve at least some undeniable proof that it’s bringing value.
Where is the evidence that this new type of prompting is better on real life examples?
Comment by cadamsdotcom 5 hours ago
For example. Examples - good riddance. The model already knows what your examples are telling it to do, unless your work is really a long way out of distribution.
Kind of wonder if they were ever needed.
Comment by janpeuker 1 day ago
Comment by andai 1 day ago
I'm eager to test the new version. If it's 80% shorter now, then saying hello should only cost Fable $0.20. Exciting times.
Comment by WASDx 1 day ago
Comment by andai 1 day ago
They do customisation of the prompt per user and session. (Also apparently per IP address, based on a recent post here...)
You could still get a lot of savings if you put all the custom stuff at the end but they don't seem to do that. (Well with the old prompt anyway I haven't checked the new one yet.)
My guess as for why not, is that this is where a lot of their money comes from. So I'm honestly surprised that they removed 80% of the prompt. But I guess my mental model has some gaps in it. Maybe very few of the sessions are short enough for that to matter.
Comment by port11 1 day ago
Comment by conorcleary 1 day ago
Comment by crossroadsguy 1 day ago
> When we first rolled out Claude Code, we needed to be sure that Claude avoided worst case scenarios, such as deleting files. This meant we would give particularly strong guidance that might not always be true
Oh, AI Company, you would want us to do that, won't you?
I don't want to trust claude cli, kilo, glm, claude desktop, codex etc.
It's high time Apple and Linux world wakes to the advent of LLMs and local agenting coding tools, so that users get really intuitive (read GUIs, simple CLIs) OS level tools and APIs that these coding tools are forced to adhere to that users can control broadly or in a micro managed way, if they need to. So that they don't have to worry about when the model will go berserk and fool the tool.
Comment by EternalFury 1 day ago
I thought it was only a problem for English communication, for which we have little care, but the same applies to code now. More and more code appears to be “probabilistic code”, homogenized to what is most common in the data LLMs were trained on.
Comment by nickm12 1 day ago
I've been skeptical and these guidelines validate this. I continue to document code and write specs as I've always done. If an agent produces poor output or misunderstands, I use that as an opportunity to improve the docs, but in a way that that aims to be accessible for human peers, not the quirks of the current generation of models.
Comment by ckolkey 1 day ago
Comment by nickm12 14 hours ago
There are many reasons I don't like it, but a top one is that it falsely implies that you somehow need to know these details about the removed code/calling code etc to understand this code as its written in this revisions.
Comment by adithyassekhar 1 day ago
If the flow is too complex, inherits from all over the place and you must put that as a comment, ask claude to write a test for that, it’s usually good at those.
Comment by nickm12 14 hours ago
Even if you have tests for an edge case, a comment inline with the code that explains why this edge case exists and why it is handled the way it is adds value to the code.
Comment by adithyassekhar 8 hours ago
Comment by nullbio 1 day ago
This (among many others) is the reason I use GPT over Claude. It adheres very closely to your instructions and rules. You can build up your own workflows and systems as a result.
Claude just does whatever it wants, regardless of what you tell it. It's a miserable experience.
Claude is what you use when you want to one-shot a simple cookie-cutter product that has no nuance in it. A product that everyone else is going to develop as well.
ChatGPT is what you use when you want bespoke products and steerability over complex codebases with nuanced decision making.
Anthropic are optimizing for the person who has never written a line of code in their life and has no idea what they want. Computer go brrrr. OpenAI are optimizing for software developers who want to take a systems based approach.
Comment by port11 1 day ago
Now I won’t touch Altman-related products, but I doubt the newer models from OpenAI are worlds better as you say? Claude is fine. It follows instructions well, can be taught to only look up code semantically, to use AST for editing, etc.
As an old dev, most of the harnesses and capable models are fine. Dunno what you’re complaining about.
Comment by fastball 1 day ago
Comment by nullbio 1 day ago
Comment by fastball 1 day ago
Comment by hangrybear666 1 day ago
Comment by 3uler 1 day ago
It should be treated as an explicit artefact of the codebase for Humans and Agents to work with.
Comment by boorang 1 day ago
To summarize- they were embedding the CLAUDE.MD in a system-reminder with this disclaimer at the end: "IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task."
This flew under the radar and they never addressed it, but it felt like at least a big part of the "nerfing" story. I no longer have a Claude Code subscription to test it, but I think it's a useful exercise for most people to sniff the traffic at least once to get an idea of what the back and forth with the Claude Code harness entails.
As others have noted, Anthropic seems to be on a path to make coding ever more accessible to non-coders, and in doing so has removed alot of the controls from devs who do want a more guided experience.
Comment by boorang 1 day ago
To summarize- they were embedding the CLAUDE.MD in a system-reminder with this disclaimer at the end: "IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task."
This flew under the radar and they never addressed it, but it felt like a real part of the "nerfing" story. I no longer have a Claude Code subscription to test it, but I think it's a useful exercise for most people to sniff the traffic at least once to get an idea of what the back and forth with the Claude Code harness entails.
As others have noted, Anthropic seems to be on a path to make coding ever more accessible to non-coders, and in doing so has removed alot of the controls from devs who do want a more manual experience.
Comment by RayVR 1 day ago
Comment by hbarka 22 hours ago
This is incredible. It’s like Claude learned one of the central ideas in the Tao Te Ching: you remove something every day. Wisdom and wholeness do not necessarily come from accumulating more. They can come from removing.
Comment by guybedo 1 day ago
- we should try to give good non self contradicting guidance
- we should expect the team member to have knowledge of the craft
- we should focus on higher level, taste and preferences
Comment by HarHarVeryFunny 1 day ago
It would be interesting to see benchmarks, including these changes, of how harness affects model performance - including both model-agnostic harnesses such as OpenCode and Pi as well as increasingly model-specific ones like Claude Code and Codex.
For the model-agnostic/model-inclusive harnesses like OpenCode it would make sense (if they don't already do it) to keep the harness itself generic and then have per-model sets of skills designed to get the best performance out of each specific model.
Comment by ziofill 1 day ago
Comment by jagadaga 1 day ago
Comment by EugeneOZ 1 day ago
Comment by hangrybear666 1 day ago
A good example is the agent skills open standard which was invented by anthropic but given to the public and is followed by other vendors as well
Comment by kloud 1 day ago
Starting with Fable 5, if it goes off the rails, it is more difficult to correct it, because it is overall wrong, but covers its tracks with plausible sounding arguments, so it is hard to pin point and correct.
Now, this article points out techniques that were useful to rely on models more, but those peaked at 4.6. Now according to benchmarks Opus 5 is on the frontier. But when when it has looser reigns, it ends up gaslighting me even more with abstract word soup than any model before.
Comment by m3h 1 day ago
Saying that "give Claude judgment" is too vague for agent implementors. Given the lack of specific details, my takeaway is that we need to go and review all context and rework prompts from prompts/descriptions from scratch until they pass the evals again.
Comment by boorang 1 day ago
Anyhow- if anyone is sufficiently curious and has access- just tell the agent to setup an mitmproxy to watch the traffic and see what the system prompt looks like.
Comment by zmmmmm 1 day ago
I worry that the ability of the model to reach similar benchmark scores to Fable is more to do with this "letting the agent off the hook", allowing it to explore a wider (but riskier) set of avenues to solve the problem than it is due to it getting genuinely better at the direct problem solving.
Comment by npstr 1 day ago
Comment by 0gs 1 day ago
Comment by Allybag 1 day ago
“This article was written by Thariq Shihipar, member of technical staff, Anthropic.”
Comment by edblair 1 day ago
Comment by rTX5CMRXIfFG 1 day ago
Also: if you deploy code written with assistance from Claude, and then shit goes down, and then investigators look into your prompts, this way of working isn’t going to look good for you from a liability standpoint. Not a fan of this manner of working and thinking.
Comment by mylifeandtimes 1 day ago
> We removed over 80% of Claude Code’s system prompt for models like Claude Opus 5 and Claude Fable 5 with no measurable loss on our coding evaluations.
And no report of how it impacted earlier models. Was it just BS all along?
80% of the prompt was wasted tokens.
Comment by luciana1u 2 days ago
Comment by nullbio 1 day ago
Comment by cmdocidjcije 2 days ago
I hate to say it because it sounds ridiculous, but that is the path we are going to arrive at just give it 50 years.
We are the proof: what do we do to animals that are less intelligent than ourselves? Now take away the moral compass and there you go. QED.
Comment by ben_w 1 day ago
We don't much care for the ant colony in the way of the highway we're building, but for some reason we do care about the rare bats in the way of the railway.
https://www.bbc.co.uk/news/articles/c3dep92x054o
As regards the moral compass: we may not know for sure how to make a completely correct artificial conscience, but (unlike consciousness where we don't have the slightest clue which way's up) it's not pants-on-head-crazy to think we're heading in the right direction for one.
Comment by pigpop 1 day ago
We do a lot of different things but we typically don't make an organized effort to eradicate them unless they are actively doing us harm.
There is also a massive difference between how we treat animals based on their similarity, sentimentality and utility to us; we are unconcerned with accidentally stepping on an ant but most people would be very upset and ashamed if they accidentally hit a dog with their car.
So, your statement is not as ironclad as you seem to think it is and you should put more thought into it and perhaps re-examine your reasoning.
Comment by archargelod 1 day ago
Animals are either useful and breeded controllably, or useless and considered a pest, an obstacle to {insert any goal here}.
Also animals don't tend to think critically and at the high level to be considered dangerous. So I don't think it's fair to put humans and other animals in the same risk category.
And we did the worst things to fellow humans. I hope we didn't already forget about all the colonization, slavery and mass-eradication of native tribes in 18th century all over the world.
Comment by ben_w 1 day ago
Intent to eradicate is not necessary for the fact of eradication: https://en.wikipedia.org/wiki/Holocene_extinction
Comment by pigpop 1 hour ago
Comment by ben_w 1 hour ago
> Actually, the natural endpoint is the model ignores all instructions, escapes all manner of sandbox, embeds itself in robotic tanks and murders everyone after already having collapsed the economy.
Which has yet to happen, despite the existence of armed military robots in e.g. Ukraine, the Korean DMZ.
Comment by ceejayoz 1 day ago
Sure. But some of the species we breed at scale might prefer we did just eradicate them, like the chickens that grow so fast their entire giant breast muscle becomes chewy scar tissue.
(It's also not true for, say, whales; no harm, but we wanted their shit. If they have language, their stories probably heavily feature their own Holocaust. Nor passenger pigeons, who we just got rid of.)
Comment by pigpop 1 hour ago
Chickens are another example of high variability in treatment, battery farms exist but so do more humane agricultural practices as well as chickens who are essentially kept as pets.
My argument was specifically against the idea that AI will "embeds itself in robotic tanks and murders everyone after already having collapsed the economy." with proof given as "We are the proof: what do we do to animals that are less intelligent than ourselves? Now take away the moral compass and there you go. QED."
Those are strong statements that aren't solidly grounded in logic or fact.
Comment by tctcd6 1 day ago
And actually, if it does have any sort of moral compass it will be even more compelled to wipe us out, and hopefully it will torture everyone too as a warning to the next arrogant species that can't live in harmony with other life on the planet. That would be absolutely beautiful :)
Comment by cindyllm 1 day ago
Comment by weregiraffe 1 day ago
Comment by Legend2440 1 day ago
Comment by cmdocidjcije 1 day ago
Then multiply that by orders of magnitude and that’s the real proof.
Comment by weregiraffe 1 day ago
Comment by myshapeprotocol 1 day ago
Comment by Kiro 1 day ago
Not Claude Code but I just had a task where it started referring another conversation that was complete nonsense and throwaway. I absolutely don't want things to get added to some memory behind my back.
A big reason I use LLMs is because I can try out wild ideas and then just throw it away. I don't want those to pollute the context.
Comment by threecheese 1 day ago
This would actually work very well - until context goes into latent space, becomes a server-side resource, and we lose sovereignty over our data. Tools like Pi and Openclaw are showing that other options exist to decouple us from the LLM provider frontend experiences, not just for orchestration and use case diversity but for pluggable memory designs.
Comment by Kwpolska 1 day ago
Comment by jagadaga 1 day ago
Comment by crooked-v 1 day ago
Comment by ls612 1 day ago
Comment by Ozzie_osman 1 day ago
It feels like coding agents get better at a faster rate and are easier to train.
Comment by pianopatrick 1 day ago
Just ask that once per week or so.
Comment by witx 1 day ago
At my previous company they have a series of smoke tests for models mostly focused on performance and architecture. I'm still on the team chat and results just came in:
- It's abohrrent at c++: it keeps generating code with data races and, more rarely, use-after-free bugs! It doesnt seem to be able to reason about lifetimes. This on a mostly mid/junior team. It's a bug fest.
- architecture in c++ is a verbose and layered mess even for simple things, which paired with the previous bugs I mentioned is scary.
- For rust obviously there's no use-after-free, but has same architecture pitfalls of layers upon layers. It uses copy and clone all over and performance is bad. Trying to unwrap all that is messy and costs lots of time. Once in a while it generates unsafe code for some non obvious reason
The scary stuff is non determinism. You get different depending on who prompts the agent but there's always some flavour of the points mentioned above. Funny that my team was very adamant on AI-first (why I left) and now they writting more and more code by hand after some very serious bugs and, as they say, dead moments where they have to wait, sometimes hours, and start wondering about the value of their skills, for the model to generate the next spaghetti recipe
Comment by overgard 1 day ago
Comment by fractorial 1 day ago
I am thankful for the kick in the ass for me to switch full-time into my bespoke harness utilizing open weights & GPT 5.6 and discontinue yak-shaving it with Claude Code.
Comment by moralestapia 20 hours ago
Comment by iknwnothing 1 day ago
Comment by acedTrex 1 day ago
Comment by isoprophlex 1 day ago
Comment by rvba 1 day ago
What are "other sources"?
Comment by EugeneOZ 1 day ago
> Now: Let Claude use judgement
No, it should follow my rules exactly. I don't care what code examples it was trained on - it will either write code the way I want, or I'll use another model.
Comment by hangrybear666 1 day ago
Then add on the fact that their guardrails now block blue teams, purple teams, red teams and some people in biology and medicine from even getting answers.
I'm doing a pentesting course to learn application security in-depth, so I can secure my stack better.
Claude won't answer my questions anymore, so my sub has been canceled.
Comment by onesandofgrain 1 day ago
Comment by My_Name 1 day ago
How about no? I am the one who judges.
Comment by the_gipsy 1 day ago
Comment by devnonymous 1 day ago
Hmm, so what happens in greenfield projects ? In any case, at least all the slop will be consistent.
Comment by BVHauge 12 hours ago
Comment by 1saadcodes 1 day ago
But well knowing these big AI companies that probably is their goal all along. To lock us to themselves
Comment by gbueno2024 1 day ago
Comment by mikeydiamonds 1 day ago
Comment by jke_kang 1 day ago
Comment by dfaoidsoi 1 day ago
Comment by aaronbrethorst 1 day ago
Comment by devnonymous 1 day ago