Engineering management after the cost of code collapsed
Posted by kiyanwang 2 days ago
Comments
Comment by dbingham 2 days ago
The assumption is that LLMs should be writing the code and human engineers reviewing and verifying the LLM output. And that this pushes the cost of producing down. And I fundamentally disagree with that.
Every time I ask LLMs to write code, even with Opus 4.8 (haven't tried it with Opus 5 yet), what I get ends up being totally rewritten. LLMs still aren't good at writing maintainable code. Can they write plausibly functional code? Yes. But it won't survive the long term. People using LLMs to write all their code are gambling on them eventually getting to a point where the LLMs can fix their own code. It's possible, but I wouldn't necessarily bet on it.
Where I have found immense value from LLMs is in code review. Repeated review by LLMs catches an amazing amount of potential issues. They really shine on security review, but are very effective with any kind of review.
The other thing that the "LLMs write code camp" misunderstands is that writing was never the bottleneck. Understanding was. And understanding the code is still the bottleneck. But understanding is truly gained during the writing loop. The understanding you gain from pure reading or code review is marginal compared to the understanding you gain while writing.
Most of the time previously spent writing was actually spent updating and deepening our understanding of the system under development. There's no replacement for that understanding in a world where LLMs are doing the writing.
But if you flip it: humans write, LLMs review, then you still get a major gain -- not in speed, but in quality. And you keep the understanding loop intact. I would propose that this might be the best way to deploy LLMs.
Comment by SoftTalker 2 days ago
But to be fair human code reviews have the same problem. It's like reviewers feel they have not done their job if they don't find something wrong.
Comment by sheept 2 days ago
I do the same for LLM code review comments: some changes are out of scope or could be moved to a separate PR; some edge cases don't happen in practice and should just fail noisily instead of writing more code to maintain. When these are the only issues it's raising, then I know it's done.
Comment by YZF 2 days ago
Comment by davidkoenit 1 day ago
Comment by RyanHamilton 2 days ago
Comment by davidkoenit 1 day ago
Comment by bluefirebrand 2 days ago
Comment by clbrmbr 2 days ago
Comment by dbingham 2 days ago
Comment by SoftTalker 2 days ago
I wonder if the same trick works with AI?
Comment by clbrmbr 2 days ago
Comment by wonnage 2 days ago
Comment by wuschel 2 days ago
Perhaps someone more knowledgable could jump in here to clarify?
Comment by mike_hearn 2 days ago
Comment by chucky_z 2 days ago
To be clear I mostly only use Opus and Gemini Flash but this might work for others too.
Comment by chrisjj 2 days ago
Disabling your C compiler warnings works too. You get to ship then leave work early!
Comment by gste 2 days ago
Comment by tripleee 2 days ago
Comment by therealpygon 1 day ago
How much of a grasp does one need on the UEFI code operating their computer? The drivers that operate a hard drive? The communication protocols for their monitor?
This seems to me to be the fallacy — that all code must always be understood to operate properly or to be useful. If the AI is sufficiently intelligent enough to understand the code, at some point (most) people won’t need to. Not saying that’s now, but that is likely to be where things are headed in the long horizon. That said, I don’t have a crystal ball, nor do the people claiming that won’t be the outcome, so right now it seems more like people arguing about which forks should be used while the cake is still in the oven.
Comment by dalemhurley 2 days ago
Comment by preg_match 2 days ago
But, with AI, the cost of code has gone down more than it already has. Well, cheap things are easy to throw away. So you prototype, prototype, prototype, and close the loop as much as possible with the customer. True agile development, not big A Agile.
The problem is this requires alignment from management, and we're just not seeing it at many company. They can't grasp that things have changed, and that throwing away code is free. They don't trust engineers to close that gap, so customers and stakeholders are still waaaaaay over there and we're delivering features they don't want.
Comment by eikenberry 2 days ago
Comment by prymitive 2 days ago
Comment by addaon 2 days ago
Comment by bluefirebrand 2 days ago
Until managers give up on Agile, we'll never get time to actually write specs
Comment by whatevaa 2 days ago
Comment by written-beyond 2 days ago
Comment by gste 2 days ago
To be honest, the models are getting so good that they do most of this unprompted now.
Comment by theptip 2 days ago
At this point if you can’t get the agent to write good code then either I) you are in a very specific niche (like Karpathy trying to write NanoGPT) that is extremely out-of-distribution, or II) skill issue, you need to learn how to prompt better.
It’s fine to have a skill gap! Just don’t delude yourself that the tools are bad and everyone claiming they are good is wrong.
Comment by dbingham 1 day ago
This has been the pro-"write with LLMs" argument the whole time. "If you're not getting good results, you're not prompting well. FAANG are doing x% of their development with AI."
Yet all of the companies that I can think of that build on planet scale infra and have gone all in on LLM development have seen a marked drop in quality and stability since adopting LLM driven development.
So yeah, I don't buy it. When I hear that engineers are writing their with LLMs and I see quality improve or stabilize, then maybe I will start to question whether its a prompting issue.
Comment by piker 2 days ago
Comment by theptip 2 days ago
Planet scale infra often requires some novel ideas at the architectural layer (the engineers can input those) and then it’s mostly in-distribution C++ / Rust / Go; most of the hyperscalers open-source their stacks, for one. But for two, a ring buffer, look-aside cache, deterministic hash, b-tree, lsm-tree, etc. are all well-known patterns.
Hyperscaling is also often a very clear objective function; i need this code path to run in this many microseconds/nanoseconds, so i can hit the scale numbers I need. Claude/Sol can extract prod logs and build a representative micro-benchmark, and then hill-climb on it autonomously. Thats straight up the fairway for the training set, even if it often has a high bar for finishing and requires sophisticated Workflows or lots of tokens to explore the search space.
E2A: sorry sorry, typo, I meant microGPT. I can see why this would be confusing.
Comment by potatolicious 2 days ago
So despite its importance much of it is actually pretty in-distribution.
Comment by piker 2 days ago
Comment by tripleee 2 days ago
Comment by bigstrat2003 2 days ago
Comment by gamblor956 1 day ago
Comment by disgruntledphd2 1 day ago
Personally, I'm more on the LLM code is pretty bad from a design point of view, particularly when much of the code is already LLM generated.
Comment by oasisaimlessly 2 days ago
Comment by dwayneII 2 days ago
Comment by sanderjd 2 days ago
Comment by lioeters 2 days ago
Well put. This insight is worth repeating in every discussion on the subject, from software engineering to mathematics.
The problem is that understanding is not the product being sold. The business model is for everyone to become consumers of what the magical genie generates, where the "understanding" is kept on the side of the model providers. This ensures a future generation of consumers dependent on someone else to provide the understanding.
Otherwise, you can create your own answers based on actually understanding the code, theorem, proofs, etc. Smart consumers of LLMs will use them to increase their own knowledge and understanding, so that the service is augmenting their intelligence, not replacing it.
Comment by clbrmbr 2 days ago
The key I’ve found is human peer review. The reviewer jumps on a live call with the developer, pulls up the PR with transcription on, and asks questions. At the end of the call, the transcript passes back into the coding agent and the PR is polished up, becoming more self-documenting, and the humans are left with some degree of common understanding of what’s going on.
I’ve been operating my team of ~15 this way for 9mo to great effect… there is simply no going back to the stone ages.
Comment by oogali 2 days ago
Can I watch/observe one of your review sessions?
Every few weeks, I hear the beginnings of a great approach towards working with LLMs but I rarely see it in practice.
If you're open to this, remote or in person, ping my username at gmail.
Comment by tracerbulletx 2 days ago
Comment by davidkoenit 1 day ago
Comment by disgruntledphd2 1 day ago
I think for refactoring, you're definitely right as if you give them a good high level sketch you can get them to do all of the more tedious implementation.
Comment by supriyo-biswas 2 days ago
Comment by lukeschlather 2 days ago
Comment by lordnacho 2 days ago
This was my stance a couple of years ago, but now I've given it up.
It turns out writing actually was the bottleneck. You can understand perfectly well what you want, but writing it is long and tedious to the point where you find excuses not to do it. Particularly with version 2, the step where you have an OK system and you want to improve it. Quite a lot of changing the code is just useless busywork: re-wiring old functions, moving imports around, searching for locations that benefit from extracting a common piece of code. And each time you do one of those, there's a decent chance you did something even more trivial like forgetting a semicolon or calling the wrong function.
Now that I have an LLM helping me, I can see why. The critical decision is a terse declarative like "we need to have several TCP connections instead of one, and just use the sequence number to arbitrate". A human junior programmer could perfectly well understand what this meant, but he would have to go through all of the above to get to the final product. Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
This means I can use my attention on the things that matter. So instead of spending today thinking about how to arbitrate between the TCP connections and tomorrow thinking about pre-calculating my outgoing orders, I can just do both today. I don't waste the good waking hours chasing minor bugs, I just think about the large structure.
I get the feeling the best programmers of years past were actually masters of the little things, which led them to be able to look at the big things. Essentially it was cheaper for them to get to the top of the mountain, where you can see the landscape. Kinda like how the kid who was good at mental arithmetic in primary school was also good at calculus at the end of high school: if you don't have to concentrate on the little things, you have time for the big things.
Comment by twister2920 2 days ago
how did you decide to pick the most trivial kind regression for this example? do you compile your code before checking it in?
> A human junior programmer could perfectly well understand what this meant, but he would have to go through all of the above to get to the final product. Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
the main efficiency you have described here is offloading the verification of a change onto the LLM. that is the bottleneck. readers can decide whether a non-deterministic statistical model is a good tool for this job
> I get the feeling the best programmers of years past were actually masters of the little things, which led them to be able to look at the big things
the best programmers understand that their job is to automate workflows, and that includes their own. if you're worried about missing a semicolon, I'm sorry to say that's a skill issue
Comment by lordnacho 2 days ago
Why would this be a regression? You might just be writing a new line of code.
> do you compile your code before checking it in?
Well obviously. That is generally how you discover that a semicolon is missing.
> the main efficiency you have described here is offloading the verification of a change onto the LLM. that is the bottleneck. readers can decide whether a non-deterministic statistical model is a good tool for this job
No, it's the time between you deciding something needs to be done, and it being done, that is the bottleneck. You cannot avoid trying to compile the code and testing it. Now you can get to that test without paying attention, which is time you can use productively.
> readers can decide whether a non-deterministic statistical model is a good tool for this job
Somehow, the non-deterministic model has built me the deterministic code that I want, very fast, pretty much all the time. A year ago it would get stuck. Now it doesn't, for me at least, and for competent programmers that I know.
> the best programmers understand that their job is to automate workflows, and that includes their own. if you're worried about missing a semicolon, I'm sorry to say that's a skill issue
Well yeah, and I've automated my workflows completely. I don't have the problems I used to have. If you haven't caught on to the new way of working, well, that's a skill issue...
Comment by maccard 2 days ago
I still find the models get stuck or go on _massive_ side quests. Just today, I asked claude to write a hello world C++ program using import std; I interrupted it when It decided I needed a new toolchain installed, and started checking for docker installations. This is super basic stuff, it hadn't even generated a plan, it just started searching for LLVM versions rather than running clang --version.
> If you haven't caught on to the new way of working, well, that's a skill issue...
Honestly, it feels like the emperor has no clothes on this topic, and the crowd defending LLMs to death are way too quick to call it a skill issue.
Comment by lordnacho 2 days ago
I feel it's the other way around. The LLM skeptics are unwilling to admit that these things can get you there faster than you would on your own, in the face of clear evidence.
Comment by bigstrat2003 2 days ago
Comment by twister2920 2 days ago
I use LLMs, but they're just a tool in the workflow, and I make sure to review the output. they might remember semicolons but they make much more pernicious mistakes that are harder to detect
Comment by lordnacho 2 days ago
Comment by twister2920 6 hours ago
Comment by leptons 2 days ago
Speak for yourself. Writing code has never been a bottleneck for some of us. I can't speak for everyone, and neither should you.
>Now, I can just tell the LLM and I will get what I want, even with the things I didn't explicitly state, without spending attention.
This should worry you. All too often the LLM invents things I didn't ask for and implements things I didn't need. YMMV, I guess. If slop gets the job done, and nobody notices, then who should care?
Comment by Valectar 2 days ago
If other people are dissatisfied with LLM output quality while it seems to work fine for you, you might want to consider that the quality of code you produce is closer to the quality of code the LLM produces than what those other people are producing.
What you posted there, for example, about most of changing code being busy work is a pretty big red flag for a codebase. One of those "large structure" things that you're supposed to be paying attention to is the architecture of the code. There's always the chance that some change you need to do goes against the grain of the solution you architected, and you need to make changes all across your codebase to fit it in, but in general the point of modularity and good architecture is that when you make a change you just have to make that one change, ideally just changing the logic of the one responsible function with only minor changes required anywhere else in the codebase. If you're consistently having to hunt throughout the code for related functions that you need to rewire that's a sign that your architecture does not fit with the direction your codebase is evolving, or alternatively that you don't have much of an architecture to begin with and your code is highly interconnected.
Actually one habit you mention at the end of that quote can worsen this issue: "searching for locations that benefit from extracting a common piece of code". Tautologically this is a good thing as you define it as only working on locations that will benefit, but given the frequent need for rewiring of functions I would hazard to guess that you've "deduplicated" code a bit overzealously. Just because two functions share some common code does not necessarily mean it is appropriate to pull that out into a function. Deduplicating is good if conceptually the code is a single thing that you would always want to keep in sync, as it means that when you need to make a change to it you don't have to hunt down all the places it's used. On the other hand, if you find yourself frequently needing to delve in to these functions to rework them because you need to make a change to how it's used by just one caller, your "deduplication" has added to your workload, and probably created some overcomplicated code in the function that is in reality handling multiple distinct needs.
I hope this doesn't come across as too condescending, and if I've just wasted your time explaining principles you already understand I apologize. I don't know you or the code you're working on so I can't exactly confidently judge your work solely on a few paragraphs. It's just that your mention of how your experience of coding has been different from what others have described, and specifically that, for you, writing has been the bottleneck rather than understanding, combined with the specific issues you describe facing, imply to me that you may not realize that the approach you are taking to producing code yourself may be significantly different from how other Software Engineers are producing code, and that may account for some of the differences you note in your personal experiences programming.
Comment by lordnacho 2 days ago
1) It was good for me to spend years learning the little stuff. Loops, variables, if conditions, how to import stuff, git, debugging things, reasoning about the flow of control. Classic coding.
2) I had a false dawn at about 10 years in. I thought I understood a lot.
3) I learned I had a lot to learn. Very wide areas of programming I'd never touched, ways of thinking that started to click.
4) I spent another ten years covering holes, building a different type of experience. My guesses about how to do a project are much better now. My guesses about what really matters have changed.
5) Now the small stuff is actually just bothering me. I'm not going to learn much more from staring at little things. There are larger architectural things to think about, and the little things are just friction.
So that's where I'm coming from. I get that a lot of pushback is going to be from 10-year-me, who thought he'd gotten to a high level of understanding by slogging through the little stuff.
Comment by luaKmua 2 days ago
But that doesn't mean they're useless either. I use them all the time for review as you mentioned or to knock out one-off scripts that don't go anywhere near source control. There's just no world where I don't need to understand every line of code that I'm responsible for getting into our project.
Comment by lukan 2 days ago
Debugging code step by step is how I understand complicated code.
Comment by KronisLV 1 day ago
Anyone who's seen tech debt where each item is well known and understood but just big in scope knows that this isn't true universally - depending on your team size and composition, any damn thing can be the bottleneck, often at different times too. I'll take everything that helps me resolve them with reasonable trade-offs, even if I need to come up with ways to mitigate the issues created by those trade-offs (like enforcing >90% test coverage as a starting point).
Comment by YZF 2 days ago
Are they as good as handcrafted code by 0.1% of top software engineers. Generally no. But neither is 99.9% of real code.
LLMs also are good at code reviews. What they'll miss is often the big picture but they can still catch plenty of issues. I still want to see a human in the loop in my domain.
Totally agree that writing the code was never the bottleneck. We're not seeing massive productivity gains even if some code is written faster. It's not just about understanding but also various other activities that happen in large companies and teams.
Also agree LLMs can be used to gain quality but realistically most orgs are going to aim for "fixed or decreasing" quality at lower costs.
Comment by nicce 2 days ago
They are getting more and more hostile for making any security assesments. I wonder will they even write secure code in the future if they can’t point vulnerabilities from existing code.
Comment by himata4113 2 days ago
However, sometimes then I tell it to write an app with detailed instructions and it spits out garbage so your mileage might vary.
Comment by davidkoenit 1 day ago
Comment by win311fwg 2 days ago
I find that depends on the target language. They can be good at writing maintainable code, but not consistently across every language.
The languages beginners usually gravitate towards are especially hard for LLMs to produce quality output for. Presumably this is due to the training data including all the unmaintainable codebases written by beginners in those languages, which hasn't allowed the LLM to converge on recognizing what a maintainable codebase looks like in those languages.
Comment by ManuelKiessling 2 days ago
If LLMs are, as you stated, really good at catching potential issues, then they are, almost by definition, really good at producing code without potential issues, if guided correctly: all they need to do is inspect and iterate, until they do not find any more potential issue in the code they produced.
Comment by layer8 2 days ago
Comment by booleandilemma 2 days ago
Comment by ilovefood 2 days ago
Comment by bcrosby95 2 days ago
I've been using it for a unity game for the past few years. Nowadays it will go sleuthing into packages and assembly and make decisions based upon what it sees there.
It will make comments about why it's doing something based upon a function call 3 methods deep.
God forbid any of these details change in a minor version update.
Comment by hkpack 2 days ago
Comment by drTobiasFunke 2 days ago
Comment by layer8 2 days ago
Comment by consumer451 2 days ago
I do know what good code looks like, but does that even matter anymore? All I know is that now, I get to focus on endless UX polish, which is the only thing the matters.
I feel like we are living through something like the Protestant Reformation, where priests once spoke Latin, and then started to speak in plain local language. The old guard did not like this.
Comment by Silhouette 1 day ago
AI pricing is mostly based on tokens consumed right now. Shouldn't that mean being able to quickly and reliably analyse existing code and to make only small local changes to implement new functionality is as valuable as ever - if not more so - if you're relying on agentic LLMs to do the grunt work?
A lot of things about writing clear specs and developing systematically and employing lots of different kinds of checks and controls to ensure quality and performance have always been true but used to get brushed under the carpet by a lot of cheap/lazy development teams. If LLMs really do accelerate everything about development - including negative behaviours like acting undesirably based on flawed or ambiguous information and doubling down on mistaken assumptions - then the pattern across all of these areas is that doing things the right way is more important than ever if you want to get good results from AI assistance.
Comment by bigstrat2003 2 days ago
Maybe so, but it doesn't.
Comment by broast 2 days ago
Comment by rufius 2 days ago
Comment by vatsachak 2 days ago
Comment by marginalia_nu 2 days ago
If programmer productivity was something we actively optimized for, we wouldn't have crammed programmers like sardines in warm and noisy open floor offices with 2000 ppm CO2 levels and then further constantly interrupt them with emails and slack pings and meetings all day long, Jira rigmarole wouldn't make up a significant portion of what they did, programmers would have instead mostly been thinking and programming.
We've always had the ability to 2X if not 10X the output of each and every one of those poor souls. You don't end with this sort of programming purgatory because it's a productivity optimum, it very clearly isn't, but because it's a billable hours optimum and/or an org chart clout optimum and/or because of Jevons paradox got hands even in business management and the IT department was allocated too many dollars.
Comment by raverbashing 2 days ago
I disagree. Kinda
What AI has made much simpler is that you don't have to waste time checking docs and have the best autocomplete system by a long shot - this was a bottleneck unless you were doing Java or some other language with "perfect" AC
What AI made "kinda easier": solving for usual problems. The stuff you would search Stack Overflow, or think a couple of minutes for an optimized solution - not a bottleneck but not 100% smooth neither
You still have to test and validate your code. AI made this easier-ish but this is still where I see manual work being needed (even if you are automating tests - you still have to think on what you want the code to do)
Comment by bitlad 2 days ago
I think now, code is the bottleneck. Just because you can generate million lines of code, people with different skill level think they are accomplishing the task, testing, merge conflicts, trust has become the bottleneck.
Comment by parpfish 2 days ago
The “old way” would be lots of debate (both bike shedding and useful) among engineers during design phase, and then you’d implement.
Now it’s shifted so there are no design docs and there is only the generated prototype. People trying to do their design review while there’s already a functional-ish prototype and it goes nowhere. There’s an anchoring effect in place because the first thing already exists and management says “this seems to work, just use it and move on”. The result is that useful debates about substantive issues don’t happen and bikeshedding is all way get to do
Comment by drTobiasFunke 2 days ago
Comment by skydhash 2 days ago
And then you get paged at 2am because prod is down and the support channel is more active than the team's one.
Comment by raverbashing 1 day ago
A working prototype beats endless discussions over paper
Comment by dns_snek 1 day ago
Comment by drTobiasFunke 1 day ago
Comment by parpfish 1 day ago
Comment by fragmede 2 days ago
Comment by marginalia_nu 2 days ago
In larger organizations, quite often it's the business that is holding back development. They can only handle so much change and speed needs direction to be velocity. Drafting requirements is generally much slower than implementing them.
Like the number one complaint from programmers has been that they don't get to do programming. They want to write code, not update jiras or spend hours in meetings.
Comment by sdevonoes 2 days ago
LLMs help, but they haven’t been trained on our own repos. I don’t need the LLM to help me with algos that are available online… I need them to help me with custom business logic
Comment by baron3dl 2 days ago
There are a lot of (excruciatingly) long-form posts about what folks are pioneering but not a whole lot of follow up about what failed. Where are the short posts on the negative space? How did halving your staff work out? Flattening your org? All those dark factories, what haven't they produced? How about all the other things tried, failed, and unceremoniously scrapped?
We need to explore and communicate the negative space more efficiently. Don't repeat the same mistakes, and don't make me read 2653 words when 300 do it better.
Comment by VeninVidiaVicii 2 days ago
Comment by baron3dl 2 days ago
Similar to a functioning side project in the 5-10k LOC range. Announcing something that worked a year ago, was laudable, even if not profitable.
I vibe coded 15k LOC this morning and read 20k words of AI generated text while doing so. No longer are either noteworthy or valuable public contributions just by virtue of having been done. I don't think that's widely recognized yet.
Comment by tempodox 1 day ago
The recognition is represented by one word: Slop.
Comment by w10-1 2 days ago
Isn't that the essence of "attention is all you need"?
Comment by coffeebeqn 2 days ago
A million times this. Can we please RL the next models to learn the “if I had more time I would’ve written a shorter letter” method please.
I see it every day in tickets, many communication channels, PR descriptions, comments, documentation. All have at least 70% verbose fluff which is so taxing and makes it very hard to keep track of the one important thing they’re trying to communicate in the message
Comment by baron3dl 1 day ago
Comment by sanderjd 2 days ago
Comment by mgaunard 2 days ago
Comment by leptons 2 days ago
Comment by pfannkuchen 2 days ago
I’ve been wondering whether 2006 Google, Amazon, Facebook would be acting like this if LLMs were launched back then, or if these companies are just due for getting replaced and this is just their big mistake making phase. Prior big companies also entered a big mistake making phase, the mistakes just looked different because different era etc.
Comment by thih9 2 days ago
New tech stacks will appear and they will handle the mess to some extent.
Comment by oenton 2 days ago
With this and the pipe dream of continuous generation and deployment of code without a human in the loop, I can't even think of a problem it's solving or attempting to solve. The closest thing I can think of is library or language upgrades... but we already have automated solutions for those e.g. codemods and migration scripts.
Comment by leptons 1 day ago
React has about 3 to 4% of the top 10 million.
Comment by leptons 1 day ago
Comment by thih9 1 day ago
[1]: https://www.rocketsoftware.com/en-us/insights/what-is-cobol
Comment by leptons 1 day ago
Well that's some cherry picking. And really off-topic, because nobody here mentioned COBOL.
If someone were shitting on COBOL here, it would be right to point out that COBOL is still getting the job done at 90% of fortune 500 companies.
Comment by swat535 2 days ago
If your engineers are "wasting time" optimizing artisan code whilst the competition has released their next version, they'll be told to use AI.
The other reality is that by the time you figure out the right abstraction, business has already pivoted, or your feature will be rewritten , or dumped all together. Obviously there are niche industries where this is not the case but in majority of companies, the churn is exhausting.
My point is this: we can shout all we want about "maintenance" and "technical debt", but it's guaranteed to fall on deaf ears.
LLM has validate upper management's assumption that engineering is nothing but a cost center.
I think that our industry is fundamentally shifting.
Comment by Matumio 1 day ago
Imagine a world where those only get done for a business incentive by people with this hard-reality business mindset. Or not done at all, reinvented at every place, no transferable knowledge for developers.
Comment by mgaunard 15 hours ago
Comment by VeninVidiaVicii 2 days ago
Comment by 650 2 days ago
Too many teams and organizations have business types, mostly PM's who seek to lord over their area of know how and see themselves as delegators and mini CEOs, actively avoid looping engineers in to validate themselves. Engineers need to take on PM roles, and the PM role needs to be 1:50+ eng or go.
Comment by georgeburdell 2 days ago
1. Pursuing polish and quality beyond previous norms
2. Replacing $100/mo/seat SAAS with something coded by a junior costing $200/day to develop over months.
The cost of code approaches zero, but the cost of having accountability, and hosting remains the same, and so individuals need to only coordinate to the extent that those things remain finite resources. Management needs to stop insisting that their directs adopt each others vibe coded tooling.
Comment by sanderjd 2 days ago
Comment by antonvs 2 days ago
Not everything has to be written as though it’s a middle manager’s idea of what makes for a good TED talk.
Comment by nvme0n1p1 2 days ago
Comment by aplummer 2 days ago
Comment by nwah1 2 days ago
Comment by ilovefood 2 days ago
> What follows is a cleaned-up version of notes I accumulated over the past year. Gemini 4 helped with the editing.
The image is made by an AI image generator on fal.ai. It's better I spare you all my design skills :)
Comment by hgomersall 2 days ago
Comment by sodapopcan 2 days ago
Comment by jboss10 2 days ago
Does this guy have access to Gemini 4 already?
I'm guessing Gemma 4 was happy to be mistaken for Gemini and didn't catch this mistake.
Comment by trollbridge 2 days ago
Comment by dualvariable 2 days ago
I find it exhausting trying to read anything that smells remotely like linkedin AI slop at this point.
Comment by ilovefood 2 days ago
Comment by dataplumb3r 2 days ago
The smaller models can be sufficient for coding but for document writing not highly specific I've yet to be satisfied with AI output. I certainly wouldn't expect gemma to produce good outputs.
Comment by ilovefood 2 days ago
Thank you for the suggestion.
Comment by add-sub-mul-div 2 days ago
Comment by ilovefood 2 days ago
Comment by j45 2 days ago
LLMs have brought a different unlock, and for everything we're seeing become easier, it allows people learn to use the tools to take on solving problems that couldn't be approached before.
Comment by glimshe 2 days ago
Comment by j45 1 day ago
The focus is placed by those who ask question and have a platform, and most don't have a tech background let alone implementing tech or software in businesses.
Comment by chickensong 2 days ago
This is a long-standing problem related to operational excellence and politics. I expect this will improve with AI adoption and integration. You can't get an exec to create a decision record and commit it to git. Managers have incentive to sequester information.
Engineering already has the discipline (maybe) and abilities to solve the problem. Version control, change control, ADRs, logging, structured docs, etc... We can trace an inbound packet or call through the entire stack. Management can't/won't do anything remotely close. 1-to-1 emails, meeting minutes, stale Word docs is the standard for most.
Inserting LLMs as the interface, and/or plugging into existing interfaces like email, is going to change things. Finally it will be possible to capture more institutional knowledge, without trying to teach an old dog new tricks.
Comment by raffraffraff 2 days ago
I work at a company where the biggest problems are not 'writing code', they are:
- Organising teams
- Designing the system
- Prioritisation of work
The fuckups that we make on a daily bases are not 'code errors' they are failures in THOSE three things. I'll go into detail if anyone cares.
Comment by treetalker 2 days ago
Comment by convolvatron 2 days ago
Comment by sanderjd 2 days ago
Comment by siliconc0w 2 days ago
Comment by ilovefood 2 days ago
Comment by mikewarot 1 day ago
Words have meaning, and I strongly feel that clarity here is helpful.
An engineering approach to software development would include rigorous testing and require individual signoff for every library and module. It is quite clear that an LLM would not be able to meet this requirement.
When casting about for ideas or prototypes, the throw away nature could allow their limited use.
Comment by tiago_human 2 days ago
AI is doing a good job on writting code these days! Nothing against it; I use it every day, but the context switching is costing us a lot!
Comment by OutOfHere 2 days ago
As I understand it, the purpose of management is to match financial resources with material+human resources to perform feasible tasks. There is nothing here I see that can't be done by an experienced token generator. If anything, automating management seems easier than automating engineering.
As for leadership, it can be done by the investors.
Comment by pillefitz 2 days ago
Comment by OutOfHere 2 days ago
Or if you're saying you still do engineering/product work, then perhaps it shouldn't come with a management title. Whispering task allocation ideas to the AI is for everyone to do.
Comment by aabdi 2 days ago
As we’ve collapsed the cost of the operations then largely the point of such an llm is constructing above.
If you’re an engineer why wouldn’t you pay more for better abstractions in an llm? That’s half the work anyways.
To put some more context into this: a manager defines the shape of a group of people. They own amorphous blob of responsibilities and various services. This leads to context confusion and lack of ownership and diffuse ability to operate services.
To fix the manager goes: Team a is responsible for say the device platform. Team b is for the applications platform. Each is responsible for their domain and the abstraction of team. Each is then responsible for their own ops, roi, quality, else. This attempts to maximize consistency of context, incentives, and scaling. Whether it works or not is frankly up in the air.
But notice this is basically the same as deciding in your intra service modularization and how you define the interfaces. Can you objectively say that whatever abstraction you usually write is correct? No. You just hope with experience and pragmatism.
Comment by visarga 2 days ago
Comment by OutOfHere 2 days ago
Comment by pillefitz 2 days ago
Comment by OutOfHere 2 days ago
Comment by whinvik 2 days ago
Comment by 0gs 2 days ago
Comment by trollbridge 2 days ago
Comment by CharlesW 2 days ago
Pangram's "100% AI generated" claims are right 65% of the time. https://link.springer.com/article/10.1007/s40979-026-00226-w
Comment by jtorsella 2 days ago
Comment by CharlesW 2 days ago
Comment by jtorsella 1 day ago
Comment by CharlesW 1 day ago
Comment by jtorsella 1 day ago
Comment by generic92034 2 days ago
Comment by selimthegrim 1 day ago
Comment by bonzini 2 days ago
Comment by ilovefood 2 days ago
> AI is good at coding if there's an oracle. If the system is ancient, unreadable, untestable, that's exactly the opposite. It won't get the exact set of corner cases.
I sort of mention this in the article, so I'm sure we're somewhat aligned on the core. How would you have worded things?
Comment by bonzini 2 days ago
The thing that I like the least is the headings and the way LLMs always try to put a punchy line. "Correctness time splits in two" says nothing if you don't know what it splits in. Maybe "making it correct vs. describing what's correct"?
Another trope is short sentences: "Good engineers used to say this before LLMs, and now nobody can argue about the sunken cost of having written that code." instead of the longer and redundant "Good engineers said this before LLMs. It was true then. It is enforceable now in a way it was not, because nobody can argue that writing more code was the hard part".
Another clearly AI paragraph is "Plumbing time collapsed. Scaffolding a service, generating tests, translating between frameworks, writing the first draft of a migration: all of this is fast now, and any timeline built on those costs deserves compression." Instead: "The time to bring up a proof of concept or refactor old code has compressed, and you should take that into account when planning your timeline".
There are videos on YouTube about AI style, you just need to learn them and undo them when they're the most blatant.
Comment by Mohiuddin7 2 days ago
Comment by cineticdaffodil 2 days ago
Comment by arjie 2 days ago
Comment by mathgeek 2 days ago
Comment by cineticdaffodil 2 days ago
Comment by cineticdaffodil 2 days ago
Comment by coffeebeqn 2 days ago
Comment by 0gs 2 days ago
Comment by teyc 1 day ago
Comment by 1over137 2 days ago
Comment by chrisjj 2 days ago
Who wants merely plausible code?
Comment by devin 2 days ago
Comment by aivengo_mk 1 day ago
Comment by jdw64 2 days ago
Honestly, when people say AI code quality is bad, Linus himself has said it's now genuinely useful. AI is useful and writes better code than most people. Even in competitive coding, tourist lost to AI. And in the most logical field of all, mathematics, AI is churning out an enormous number of theorems.
Looking at all this, it's fair to say AI is at least at a PhD level of technical ability, and most people would admit they don't have PhD level skills. Of course, there are still many people who code better than AI. But at least when it comes to unfolding logical structures, AI has a higher chance of being more logical than humans. Within a given framework, AI constructs much more logical structures.
That's why I think the article's use of the word 'semantic' is right. It's humans who form the framework, and that's the semantic, while AI fills the empty spaces inside it. If you feed it a flawed framework, it fails.
And the fact that AI is more logical than humans is paradoxically a greater risk. Human developers can rely on tacit knowledge to make reasonable compromises even when the requirements, the framework, are sloppy. AI can't do that. If there's a logical gap in the framework humans design, AI will exploit that weakness and expand the state space into regions we can't cognitively grasp.
Programming is ultimately about how you occupy state space. The problem is that as the program grows, the cognitively inaccessible territory keeps expanding. So we distribute trust across reliable points, libraries, frameworks, and for my own code, once it exceeds tens of thousands of lines, I rely on tests and gates.
Honestly, the idea of understanding everything in a program is a purely academic claim. Once the program gets large, it's impossible. No one can know every external factor, test bug, or unexpected interaction.
The issue is that with LLMs, when the prompt input goes deeper into the semantic space, it also reaches into areas I don't understand, producing code at a depth that's untestable.
For example, I might be an expert in domain A but a beginner in domain B. If I inject expert level knowledge for domain A into the AI, the AI will try to match that level in domain B as well. That results in code I can't understand or modify, and eventually, I'm left with no choice but to replace all the code with AI generated code.
So I'm wondering what to do about this. Should I focus on gaining empirical experience in handling black boxes? Or should I stick with smaller, human written codebases?
But realistically, the current situation, where I can build bigger and touch more things, is more enjoyable to me. I think what I actually enjoyed wasn't programming itself, but the act of creating something.
Comment by chasd00 2 days ago
i think this is the real difference between the pro-ai and anti-ai crowds.
Comment by matthorse 2 days ago
Comment by 2596-ANXC 2 days ago
Comment by vips7L 2 days ago
Comment by phrones1s 2 days ago
Comment by lardosaurusrex 2 days ago
Comment by CurbStomper 2 days ago
Comment by happytoexplain 2 days ago
Comment by doug_durham 2 days ago
Comment by SoftTalker 2 days ago
Once you know what you are building and can clearly describe it, the code isn't the hard part. Or at least that's how it has always seemed to me.
Comment by mathgeek 2 days ago
Comment by mohamedkoubaa 2 days ago
Comment by stefangordon 2 days ago
I think perhaps the assumption that engineering managers should have any employees may be outdated.
I can imagine average and mediocre engineers equipped with tokens could create chaos and debt on a scale never before imaginable, so it’s easy to see how orgs who still have these employees around are struggling with the transition.
The reality is you need to get rid of them all, and replace them with the most experienced highest paid person you can find. In the near future that person will become obsolete too.
Comment by chasd00 2 days ago
to me, the latest models and harnesses turn all devs into an engineering manager with one direct report. Some devs naturally take up the role and great things happen, for others it's like trying to get a fish to ride a bicycle. This is how it was pre-genai too so i don't think there's a right or wrong answer, some will take off with the technology and some will struggle.
Comment by sdevonoes 2 days ago
Sure thing there are engineers out there that know some about all of the above (I personally do all of that on personal projects) but you still need specialists, otherwise it’s you alone with all the unknown unknowns that the llm may claim to solve, but you cannot verify
Comment by chrisjj 2 days ago
Try air traffic control?