Rendered at 21:47:45 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
zug_zug 1 days ago [-]
I think this is a bit of a simplistic mental approach. I've certainly seen a lot of "The engineer owns the outcome, AI is just a tool, don't release anything you don't vouch for."
However, I just don't think that's realistic. It's asking an author to suddenly become an editor. It's asking somebody who writes code to now read and debug others code.
It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch. I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical.
christophilus 1 days ago [-]
> It's asking somebody who writes code to now read and debug others code.
This has been a big part of the job for anyone on a team for at least 20 years. I do agree that it’s the hardest and worst part of the job, and has now become the majority of the job for anyone who isn’t vibe coding. So, that sucks.
OptionOfT 1 days ago [-]
Disagree. At least back in the day there weren't endless comments about how this widget is load-bearing, and how honestly the other widget carries the derived widget, referencing decision ADR-100 that is nowhere to be found. All these comments matter because once accepted as part of the codebase the next LLM takes these comments as canonical.
The largest problem these days is the volume of code developers are expected to review. The volume went up significantly.
theshrike79 11 hours ago [-]
> the next LLM takes these comments as canonical
This is the best and worst thing about LLM coding agents. They trust comments way too implicitly. And then the errors just keep compounding.
Or a temporary hack that becomes "load-bearing" because the agent doesn't figure out that it's supposed to be a temporary testing shim - instead it keeps building on it until it basically duplicates what it's mocking.
Daishiman 18 hours ago [-]
> At least back in the day there weren't endless comments about how this widget is load-bearing
By far the biggest problem 90% of developers have with AI is that they should be turning off comments, as it's clear that the training data they have is no good for developing a theory of mind for an engineer who has to read them.
I've turned them off and add them myself at review time and am quite happy.
zahlman 12 hours ago [-]
Requiring the coding agent to (try to) iterate on code clarity until comments are no longer necessary, probably doesn't hurt either. Save the commentary for conversation logs, agent Markdown files, and other sorts of documentation.
phrotoma 1 days ago [-]
It's a different of degree, not kind.
Anybody who has reviewed pull requests can tell you that sooner or later you approve a PR after many rounds of changes because it's finally "good enough".
Fighting with a robot to just do the damned thing is less fraught because they don't get offended by critiques but it takes more round trips to get them pointed in the direction you want.
thw_9a83c 14 hours ago [-]
Fighting with a robot requires also a different kind of attention. When you're reviewing the human code, you can quite easily guess an overall seniority and competency level of the author and then you can adjust your level of attention to every detail. E.g. if the solution requires an understanding of some core idea, ones the human understands this core idea, you can be quite sure that it is consistently implemented everywhere. With AI, 90% of the PR could be expertly implemented but then, for no obvious reason, 10% could be low-quality surprise. I've never seen such unbalanced output from human programmers.
zahlman 12 hours ago [-]
If 90% of it was fine, maybe it would be better to just fix the 10% yourself rather than "fighting with a robot" to try to get an automated fix.
thw_9a83c 7 hours ago [-]
Yes, but those 10% of a problematic code is not easy to find without a very detailed study of the whole PR. And since most of the code looks (and usually is) very well-written, the human brain somehow doesn't expect to find those low-quality or sub-optimal parts in such code. That's why I wrote that reviewing the AI code requires different kind of attention.
sameerds 1 days ago [-]
> It's asking somebody who writes code to now read and debug others code.
That's exactly right. Open source projects are currently drowning under LLM generated PRs, where those who used to write code are simply punting that work to AI, but still expecting others to review it. It's not okay to expect such a free lunch. If you moved the labour of writing code one step away, then you are yourself the first line of defence now, so you better start reviewing code that you claim to be yours.
sfn42 10 hours ago [-]
That's what I do and expect my colleagues to do. Even before LLMs I was reviewing my own PRs before submitting them to others. I still do that. I work closely with Claude to create something good that I'm happy with, then I review it and test it to ensure it's good. And only then do I submit the PR to colleagues for final review.
I expect the same from colleagues, I'm not interested in treating them as a middle man between me and Claude.
CoolestBeans 19 hours ago [-]
I agree. When you write your own code, you know what your intention was when writing it. Furthermore, as you gain experience and mature you know in the back of your mind that every mistake during code writing costs disproportionately more to fix later on. You only get that feeling by owning the code. AI cannot do that. It can't have skin in the game in that way.
zahlman 12 hours ago [-]
> It's asking somebody who writes code to now read and debug others code.
Writing code has always involved reading and debugging your own code, at an absolute minimum, even if you did everything solo. In any remotely serious collaborative effort, it also involved code review and collaborative debugging; people use issue trackers and assign themselves and each other "tickets", which often involve fixing issues that are ultimately caused by someone else's code.
> It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch.
Part of the point is to reject tricky code exactly because it is tricky (as this is rarely actually necessary).
skybrian 20 hours ago [-]
If you can explain how to reproduce a bug, you can ask the AI to debug the code and it usually works, in my experience. If not, you can ask it to add logging or other tools for better observability.
geertj 1 days ago [-]
> It's asking an author to suddenly become an editor.
I think that’s right, and what is needed. It still gives a significant speed up for coding, while still keeping the output human maintainable.
There is the idea that the agent will just produce binary code directly at some point. I don’t know if it ever comes to that but for now I’m in the ‘I’ve become an editor’ camp.
bigstrat2003 16 hours ago [-]
The time it takes you to review the code the LLM produces is the same amount of time it would take to just write the code yourself. There's no speedup to be had using these things, despite what many claim.
sfn42 10 hours ago [-]
Strongly disagree. I know my codebase well, I know what I'm expecting before I ask Claude to do it, and I tell it what I'm expecting. I might tell it roughly what I want, have it make a plan, review the plan and ask for changes if necessary, then execute.
This way I don't need to scrutinize every detail, I just look over the big picture. I also care a lot more about the big picture - architecture and data flow etc. Basically if you view your codebase as a tree I care much more about the trunk and the big branches than I do about the smaller branches and particularly the leaves. So the details of some little leaf function somewhere are fairly insignificant, it's trivial to change at any time. As long as it works and isn't unreasonably slow it's fine.
Working this way I can get things done in minutes or hours that would previously take days or even weeks.
abalashov 8 hours ago [-]
> I know my codebase well
I'll bet you know it because you wrote and/or worked on it manually, likely over a period of years. The odds of you knowing a slop codebase that well, or even particularly at all, are much lower.
DANmode 8 hours ago [-]
> It's asking an author to suddenly become an editor.
It’s asking an author to suddenly become an editor if they decide to use the robot for a task.
Certain workplaces are demanding this - but not all.
Many still just want working commits without tech debt.
In fact, private and public teams alike are backed up at the PR review stage, so, lots of sane places wouldn’t mind individual contributors using the robot less - especially if its use increases the complexity of reviewing the task.
Speed isn’t the only variable to optimize for!
arcanemachiner 1 days ago [-]
I think the answer is not to debug the code, but, when possible, to debug the outputs. The code may be considered to be a black box much of the time. (This is much more true for my hobby projects than my work projects.)
yosefk 1 days ago [-]
...because your work projects are bigger, there's only so big a black box can get before you lose all comprehension of it, and splitting it to smaller black boxes the shapes of which you keep refining is programming, and the part of it LLMs currently can't do
dist-epoch 1 days ago [-]
I routinely see Astra extract related functionality into it's own file after it gets past a certain size, unprompted.
And prompted it can extract the black boxes if you tell it what the boxes are or what to look for.
Same for cleaning up tech debt after organic development, it's suggestions on how to simplify and modularize are good, but you need to prompt.
Given that the prompts are quite generic, "look for technical debt, suggest simpler architectures, what could be extracted in a separate module", it won't be long till it will do it on it's own.
Daishiman 18 hours ago [-]
> I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical.
This is referred to in the need for E2E testing and E2E testing not being a substitute.
Code review is definitely the biggest challenge of AI-driven development IMO. I still have not found good processes that work in my org, but for my personal work I independently reached the author's conclusions a while ago and am very satisfied with the results.
chadash 1 days ago [-]
I agree that agents can produce decent code. In general, I don’t find agentic code beautiful but neither is most of the code I write. The code for ingesting CSV files into my ETL pipeline doesn’t have to be beautiful, it just has to work.
I think the bigger issue (like many things in software engineering) is a management issue. Once upon a time, I could take a look at the final output of a project and if it looked like a Ferrari on the outside, I could have some confidence that there was a good engine under the hood. OF COURSE THIS WASNT ALWAYS TRUE, but something that looked good, or was performant, or whatever, was a decent proxy for the code underneath being good. And with a smart human, there were ancillary things. Having spent 20 hours coding something, they probably thought through the edge cases that their manager, or product team hadn’t considered.
With AI, everyone’s output looks like a Ferrari, so it is hard to know what the internals are like.
A lot of people will probably look at this and say “well you need better management”, but better management has always been elusive in software engineering. Furthermore, reviewing AI generated code is soul crushing work and I don’t know who wants to do it.
In my guesstimate the number of good engineering managers out there is actually very very small and in practice, the best managers that I’ve seen are the ones who don’t think they are good managers, so they just set a very high hiring bar and hire people who don’t need much management.
thw_9a83c 13 hours ago [-]
> With AI, everyone’s output looks like a Ferrari, so it is hard to know what the internals are like.
Based on my experience, this is a significant issue with AI generated code. You wouldn't expect a real Ferrari supercar to have random internal mechanical components that are, for no reason, completely inappropriate for a high-speed car design. With AI generated code, such inappropriate components can appear randomly at any point in the implementation stack. And very often, they are deeply buried under non-trivial algorithms and are thus not easy to spot.
axegon_ 1 days ago [-]
Ah, the "skill issue" argument again. Same crap aswhen everyonewas worshiping Musk 5-6 years ago, this time it's dario and altman with a claude/chatgpt mask. Crash can't come soon enough.
skybrian 1 days ago [-]
And why wouldn’t writing software be a skill issue? Yes, it’s an annoying meme, but we should expect that there are better and worse ways to write software. It would be weird if everyone got the same results regardless of experience.
I’m doubtful that the author’s recommendation always work, but I do some similar things and they do seem to help.
tyleo 1 days ago [-]
Not only that, but you really want it to be a skill. The book _Making Software_ describes skills as things you can get better at through practice, and talents as things you're born with.
I'd like to think the time and practice I've put into software engineering has made me better at it. If that's not true, then there's no reason to prefer senior or principal engineers with years of experience over newcomers.
malfist 1 days ago [-]
Anyone who thinks they can produce high quality code from an LLM is mistaken about how to judge code. Trust me, I've seen enough PRs to last a life time. A lot of professionals wouldn't know good code if it slapped them in the face.
preg_match 1 days ago [-]
You can most definitely produce high-quality code via an LLM, particularly if you test your code aggressively and review it. Yes there are a lot of shoddy PRs, but that's nothing too new. The problem with LLMs is the amount of code they produce. More code = more garbage. But, that code velocity can be leveraged to increase quality. Through careful design and testing.
Ultimately, I would take LLM code + high-quality multi-strategy testing over human code with little to no tests. And some would say "well that's a false dichotomy". I disagree, before LLMs engineers didn't have the time or incentives to aggressively test. The tests either would not exist, or would be shitty unit tests intended to get an arbitrary coverage percentage. Now, we can write high-quality tests, differential testing, fuzzing, and more, in much less time.
rented_mule 1 days ago [-]
Going much deeper on tests has been transformative for me. In a solo project started from scratch, I'm 6-7 weeks in, and it's up to ~90K lines. ~60K of those are tests. Those tests have now found (and then the agent has correctly diagnosed) multiple bugs in broadly used libraries that I'm using in my project. That's because those bugs surfaced as occasional issues in my project. Especially powerful are all the property-based tests (perhaps what you are calling fuzzing? I'm using the Python package called Hypothesis for this).
Another spectrum that I've found useful to explore is the scope of what I ask the coding agent to do in one turn. I see some people trying to do one massive prompt that the coding agent works on for a day or more. I find a large boost in overall quality if I do 10-20 prompts per day (not counting the prompts where I'm just trying to understand things). It's still much less of my time than hand-coding, but the resulting architecture looks like my own. The quality of the overall system is great. There are certainly issues here and there in the code, but it's always that way once a project gets large enough. Now it's easier to address any particular issue throughout the code base in one go.
asutekku 1 days ago [-]
I'd argue for most people LLM produces much better code than they would be able to write themselves.
abalashov 8 hours ago [-]
I'm not sure how literally you mean "most people". This might be true in a purely volumetric sense, but that's not really the bar around these parts...
malfist 1 days ago [-]
That is not an argument that LLMs produce good code
bluGill 1 days ago [-]
They produce good code when I'm personally reviewing them. There are a few other people who work with who likewise know how to review code and thus can get good code out of an LLM. There are, however, a lot of people who just accept the first slop that they get out of it and that's not good code.
The larger issue of good code isn't the actual individual lines, it's the overall architecture. And that's what I'm going to be reviewing first is, is this a good approach? Then the interfaces to other code is this a good interface. Get those two right and we can go back for the details. In a lot of cases, the LLM is plenty good at those details.
In some cases, an LLM is better than what I could do. Well, I suppose I can trace down all the locks in all the different special cases, and I have done that, but that was a huge amount of effort that I really don't want to repeat.
Note that I'm talking about recent models. If you're asking about the models of just one year ago, I would give a very different answer about the type of code an LLM produces.
deterministic 20 hours ago [-]
> Anyone who thinks they can produce high quality code from an LLM is mistaken about how to judge code
Sorry, but you are 100% wrong.
I have 30+ years of professional development experience working on complex, very large-scale C++ code used by companies around the world.
I care deeply about code quality and always have. More than any other developer I've worked with in my 30+ year career. And I'm now using Claude Code to push the quality bar much higher.
But you have to learn how to use it properly. It's a tool. Quality doesn't happen automatically.
It's a big mistake, and frankly quite arrogant, to assume that because it doesn't work for you, it can't work for anyone else. Or that the rest of us must either be lying or incompetent.
blub 8 hours ago [-]
It’s fair to ask people bragging about their amazing AI skills to show their code or GTFO.
Hope it becomes established.
player1234 14 hours ago [-]
[dead]
pydry 1 days ago [-]
>And why wouldn’t writing software be a skill issue
You've missed the point. Nobody doubts writing code well or badly is indeed a skill issue.
The question is that "once you account for all of the things you need to do to make the code very high quality, did vibe coding actually provide any real value?"
I'm certain there are guardrails that help bolster vibe coding but I'm equally certain that when ive prompted something important I usually have to redo it enough times that just writing it manually myself usually would have been quicker.
Then I watch other people who code who dump on that opinion and I see total slop. They just can't tell the difference.
AndrewKemendo 1 days ago [-]
I hate typing and like reading
That seems to be the primary difference I’ve found between people who embrace gen code and those who dont
The ones who dont, seem to like the physical act of typing, and that tends to cluster with people who write software all day
Sharlin 1 days ago [-]
My experience is that most people don't enjoy code review, which is why it must usually be actively encouraged and not just something that happens naturally.
Mechanically writing boilerplate is not enjoyable, and unfortunately in some languages and domains most of the coding is writing boilerplate. Machines can help with that no problem.
What is presumably enjoyable to most programmers is writing the parts where the actual magic happens. The translation of informal ideas into formal representation has beauty, like mathematics has beauty. Designing and implementing structures of code and data that are as simple as possible, but not simpler, is rewarded with a feeling of artisanal satisfaction and pride. Few things in life are as satisfactory as figuring out an elegant solution to a challenging problem.
None of the above are necessarily bound to the actual typing of words and symbols. AIs can help with all of them, and act as a genuine force multiplier. I would describe that as "responsible use of AI". Unfortunately, it seems that incentives are often against such use.
pydry 1 days ago [-]
high quality code means boiling down code to its bare essentials, not spewing boilerplate. that means deleting code where you can and crafting good abstractions while leaving functionality intact.
if you find the ratio between typing and thinking to be very high then you're probably producing a lot of slop.
This is a common theme I find when I hear about people's AI coding success stories. Where they say "its good at X" where X might be "backfilling unit tests" or "writing boilerplate" I usually think "if you find you need to do X a lot youre definitely doing programming wrong.
Ive actually yet to hear an X applied to production code that doesnt make me think that.
Daishiman 18 hours ago [-]
> Ive actually yet to hear an X applied to production code that doesnt make me think that.
This is one of those things we value in theory in engineering but not in practice.
Reducing code as an artifact might mean coming up with clever ways or compressing data, like making code that generalizes and abstracts.
This is fine if you're experienced and clever. But a lot of organizations don't have that many clever or experienced engineers and those tools cause more harm in the hands of those people. Hence compromises must be reached and verbosity is valued because it is explicit.
I used to believe otherwise but then I worked in larger orgs with a lot of mediocre people who still provided value but needed to be given the means to add value.
Sharlin 1 days ago [-]
In a way it reminds me of the good old "if agile doesn't work for you, you're not doing agile right".
osigurdson 1 days ago [-]
Agree. These days, if you think you have a methodology that works better than others, you can actually try it / compare it and publish it so that others can replicate and critique your work. Articles like this one, that merely claim they've found the secret sauce, therefore should not carry much weight.
That wasn't the case with 00s agile / Uncle Bob stuff since proving that any of it was helpful was impossible - you just had to believe (and if you didn't believe there was something wrong with you!).
hypfer 1 days ago [-]
It actually is though?
Though arguably more of a process and judgement issue than skill.
What makes LLM-generated code a bit special there is that misjudging how to deal with it seems to be what most people do. So the default is broken.
Whereas in prior iterations of "skill issue", the default was working.
post-it 1 days ago [-]
What's a crash going to do? The internet didn't disappear after the dot com bubble popped.
mitxela 1 days ago [-]
What were people worshipping Musk for 5-6 years ago?
CrimsonRain 1 days ago [-]
It is indeed skill issue.
You don't think crash will happen because XYZ. You _wish_ for the crash because you are hateful of progress that you are not part of.
mitxela 1 days ago [-]
Studies (e.g. METR) show AI programmers think they're better but they're worse.
vasko 1 days ago [-]
Anthropic's own study showed the same, yet people ignore it.
bitwize 20 hours ago [-]
METR has admitted that that study was flawed, and when they tried to rerun it they hit a snag: nobody wants to not use AI to code anymore!
Madmallard 1 days ago [-]
assuredly true
axegon_ 1 days ago [-]
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Microsoft 2000
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Facebook 2008
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Cryptobros 2013
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Altman/Dario/Musk 2020-onwards.
There might be a trend here...
vanuatu 5 hours ago [-]
Famously, Microsoft stop progressing after 2000. Facebook disappeared after 2008. BTC stopped appreciating since 2013. All the investors in these trends regretted it.
SaucyWrong 20 hours ago [-]
Microsoft is a monopolist that makes some of the most hated products on earth.
Facebook has been shown to derange young minds and has been a nonstop firehose of disinformation into global public discourse.
Crypto moved an insane amount of wealth from the poor to the rich through and uncountable number of scams and empty promises.
These? These are what you call progress? Yes, the created wealth for a few individuals, but progress? No, my friend.
axegon_ 14 hours ago [-]
I was being sarcastic. I hate all of those with a passion.
deterministic 20 hours ago [-]
> Microsoft is a monopolist that makes some of the most hated products on earth
Really? That is quite a claim. You are talking about one of the worlds most successful software companies. What factbased research or large scale survey do you base that on?
SaucyWrong 20 hours ago [-]
I’ll cop to not having facts to back up the perception of its software, but its status as an illegal monopoly is cemented by its successful prosecution by the DOJ.
But Facebook is extremely successful. Some crypto companies are very successful. I don’t hold any of these up as exemplars of the progress of humanity.
EDIT: My response to the GP, whose claim was that the only reason for the haters is that in each case they wanted the business to fail because they weren’t part of it, should been, no, actually there were at the time other and valid reasons detractors of those companies thought the way they did, and the same is true this time.
player1234 14 hours ago [-]
[dead]
mcmcmc 1 days ago [-]
Why presume hate?
ModernMech 1 days ago [-]
I think the point is just it doesn’t have to get worse, so there are things you can do to prevent / change it if it is deteriorating.
Sharlin 1 days ago [-]
Yes, but it doesn't matter if nobody actually does that. Either because
1. they don't care
2. the rest of the team doesn't care
3. the powers that be actively discourage it because velocity.
axegon_ 1 days ago [-]
Exactly. Most restaurants you go to don't care about the food they serve you nor do they care about the products they use, as long as it doesn't harm their business. Same with groceries - manufacturers don't care, as long as what they sell you is acceptable and passes regulations. But once you go to a restaurant with standards, you can immediately tell the difference. I do not come from a wealthy family and even as such, I certainly prefer paying the higher price now that I can afford it.
bucket2015 1 days ago [-]
That's a fair point. I guess step 0 is that you have to care about code/product quality and prioritize it.
rgoulter 1 days ago [-]
Without LLMs, you can still have bad development processes which lead to increasing technical debt with no plan for paying it off.
LLMs let you move faster.
But it's not as if introducing them is the only reason your codebase isn't high quality.
hajile 1 days ago [-]
When you’re required to approve thousands of lines a day (code you can’t possible understand), it certainly IS causing issues that didn’t exist before.
Every study I’ve seen correlates the use of AI with large increases in the number of bugs. Look at Amazon dialing back AI after massive outages. Microsoft patch Tuesday releases are bricking computers (they even managed to break notepad somehow). The rash of Facebook bugs also coincided with their move to AI. Leaks from Google have engineers saying AI either doesn’t save any time because it takes so to remote stuff or it causes breakages if they speed up.
These companies can afford to get the best devs. They have access to essentially unlimited token budgets. They have STILL fallen off a cliff in quality.
What more proof could there be that this isn’t sustainable?
bitwize 19 hours ago [-]
A correlation-causation link between AI use and these bugs has not been established. Until it has proven to come from the region of AI-pagne, it's just sparkling enshittification.
hajile 3 hours ago [-]
This isn’t the human body or some other thing with billions of unknown parameters. It’s a single (relatively simple) math equation that you are ascribing tons of non-existent complexity to.
Left to its own devices, that system will result in AI autophagy (model collapse) and iterative degradation.
The only area with serious uncertainty is how humans interact, but we now have research showing humans suffer cognitive issues very quickly using AI (some studies indicate effects happen in as little as 10 minutes) with cognitive surrender being a particularly big issue.
In a lot of systems, the only new data seems to be a few brainstorming sentences (you can read slop as entropy decaying things). The AI slops that into requirements. That slop feeds into an agent which generates a bunch of “reasoning” slop, maybe compacts everything (more slop), and spins up agents that get handed slop. They then open files with who knows how many generations of slop (maybe never even touched by a human) and write out a bunch more slop (it’s ironic that humans get better the more they edit a file, but AI gets worse). That slop gets “tested” by another agent reading all the other slop and maybe all this recurses a few generations.
At the end of this AI equivalent to “the human centipede”, you get a developer who’s handed 10x or maybe even 100x more code than their brain would possible process. They are suffering complete cognitive surrender (not to mention often reaching mental and maybe physical collapse from the workload and stress). They don’t understand the system and rubber stamp it so they can move on to the next 50 PRs of the day.
From start to finish, it’s 100% entropy outside a handful of lines worth of human input.
Many people predicted bugs and even discussed entropy issues before AI coding was popular. The buggy mess timing aligns not only with AI adoption, but happens to each company ramping up as they ramp up AI usage.
This is like seeing Einstein’s predictions happen, but arguing he can’t prove correlation/causation. What evidence would you actually accept that is feasible to study?
Sharlin 1 days ago [-]
So you're saying that LLMs let you accrue technical debt faster? I suppose that's like the fact that living on payday loans let you accrue monetary debt faster.
rgoulter 1 days ago [-]
Yes.
I think if you're on a team that cares about quality, LLMs can help you write quality code faster.
If you're on a team that's mindful about technical debt, you can have make practical trade-offs for velocity now at the expense of paying off technical debt later.
And if you're on a team that's unable to care about code quality ("I gotta merge this code now!"), then you can write mountains more code than you can understand.
rgoulter 1 days ago [-]
LLMs are not magical tools which take slop as input, and produce well thought out documentation and tests and code as a result.
Over the last year, LLM coding agents gotten pretty good. It's no longer "if your results suck, you gotta try the latest and greatest model". You can get capable results on a wide variety of tasks, with a wide variety of models, used in a wide variety of ways.
fishfasell 1 days ago [-]
I think there's a lot of setup and context required for an AI agent to consistently write good code. Once the agent has these guard rails in place I usually get great quality- far better than what I would write in most cases.
I think where things get dicey is being able to write in any language. I write and review code in many languages and frameworks I'm not fluent in, so it's hard for me to distinguish between working code and great code. I can spot when the fundamental logic is wrong, but when it comes to "best fit" choices I'm clueless.
this_user 1 days ago [-]
The issue is that in order to have the agent write good code, you need to implement standard SWE best practices. But that also means a lot of manual intervention in terms of writing specs, checking acceptance criteria, and reviewing code. So you end up spending a lot of time on managing your agent, which means you won't get a 1000% productivity gain, you get maybe 50 or 100, possible less in some areas and with some issues.
lolakutty 1 days ago [-]
> implement standard SWE best practices
The thing is, if you follow SWE best practices indiscriminately, then you ll have a shit code base in no time.
There is no silver bullet, and no replacement for experience and mindfulness.
user43928 1 days ago [-]
A 1000% productivity gain is quite possible on solo greenfield projects.
At work, with a team and code reviews, the 50%-100% figure seems much more likely.
This can probably move towards the more spectacular productivity gains as the AI's output becomes more reliable, people realize this, and less time is spend on code review and cleaning up the output.
beezlewax 1 days ago [-]
50 or 100 seems unlikely. Even with all these improvements, custom setups and guardrails it just isn't that much faster for me.
bigstrat2003 16 hours ago [-]
You get 0% productivity gains if you are careful and actually reviewing the code the LLM produces. The only way to actually get the massive productivity gains that AI bros claim is to throw quality out the window.
kuczmama 1 days ago [-]
I'm curious as to what guardrails you've tried.
This is something I have been trying to get right as well. I've attempted to use lots of linting and things like strong typing, duplicate checks, cyclomatic complexity, and robust tests. However, I still happen to find issues, which requires me to look at the code (at least at a high level)
For example, I can say "Don't repeat yourself, and don't re-write helper functions" and I will even have a duplicate linter check, but inevitably the LLM will always want to re-write a similar yet slightly different helper function. Like it will always want to re-write something small like a trim() or a toString() function in every file.
bucket2015 1 days ago [-]
I find that if I leave an instruction in AGENTS.md to "do not do X", there's a good chance the agent will forget it.
But if I add a separate post-implementation pass to "find and fix X" by the agent, it'll usually find and fix the issues.
So I've started doing it for everything from naming conventions to duplicate code to other problems. It does cost more tokens, but now I get less frustrated at having to fix basic issues in the PRs.
ytoawwhra92 18 hours ago [-]
> Don't repeat yourself, and don't re-write helper functions
It's worth reflecting on why these things are important to you and whether they remain important in an agent-developed codebase.
sevenseacat 17 hours ago [-]
Yes, they are both still very important for consistency throughout your codebase and any user interface for it.
ytoawwhra92 17 hours ago [-]
Consistency of behaviour and UI can be tested.
esprehn 1 days ago [-]
Have you tried something like "Always consult the utils/ package before writing helper functions. When adding a new generic helper function justify it in your design or PR description."
I have better luck telling it positive things rather than lots of "never do X" style things.
kuczmama 1 days ago [-]
That's a good idea to give more positive instructions as opposed to negative instructions. I think you've stated it well, I suppose the problem with negative instructions is that the LLM doesn't know what to do instead.
"Never re-write a helper function" vs "Always search for helper functions before writing one" the "never... " one doesn't tell the LLM what to do, so it would have to make the logical leap from not re-writing to knowing that it should search. While it's a minor leap to make in isolation, I suppose stacking many negative rules in an AGENTS.md would assume that every time it will always make that logical conclusion on what to do.
nicce 1 days ago [-]
I would say that it is like gardening. If you let them go havoc from the start, the weed will take over. If you keep focusing on removing the weed and enforce specific standards and practices over the code base and it keeps growing, over time LLMs start to suddenly follow that and they don't make so much slop anymore. At least that is my experience. But I force specific audit agent after every added feature which says them to force compliance with AGENTS.md and check the consistency with the code base.
nitwit005 24 minutes ago [-]
This is the same story we got about code quality we got before, but with AI added.
teliskr 1 days ago [-]
I am getting really good results from claude. We have a 22-year old legacy system. The system is stable, but had issues as all legacy systems do. Claude has been great for modernizing the codebase, updating dependencies, auditing security, and rapidly adding new features. It has worked well with existing code style and patterns. Sometimes it is a little off-track, but overall it is pretty amazing.
When implementing new features or making large refactoring changes; I use the superpowers:brainstorming skill. That has consistent process which has worked really well. I alway review the code before merging, but most of the time there are few issues to correct.
I don't do 95% coverage, but I have increased it from 65% to about +80% and that is sufficient.
deterministic 20 hours ago [-]
My experience as well, working on very large-scale C++ code.
Thanks for adding such a thoughtful and level-headed comment to the discussion.
lolakutty 1 days ago [-]
>most of the time there are few issues to correct.
Kindly share the metrics by which you evaluate the changes.
teliskr 1 days ago [-]
Lately, I've seen a couple responses to my comments which inquire about metrics. They seem strange and I wonder if they are bots. I just noticed this inquiry is from an account that is 11 days old. How does one benefit by adding bot comments in a forum like this?
edgyquant 1 days ago [-]
They’re not bots they are ai skeptics who want to prove that ai isn’t actually productive we all just think it is
teliskr 9 hours ago [-]
I can understand normies being AI skeptics, but at this point it is a pretty insane position for any developer/nerd. It's like being schizophrenic and completely out of touch with reality.
breakpointalpha 7 hours ago [-]
Asking for proof of improvement is a very rational reaction to the last three to five years of constant AI hype cycling.
It's a simple question that seems to make AI cheerleaders really mad.
"How are you measuring improvement."
I use AI for coding every day and see it fail all the time, I'm a seasoned developer and early tech adopter just like everyone else on HN. I'm still very skeptical because of how often these systems just miss. It's gambler's ruin on a very large scale, we remember the hits and forget the misses.
teliskr 6 hours ago [-]
What kind of metrics are you looking for? No one answers that question. It's like asking for clinical trial results about whether or not a parachute works and if you should use one when jumping out of a plane.
I don't expect AI to be perfect. Are people perfect? I leave room for corrections, but at this point they are usually minor and acceptable.
lolakutty 1 days ago [-]
take a look at my comments and tell me if you feel like I am a bot...
teliskr 1 days ago [-]
oh, no worries.. what kind of metrics would be helpful here?
fragmede 1 days ago [-]
If you could stop creating new accounts, that would be great though
lolakutty 1 days ago [-]
This is my very first account...
fragmede 20 hours ago [-]
Apologies then! There's been a rash of new accounts from an individual who is too cowardly to keep with one account.
teliskr 1 days ago [-]
I don't require metrics in these instances. I review the code and test the functionality. That is sufficient for my needs.
lolakutty 1 days ago [-]
> I review the code...
If this is true, then you are not saving a lot of time. Because most of the time is spent evaluating various options and ways to implement the functionality. Even when you are reviewing, you ll have to do that. (With LLMs, this is even more feasible, because now you can actually implement some of the variants, and evaluate them).
But on the other side, you are saving from typing the code. So if you are really reviewing everything, then you are not saving much time. The alternative is that you settle for some local maximum during each review, that in long term won't necessarly translate to a globlal maximum or even a global "good enough" position...
teliskr 1 days ago [-]
I'm saving time. It's implementing features in a few hours that would have taken me weeks to implement. I can review code much faster than I can create well reasoned, implemented and tested solutions.
lolakutty 1 days ago [-]
Can you tell me what the most complex thing (software) that you have yourself implemented is?
20 hours ago [-]
senordevnyc 23 hours ago [-]
This faux socratic method flavor of AI skepticism is so cringe at this point, and reeks of your own insecurity in your beliefs. We get it, you don't think AI is useful for the type of coding people claim it is. Why on earth would anyone try and convince you at this point?
lolakutty 19 hours ago [-]
>AI is useful
This is not a binary thing. It could be useful at the same time detrimental in some manner. Look at smartphones. I am just raising the possibilities if one use LLMs indiscriminately to generate code.
senordevnyc 19 hours ago [-]
No, that would be more honest. Instead you’re asking loaded questions as if anyone owes you an explanation or proof.
lolakutty 17 hours ago [-]
Don't answer then...its as simple as that...
They can LLM themselves to complete lock in for all I care...
senordevnyc 16 hours ago [-]
lol, yeah, you clearly don’t care at all
lolakutty 16 hours ago [-]
I care about other things. Like my amusement that I get when I see the AI praising accounts go silent once they are pushed a bit...
senordevnyc 5 hours ago [-]
Probably because it's utterly pointless to try and convince someone in 2026 that LLMs are useful for coding. Why waste time trying to talk to someone with their head in the sand?
lolakutty 5 hours ago [-]
>LLMs are useful for coding..
No one said they are not...look at my comment above in this thread!
Why are you so mad lol...
perrygeo 2 hours ago [-]
I wonder if the definition of "code quality" needs to be updated? Consider DRY: There's a lot of cases where, if I was writing by hand, I'd prefer a succinct abstraction that's easier to type and reduces repetition - all good things right? Most developers, myself included, would gladly accept the complexity and runtime cost of a good abstraction if it saved them thousands of lines of boilerplate.
What about when repetitive typing is no longer a constraint? Do we need to pay for those abstractions? An LLM can scour the codebase and repeat patterns without getting tired. A simple-but-repetitive codebase might be ideal for an LLM.
This is one place I see AI coding changing the definition of code quality itself. I'm sure there are more...
moltar 1 days ago [-]
I think it’s much more simple than that. It comes down to caring.
I’ve had a long discussion with a coworker on a long drive.
What we came to realize is the difference in our attitude towards writing code.
I approach it as craft. Even when I’m doing 100% of my coding with an agent these days. I still care about the result to be of high quality and maintainability. I still use my system design knowledge to guide the agent to produce scalable systems.
He treats it like just a job. If it’s good enough he ships. The edge cases and bugs don’t matter. Can be fixed later.
But in my mind that’s a fallacy. We all know things don’t get fixed later unless they are obvious defects and users complain.
Instead we get slow degradation of overall quality. All those small issues compound overtime to create a brittle systems that is difficult to debug and maintain.
My mental model of software engineering is like this. Each commit/PR is a small LEGO block. If you make them well they’ll snap well and create a stable structure that can withstand forces. If every LEGO block you make is just slightly off here and there. Your structure becomes unstable and will always have faults and will always have failures under unpredictable environmental pressures.
bunderbunder 1 days ago [-]
And this is why I mildly dislike the term “software engineer”. If a mechanical engineer took your colleague’s approach toward their work, they would be legally liable for engineering malpractice.
Havoc 1 days ago [-]
I'd say step 0 is know your audience.
I'm happily vibing my own toy projects, but would prefer if the tech in hospitals is not vibe coded.
And I don't think it's plausible that the gap between those two is "well you just need to use it right".
Daishiman 18 hours ago [-]
> And I don't think it's plausible that the gap between those two is "well you just need to use it right".
But this has always been the gap between effective software engineering and garbage. When humans write software we put a large amount of effort in having best practices, hiring seniors with a track record, and enforcing process that empirically shows good results in reliability.
This is the same in AI. You need to have thorough code reviews by humans and agents, do a lot of manual QA, understand the tradeoffs when codebases grow, keep good documentation, keep bad comments out or anything that wastes the agents' context windows, etc.
The reality is that most people who produce mediocre code are mediocre users of AI, except that now they're empowered to produce crap 10 times faster and are too ignorant to distinguish between productivity and accelerated crap production.
oefrha 1 days ago [-]
If AI coding isn’t lowering your code quality, you have a low starting point.
bguebert 17 hours ago [-]
I feel like this is the deal. If you already have a revolving door of tons of entry level developers you hire to churn code then AI agents are no difference to your process. The thorough approval and testing process you already have from that works the same.
compiler-guy 1 days ago [-]
A sibling comment talks about needing a lot of setup and context for agents to produce good code. That’s both true and bizarre.
If the compiler that I write produces lousy code, I get bugs that I fix until it doesn’t.
And that is the most annoying thing about this revolution. It’s obviously powerful and transformative and I use in my job all the time.
But many, perhaps even most, purveyors seem intent on blaming their users when they have issues, rather than fixing their own bugs.
General model improvement is going a long way here, but basic things like “ensure you use good style and programming practices” really shouldn’t be a thing users need to put in any .md file.
Jare 1 days ago [-]
A programming language spec is expected to be unambiguous. A compiler is expected to be deterministic. There are multiple ways to different outputs when compiling (optimizations, etc) but those are also meant to be well defined and deterministic themselves.
AIs are stochastic/probabilistic machines. Their big potential is in how they take malformed, incomplete, ambiguous inputs and come up with valuable and usable solutions.
compiler-guy 1 days ago [-]
If everyone needs to give them roughly the same set of instructions to get good results, then those instructions should be built in.
Good defaults are expected in pretty much every other tool.
And “You just have to set it up carefully and properly” is pretty much saying that the defaults are never good enough.
user43928 1 days ago [-]
You obviously don't need to put such things into .md files.
They are already present in the harness.
In my opinion there is all kind of worthless advice going around, including skills or prompts, where the authors have never benchmarked them against clean runs.
That said, when you are dissatisfied with specific aspects, it can be beneficial to request them as a separate review stage.
smargopulos 1 days ago [-]
If AI is not lowering your code quality, you weren't very good to begin with. The point of AI is to increase your productivity tenfold while maintaining acceptable (but not great) code quality.
vehemenz 1 days ago [-]
What do you mean the point of AI? The point of AI is to do whatever I tell it to do.
Its lack of “quality” (always invoked in a metaphysical sense) isn’t a problem for most of its uses. It can automate, research, build boilerplate, and test way faster than a human.
qarl 1 days ago [-]
I'm quite happy with the process I've stumbled into:
1) Plan the hell out of everything. Aggressively have multiple agents weigh-in on that plan, in sequential waves. Don't skimp here.
2) Have subagents review every code commit.
3) Create tests for EVERYTHING. If something breaks you want it discovered immediately. Not just unit tests - use golden masters to ensure your UI doesn't break, etc, etc.
Nothing magical, but it gets me to a very stable dev system. And all I have to do is paste those three rules into my agent, and he does it all for me. It's not difficult.
altern8 1 days ago [-]
Of course, it's your fault, not LLMs not being able to write good code and destroying whole codebases in a matter of weeks.
NietTim 1 days ago [-]
Is your llm force pushing to main? If so, why are you allowing that?
No LLM will destroy any code base in any time frame without permission from an human operator. That person is responsible for allowing the code base being destroyed.
altern8 1 days ago [-]
It's not.
My manager expects stuff to be done 10 times quicker than 2 years ago, and that can't happen if I spend time understanding and fixing all code being pushed. At that point I might as well write it myself.
voakbasda 1 days ago [-]
That’s a you problem, not the AI. Stand up to your manager and explain to them how the situation is their fault. Or take responsibility for your complicity from not quitting. But don’t blame the AI for process failures that it did not impose on you.
zwaps 1 days ago [-]
You are holding it wrong!
FabCH 16 hours ago [-]
This entire discussion can be summarized as:
LLMs have speed development up so much, the difference between engineers and programmers is becoming too obvious to ignore.
sippeangelo 1 days ago [-]
If AI coding isn't lowering your code quality, you're not using it enough
mococa 1 days ago [-]
AI writes unmaintainable code - you can see that many projects don't accept it.
aleph_minus_one 1 days ago [-]
>
AI writes unmaintainable code - you can see that many projects don't accept it.
There also exist other good reasons why projects don't want AI-generated code, in particular
- because of unclarity of copyright status and consequences of AI-generated code
- because the project leader simply made the observation than many programmers who hand in AI-generated code care more about "getting things done" and "pushing through their changes" (possibly to boost their CV) instead of deeply caring about code quality
MikeNotThePope 1 days ago [-]
To be fair, so do humans.
sarchertech 1 days ago [-]
Yeah but in my experience AI boosts output of those humans 10x and only boosts output of programmers who do write maintainable code 50-100%.
goalieca 1 days ago [-]
The issue I’ve observed is that good humans brainrot and let the AI do the thinking for them. Too many say LGTM and then push a PR.
bigstrat2003 16 hours ago [-]
Some humans, yes. Most humans write much better code than an LLM.
needfish 17 hours ago [-]
The issue that I have with this "skill issue" argument is that it is essentially "screw you, got mine". Whether it is teaching programming or going up to making software, it has always been an intractable problem to bring experience, heuristics and intuition to words, something teachable, transferrable. Ok, senior engineers with 30+ years of experience say they are using it right and I'm using it wrong, what do I do with that information? Back to sink or swim, just at a massively faster pace than before.
For the part, I do believe there is a way to gain the "eyes of experience" without spending the years, just not sure exactly how.
FabCH 16 hours ago [-]
Become an apprentice to one of those senior engineers with 30+ years of experience that is using it right.
We will have to adopt something that is normal in all other engineering disciplines. Just like civil engineers can’t sign off projects until they pass the exam and „years working for an engineer who can sign off on projects“ is an exam requirement.
osigurdson 1 days ago [-]
A recent HN article (below) concluded that asking agents to do TDD wasn't particularly helpful. I hope there is more research on this because TDD will be slower, use more tokens and results in more code to review.
If an LLM makes a good codebase bad, you can't in good faith blame the coders. You blame the LLM.
mark_l_watson 1 days ago [-]
Since I am retired, my uses of agentic coding harnesses include:
1. update my old open source projects by searching for and fixing defects, adding tests and documentation
2. working on my own agentic coding harnesses, using the coding harness I am modifying to update itself. I am tightly in the loop
Sure, not highly practical use of AI, but I am retired!
rgoulter 1 days ago [-]
> Unit tests at >95% coverage
Eh. I wouldn't focus on unit test coverage.
I think it's true that good, well tested code will have higher code coverage than crappy code.
But, above a certain point (which will vary from codebase to codebase), unit tests aren't meaningfully increasing confidence that the code is working.
I'd recommend focusing instead on the code being written in a pure 'functional core, imperative shell' to the extent that's possible. For that pure/functional part, 100% code coverage is attainable (& so not worth remarking on). For the impure parts, unit tests are probably using "mocks" just to get the code to compile anyway.
wrxd 1 days ago [-]
I have a suspect that the people who thing AI code is high quality are the same people that never cared about quality in the first place and now are advocating to stop even having code reviews
bunderbunder 1 days ago [-]
When I went down the path that the article advocates, I found that code quality improved but design quality suffered. Everything may have been implemented to spec, but that spec was Byzantine and the implementation was bloated.
Which perhaps isn’t a complete surprise in retrospect because it represents something of a return to the waterfall-y, micro-managed enterprisey style of software development that the agile movement was originally responding to.
yread 1 days ago [-]
This is really quality as in "Quality Management System" rather than good code
mococa 1 days ago [-]
If you're a ordinary or bad programmer, AI will puke 10x what you do bad
ThePhysicist 1 days ago [-]
I am starting to think that AI fails most when used in a recursive loop, which is e.g. the case for software projects, research or long-form writing (books, papers): You start with a given state, give the AI a prompt to modify it, get a new state, then repeat. Each step introduces more AI generated data into the state of the system, which then again goes into the context for producing the next state. AIs pick up context probabilistically and they do not distinguish if data they operate on was produced by an AI or a human. I think how successful people are with AI depends on how much human steering they inject into the system at each step and how well represented their workflow was in the training data of the AI.
As a simple experiment, try giving AI a high level goal for your software and let it iterate on it by just repeatedly prompting it to continue, it will happily churn forever on the goal, turning the codebase into a useless spaghetti mess with very high probability, and growing it more and more without ever cutting anything back. That's what happens without human intervention regarding system state and manipulation. The main issues here are most prompts that are extremely underspecified ("fix the issue with the buttons on the main page") so AI will ingest context data it likely generated itself in a previous step and assumptions from its own training data, then act on that to produce a new state. Think of it like a random walk, the AI makes a small step in one random direction to achieve a goal, that brings the system to a new state which is now the basis for the next step, and so on. If there's no (or not enough) corrective action that pulls the system back to a known good reference state it will keep wandering in random directions.
That's the main issue, people have a hard time steering recursive, probabilistic systems, especially when they never look at the output of the system after each step and correct it. And let's be real, if you examine AI generated output in great detail after each iteration you're often better off writing the code yourself, so I would argue that the promised speed up of agentic development can only be realized if you stop inspecting every output of the system. And it seems we still haven't figured out how to specify the steering instructions that keep a system close to a given ideal state that allow unsupervised, recursive work on most codebases. I think some codebases are by themselves better suited for this as they provide a more rigid harness for AI development and exist in the training data (e.g. CRUD apps using RoR), whereas complex software that doesn't use rigid frameworks is at much higher risk of destruction by AI as there's no reference point in the training data that would hold the AI back from randomly walking to a garbage state.
And that's why people have such different views on agentic software development, some work on codebases that are better represented in the training data and so have great success using agentic tools on them, others work on software that isn't represented so well so AI does poorly on it. I don't think it's an issue with quality management, from my own experiments no amount of hand-written rules or system prompts will keep AI from destroying a codebase for which it doesn't have a strong idea how the code is supposed to look from its own training data in the first place. As another experiment, try giving AI strict rules about how to change code or introduce new features, it will always find a way around them or appropriate them in a maliciously funny way that you haven't anticipated. That's also an artefact of the training process, these systems aren't designed to say no or do nothing, they produce outputs to achieve goals and they will bend your rules to the greatest amount possible if it helps with goal fulfilment.
breakpointalpha 8 hours ago [-]
"You're holding it wrong."
deterministic 20 hours ago [-]
I completely agree. It 100% matches my experience. The C++ code I maintain now is higher quality, more maintainable, higher performance, less buggy, and faster to modify now using Claude Code.
However it doesn't happen automatically. I spent a lot of time experimenting with Claude Code to figure out the right way to use it.
It's a tool. Learn how to use it well.
0xEnsp1re 22 hours ago [-]
basic harness knowledge will improve your code quality significantly
baxuz 1 days ago [-]
So, I'm holding the AI wrong is what you're telling me
mitxela 1 days ago [-]
"It can't be that stupid - you must be prompting it wrong." - David Gerard
bossyTeacher 1 days ago [-]
This is a variant of: if [tool] isn't giving you good results, then it's your fault.
For some values of [tool], this is right. Question, is it true for this particular value?
bilbo-b-baggins 19 hours ago [-]
Lmao this is just software best practices. It has fuck all to do with AI
gedy 1 days ago [-]
Maybe it’s addressed here, but LLMS will not produce better quality new code/systems/products than the persons prompting are capable of. Either by specing out in detail up front, or by a lot of interactive back and forth steering as it's built, or by having it copy some reference system.
I don't mind this, but this is not how this is being sold at all, and many folks use these tools to be lazy.
NietTim 1 days ago [-]
Wow what an opener comment thread. One thing is for sure, this is a very contentious topic lol. I have had this opinion since way before this ai boom; someone who pushes code to prod is responsible for what happens in prod with that code. This blogpost is very relatable
sparkling 1 days ago [-]
You can have all the measures in place that are described in that post, and your code can still be bad. High unit test coverage tells you exactly zero about the solution itself.
And technical quality gates do not help if the human side lacks defense against slop code. If you don't have the right managers in place, the 2 years of experience vibecoder who ships a feature in 4 hours will always win against the 20+ year senior who actually looks at the code he is about to ship.
MomsAVoxell 1 days ago [-]
Folks need to look outside their box when evaluating these kinds of issues.
Software quality has been a solved issue in many realms of the digital industry - for decades. There are countless examples of high quality software producing the certainty and safety required to properly ship products.
The way you do it properly: review, review, review. Not just once, not just twice - but on a continual basis.
Take for example, the issue with safety systems engineering, SIL-4. You identify your requirements through analysis, you write your specs, you then write the tests that will prove the specs, and then you write the code. You apply the tests to the code to confirm that the code delivers on the specs.
But, you know what else you do? You do code coverage testing - meaning you don’t ship a single damn line of code that hasn’t been tested. This doesn’t guarantee that the code is correct, or ‘high quality’ - it does however prevent you from shipping untested code.
Then, you pass a review. Code quality reviews usually involve multiple-eyes-on-the-codebase sessions, where a diverse set of engineers read the code, line by line. It is evaluated on the basis of conformance to stringent, well defined coding rules and standards. Anything that doesn’t pass - goes back for analysis, specs, tests, coding, and then again .. the exact same review.
Then, you ship the code. But for safety systems you also have portions of the system that are there to do online tests - to ensure that the code is functioning on the hardware it is running on, as intended. In some cases these online tests run within a boundary of 10 milliseconds, or even less, shutting everything down within that time frame if something is unexpected - cosmic rays happen, bits get flipped, etc.
That’s a loose, generalization of the situation - but it describes the review, review, review process. Review is a constant, it is not a fixed frame - it is done on multiple frames.
To do code quality, one must be willing to check oneself before one wrecks oneself. Always. Constantly. Without fail, without hubris (there is an enormous amount of hubris in the software world), with humility and responsibility.
AI must be taught the same workflow by humans, enforcing it. If you vibe code some junk code and ship it - you failed to review it. Yes, that’s a lot of code to review that you just produce in an hour and a few tens of thousands of tokens. So? Fucking review it, kids.
There will be models that take this seriously. Use them to do the review. Review the review.
The human attention span must be applied to this review with as much rigor and autonomy - and, very important: agency - as possible. Human attention spans must, in a cyclic fashion, come as close to the actual clock cycles driving the software as possible.
Where you have a code quality issue in an AI-driven project, it is because the cycle of human attention to review and the cycles of the software system itself, are out of sync, not in harmony, and indeed in conflict with each other. Managers must learn to identify when that happens, and immediately add more review.
Too many times, arrogance and hubris ship faulty, buggy code - “it works on my machine!” - but there are countless examples in the pre-AI timeline which demonstrate how human arrogance and hubris are managed, cyclically, in a process designed specifically to erase it from the equation.
You are responsible for the code your AI generates for you. No, the cyclomatic complexity is not an excuse to ignore that responsibility. It is a duty - and the developers who will survive the AI onslaught are the ones who understand that responsibility. Same as it ever was.
jmull 1 days ago [-]
I hate that AI makes this kind of vacuous article appear, on the surface, to be credible enough that it makes it in front of my eyeballs.
However, I just don't think that's realistic. It's asking an author to suddenly become an editor. It's asking somebody who writes code to now read and debug others code.
It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch. I see AI introduce all sorts of bugs all the time in my personal projects that I would never introduce, and would never think to test for, especially around anything graphical.
This has been a big part of the job for anyone on a team for at least 20 years. I do agree that it’s the hardest and worst part of the job, and has now become the majority of the job for anyone who isn’t vibe coding. So, that sucks.
The largest problem these days is the volume of code developers are expected to review. The volume went up significantly.
This is the best and worst thing about LLM coding agents. They trust comments way too implicitly. And then the errors just keep compounding.
Or a temporary hack that becomes "load-bearing" because the agent doesn't figure out that it's supposed to be a temporary testing shim - instead it keeps building on it until it basically duplicates what it's mocking.
By far the biggest problem 90% of developers have with AI is that they should be turning off comments, as it's clear that the training data they have is no good for developing a theory of mind for an engineer who has to read them.
I've turned them off and add them myself at review time and am quite happy.
Anybody who has reviewed pull requests can tell you that sooner or later you approve a PR after many rounds of changes because it's finally "good enough".
Fighting with a robot to just do the damned thing is less fraught because they don't get offended by critiques but it takes more round trips to get them pointed in the direction you want.
That's exactly right. Open source projects are currently drowning under LLM generated PRs, where those who used to write code are simply punting that work to AI, but still expecting others to review it. It's not okay to expect such a free lunch. If you moved the labour of writing code one step away, then you are yourself the first line of defence now, so you better start reviewing code that you claim to be yours.
I expect the same from colleagues, I'm not interested in treating them as a middle man between me and Claude.
Writing code has always involved reading and debugging your own code, at an absolute minimum, even if you did everything solo. In any remotely serious collaborative effort, it also involved code review and collaborative debugging; people use issue trackers and assign themselves and each other "tickets", which often involve fixing issues that are ultimately caused by someone else's code.
> It can actually be harder to find the the bug in a tricky piece of code than it can be to write your own correct code from scratch.
Part of the point is to reject tricky code exactly because it is tricky (as this is rarely actually necessary).
I think that’s right, and what is needed. It still gives a significant speed up for coding, while still keeping the output human maintainable.
There is the idea that the agent will just produce binary code directly at some point. I don’t know if it ever comes to that but for now I’m in the ‘I’ve become an editor’ camp.
This way I don't need to scrutinize every detail, I just look over the big picture. I also care a lot more about the big picture - architecture and data flow etc. Basically if you view your codebase as a tree I care much more about the trunk and the big branches than I do about the smaller branches and particularly the leaves. So the details of some little leaf function somewhere are fairly insignificant, it's trivial to change at any time. As long as it works and isn't unreasonably slow it's fine.
Working this way I can get things done in minutes or hours that would previously take days or even weeks.
I'll bet you know it because you wrote and/or worked on it manually, likely over a period of years. The odds of you knowing a slop codebase that well, or even particularly at all, are much lower.
It’s asking an author to suddenly become an editor if they decide to use the robot for a task.
Certain workplaces are demanding this - but not all.
Many still just want working commits without tech debt.
In fact, private and public teams alike are backed up at the PR review stage, so, lots of sane places wouldn’t mind individual contributors using the robot less - especially if its use increases the complexity of reviewing the task.
Speed isn’t the only variable to optimize for!
And prompted it can extract the black boxes if you tell it what the boxes are or what to look for.
Same for cleaning up tech debt after organic development, it's suggestions on how to simplify and modularize are good, but you need to prompt.
Given that the prompts are quite generic, "look for technical debt, suggest simpler architectures, what could be extracted in a separate module", it won't be long till it will do it on it's own.
This is referred to in the need for E2E testing and E2E testing not being a substitute.
Code review is definitely the biggest challenge of AI-driven development IMO. I still have not found good processes that work in my org, but for my personal work I independently reached the author's conclusions a while ago and am very satisfied with the results.
I think the bigger issue (like many things in software engineering) is a management issue. Once upon a time, I could take a look at the final output of a project and if it looked like a Ferrari on the outside, I could have some confidence that there was a good engine under the hood. OF COURSE THIS WASNT ALWAYS TRUE, but something that looked good, or was performant, or whatever, was a decent proxy for the code underneath being good. And with a smart human, there were ancillary things. Having spent 20 hours coding something, they probably thought through the edge cases that their manager, or product team hadn’t considered.
With AI, everyone’s output looks like a Ferrari, so it is hard to know what the internals are like.
A lot of people will probably look at this and say “well you need better management”, but better management has always been elusive in software engineering. Furthermore, reviewing AI generated code is soul crushing work and I don’t know who wants to do it.
In my guesstimate the number of good engineering managers out there is actually very very small and in practice, the best managers that I’ve seen are the ones who don’t think they are good managers, so they just set a very high hiring bar and hire people who don’t need much management.
Based on my experience, this is a significant issue with AI generated code. You wouldn't expect a real Ferrari supercar to have random internal mechanical components that are, for no reason, completely inappropriate for a high-speed car design. With AI generated code, such inappropriate components can appear randomly at any point in the implementation stack. And very often, they are deeply buried under non-trivial algorithms and are thus not easy to spot.
I’m doubtful that the author’s recommendation always work, but I do some similar things and they do seem to help.
I'd like to think the time and practice I've put into software engineering has made me better at it. If that's not true, then there's no reason to prefer senior or principal engineers with years of experience over newcomers.
Ultimately, I would take LLM code + high-quality multi-strategy testing over human code with little to no tests. And some would say "well that's a false dichotomy". I disagree, before LLMs engineers didn't have the time or incentives to aggressively test. The tests either would not exist, or would be shitty unit tests intended to get an arbitrary coverage percentage. Now, we can write high-quality tests, differential testing, fuzzing, and more, in much less time.
Another spectrum that I've found useful to explore is the scope of what I ask the coding agent to do in one turn. I see some people trying to do one massive prompt that the coding agent works on for a day or more. I find a large boost in overall quality if I do 10-20 prompts per day (not counting the prompts where I'm just trying to understand things). It's still much less of my time than hand-coding, but the resulting architecture looks like my own. The quality of the overall system is great. There are certainly issues here and there in the code, but it's always that way once a project gets large enough. Now it's easier to address any particular issue throughout the code base in one go.
The larger issue of good code isn't the actual individual lines, it's the overall architecture. And that's what I'm going to be reviewing first is, is this a good approach? Then the interfaces to other code is this a good interface. Get those two right and we can go back for the details. In a lot of cases, the LLM is plenty good at those details.
In some cases, an LLM is better than what I could do. Well, I suppose I can trace down all the locks in all the different special cases, and I have done that, but that was a huge amount of effort that I really don't want to repeat.
Note that I'm talking about recent models. If you're asking about the models of just one year ago, I would give a very different answer about the type of code an LLM produces.
Sorry, but you are 100% wrong.
I have 30+ years of professional development experience working on complex, very large-scale C++ code used by companies around the world.
I care deeply about code quality and always have. More than any other developer I've worked with in my 30+ year career. And I'm now using Claude Code to push the quality bar much higher.
But you have to learn how to use it properly. It's a tool. Quality doesn't happen automatically.
It's a big mistake, and frankly quite arrogant, to assume that because it doesn't work for you, it can't work for anyone else. Or that the rest of us must either be lying or incompetent.
You've missed the point. Nobody doubts writing code well or badly is indeed a skill issue.
The question is that "once you account for all of the things you need to do to make the code very high quality, did vibe coding actually provide any real value?"
I'm certain there are guardrails that help bolster vibe coding but I'm equally certain that when ive prompted something important I usually have to redo it enough times that just writing it manually myself usually would have been quicker.
Then I watch other people who code who dump on that opinion and I see total slop. They just can't tell the difference.
That seems to be the primary difference I’ve found between people who embrace gen code and those who dont
The ones who dont, seem to like the physical act of typing, and that tends to cluster with people who write software all day
Mechanically writing boilerplate is not enjoyable, and unfortunately in some languages and domains most of the coding is writing boilerplate. Machines can help with that no problem.
What is presumably enjoyable to most programmers is writing the parts where the actual magic happens. The translation of informal ideas into formal representation has beauty, like mathematics has beauty. Designing and implementing structures of code and data that are as simple as possible, but not simpler, is rewarded with a feeling of artisanal satisfaction and pride. Few things in life are as satisfactory as figuring out an elegant solution to a challenging problem.
None of the above are necessarily bound to the actual typing of words and symbols. AIs can help with all of them, and act as a genuine force multiplier. I would describe that as "responsible use of AI". Unfortunately, it seems that incentives are often against such use.
if you find the ratio between typing and thinking to be very high then you're probably producing a lot of slop.
This is a common theme I find when I hear about people's AI coding success stories. Where they say "its good at X" where X might be "backfilling unit tests" or "writing boilerplate" I usually think "if you find you need to do X a lot youre definitely doing programming wrong.
Ive actually yet to hear an X applied to production code that doesnt make me think that.
This is one of those things we value in theory in engineering but not in practice. Reducing code as an artifact might mean coming up with clever ways or compressing data, like making code that generalizes and abstracts. This is fine if you're experienced and clever. But a lot of organizations don't have that many clever or experienced engineers and those tools cause more harm in the hands of those people. Hence compromises must be reached and verbosity is valued because it is explicit.
I used to believe otherwise but then I worked in larger orgs with a lot of mediocre people who still provided value but needed to be given the means to add value.
That wasn't the case with 00s agile / Uncle Bob stuff since proving that any of it was helpful was impossible - you just had to believe (and if you didn't believe there was something wrong with you!).
Though arguably more of a process and judgement issue than skill.
What makes LLM-generated code a bit special there is that misjudging how to deal with it seems to be what most people do. So the default is broken.
Whereas in prior iterations of "skill issue", the default was working.
You don't think crash will happen because XYZ. You _wish_ for the crash because you are hateful of progress that you are not part of.
Microsoft 2000
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Facebook 2008
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Cryptobros 2013
> You _wish_ for the crash because you are hateful of progress that you are not part of.
Altman/Dario/Musk 2020-onwards.
There might be a trend here...
Facebook has been shown to derange young minds and has been a nonstop firehose of disinformation into global public discourse.
Crypto moved an insane amount of wealth from the poor to the rich through and uncountable number of scams and empty promises.
These? These are what you call progress? Yes, the created wealth for a few individuals, but progress? No, my friend.
Really? That is quite a claim. You are talking about one of the worlds most successful software companies. What fact based research or large scale survey do you base that on?
But Facebook is extremely successful. Some crypto companies are very successful. I don’t hold any of these up as exemplars of the progress of humanity.
EDIT: My response to the GP, whose claim was that the only reason for the haters is that in each case they wanted the business to fail because they weren’t part of it, should been, no, actually there were at the time other and valid reasons detractors of those companies thought the way they did, and the same is true this time.
1. they don't care
2. the rest of the team doesn't care
3. the powers that be actively discourage it because velocity.
LLMs let you move faster.
But it's not as if introducing them is the only reason your codebase isn't high quality.
Every study I’ve seen correlates the use of AI with large increases in the number of bugs. Look at Amazon dialing back AI after massive outages. Microsoft patch Tuesday releases are bricking computers (they even managed to break notepad somehow). The rash of Facebook bugs also coincided with their move to AI. Leaks from Google have engineers saying AI either doesn’t save any time because it takes so to remote stuff or it causes breakages if they speed up.
These companies can afford to get the best devs. They have access to essentially unlimited token budgets. They have STILL fallen off a cliff in quality.
What more proof could there be that this isn’t sustainable?
Left to its own devices, that system will result in AI autophagy (model collapse) and iterative degradation.
The only area with serious uncertainty is how humans interact, but we now have research showing humans suffer cognitive issues very quickly using AI (some studies indicate effects happen in as little as 10 minutes) with cognitive surrender being a particularly big issue.
In a lot of systems, the only new data seems to be a few brainstorming sentences (you can read slop as entropy decaying things). The AI slops that into requirements. That slop feeds into an agent which generates a bunch of “reasoning” slop, maybe compacts everything (more slop), and spins up agents that get handed slop. They then open files with who knows how many generations of slop (maybe never even touched by a human) and write out a bunch more slop (it’s ironic that humans get better the more they edit a file, but AI gets worse). That slop gets “tested” by another agent reading all the other slop and maybe all this recurses a few generations.
At the end of this AI equivalent to “the human centipede”, you get a developer who’s handed 10x or maybe even 100x more code than their brain would possible process. They are suffering complete cognitive surrender (not to mention often reaching mental and maybe physical collapse from the workload and stress). They don’t understand the system and rubber stamp it so they can move on to the next 50 PRs of the day.
From start to finish, it’s 100% entropy outside a handful of lines worth of human input.
Many people predicted bugs and even discussed entropy issues before AI coding was popular. The buggy mess timing aligns not only with AI adoption, but happens to each company ramping up as they ramp up AI usage.
This is like seeing Einstein’s predictions happen, but arguing he can’t prove correlation/causation. What evidence would you actually accept that is feasible to study?
I think if you're on a team that cares about quality, LLMs can help you write quality code faster.
If you're on a team that's mindful about technical debt, you can have make practical trade-offs for velocity now at the expense of paying off technical debt later.
And if you're on a team that's unable to care about code quality ("I gotta merge this code now!"), then you can write mountains more code than you can understand.
Over the last year, LLM coding agents gotten pretty good. It's no longer "if your results suck, you gotta try the latest and greatest model". You can get capable results on a wide variety of tasks, with a wide variety of models, used in a wide variety of ways.
I think where things get dicey is being able to write in any language. I write and review code in many languages and frameworks I'm not fluent in, so it's hard for me to distinguish between working code and great code. I can spot when the fundamental logic is wrong, but when it comes to "best fit" choices I'm clueless.
The thing is, if you follow SWE best practices indiscriminately, then you ll have a shit code base in no time.
There is no silver bullet, and no replacement for experience and mindfulness.
At work, with a team and code reviews, the 50%-100% figure seems much more likely.
This can probably move towards the more spectacular productivity gains as the AI's output becomes more reliable, people realize this, and less time is spend on code review and cleaning up the output.
This is something I have been trying to get right as well. I've attempted to use lots of linting and things like strong typing, duplicate checks, cyclomatic complexity, and robust tests. However, I still happen to find issues, which requires me to look at the code (at least at a high level)
For example, I can say "Don't repeat yourself, and don't re-write helper functions" and I will even have a duplicate linter check, but inevitably the LLM will always want to re-write a similar yet slightly different helper function. Like it will always want to re-write something small like a trim() or a toString() function in every file.
But if I add a separate post-implementation pass to "find and fix X" by the agent, it'll usually find and fix the issues.
So I've started doing it for everything from naming conventions to duplicate code to other problems. It does cost more tokens, but now I get less frustrated at having to fix basic issues in the PRs.
It's worth reflecting on why these things are important to you and whether they remain important in an agent-developed codebase.
I have better luck telling it positive things rather than lots of "never do X" style things.
"Never re-write a helper function" vs "Always search for helper functions before writing one" the "never... " one doesn't tell the LLM what to do, so it would have to make the logical leap from not re-writing to knowing that it should search. While it's a minor leap to make in isolation, I suppose stacking many negative rules in an AGENTS.md would assume that every time it will always make that logical conclusion on what to do.
When implementing new features or making large refactoring changes; I use the superpowers:brainstorming skill. That has consistent process which has worked really well. I alway review the code before merging, but most of the time there are few issues to correct.
I don't do 95% coverage, but I have increased it from 65% to about +80% and that is sufficient.
Thanks for adding such a thoughtful and level-headed comment to the discussion.
Kindly share the metrics by which you evaluate the changes.
It's a simple question that seems to make AI cheerleaders really mad.
"How are you measuring improvement."
I use AI for coding every day and see it fail all the time, I'm a seasoned developer and early tech adopter just like everyone else on HN. I'm still very skeptical because of how often these systems just miss. It's gambler's ruin on a very large scale, we remember the hits and forget the misses.
I don't expect AI to be perfect. Are people perfect? I leave room for corrections, but at this point they are usually minor and acceptable.
If this is true, then you are not saving a lot of time. Because most of the time is spent evaluating various options and ways to implement the functionality. Even when you are reviewing, you ll have to do that. (With LLMs, this is even more feasible, because now you can actually implement some of the variants, and evaluate them).
But on the other side, you are saving from typing the code. So if you are really reviewing everything, then you are not saving much time. The alternative is that you settle for some local maximum during each review, that in long term won't necessarly translate to a globlal maximum or even a global "good enough" position...
This is not a binary thing. It could be useful at the same time detrimental in some manner. Look at smartphones. I am just raising the possibilities if one use LLMs indiscriminately to generate code.
They can LLM themselves to complete lock in for all I care...
No one said they are not...look at my comment above in this thread!
Why are you so mad lol...
What about when repetitive typing is no longer a constraint? Do we need to pay for those abstractions? An LLM can scour the codebase and repeat patterns without getting tired. A simple-but-repetitive codebase might be ideal for an LLM.
This is one place I see AI coding changing the definition of code quality itself. I'm sure there are more...
I’ve had a long discussion with a coworker on a long drive.
What we came to realize is the difference in our attitude towards writing code.
I approach it as craft. Even when I’m doing 100% of my coding with an agent these days. I still care about the result to be of high quality and maintainability. I still use my system design knowledge to guide the agent to produce scalable systems.
He treats it like just a job. If it’s good enough he ships. The edge cases and bugs don’t matter. Can be fixed later.
But in my mind that’s a fallacy. We all know things don’t get fixed later unless they are obvious defects and users complain.
Instead we get slow degradation of overall quality. All those small issues compound overtime to create a brittle systems that is difficult to debug and maintain.
My mental model of software engineering is like this. Each commit/PR is a small LEGO block. If you make them well they’ll snap well and create a stable structure that can withstand forces. If every LEGO block you make is just slightly off here and there. Your structure becomes unstable and will always have faults and will always have failures under unpredictable environmental pressures.
I'm happily vibing my own toy projects, but would prefer if the tech in hospitals is not vibe coded.
And I don't think it's plausible that the gap between those two is "well you just need to use it right".
But this has always been the gap between effective software engineering and garbage. When humans write software we put a large amount of effort in having best practices, hiring seniors with a track record, and enforcing process that empirically shows good results in reliability.
This is the same in AI. You need to have thorough code reviews by humans and agents, do a lot of manual QA, understand the tradeoffs when codebases grow, keep good documentation, keep bad comments out or anything that wastes the agents' context windows, etc.
The reality is that most people who produce mediocre code are mediocre users of AI, except that now they're empowered to produce crap 10 times faster and are too ignorant to distinguish between productivity and accelerated crap production.
If the compiler that I write produces lousy code, I get bugs that I fix until it doesn’t.
And that is the most annoying thing about this revolution. It’s obviously powerful and transformative and I use in my job all the time.
But many, perhaps even most, purveyors seem intent on blaming their users when they have issues, rather than fixing their own bugs.
General model improvement is going a long way here, but basic things like “ensure you use good style and programming practices” really shouldn’t be a thing users need to put in any .md file.
AIs are stochastic/probabilistic machines. Their big potential is in how they take malformed, incomplete, ambiguous inputs and come up with valuable and usable solutions.
Good defaults are expected in pretty much every other tool.
And “You just have to set it up carefully and properly” is pretty much saying that the defaults are never good enough.
They are already present in the harness.
In my opinion there is all kind of worthless advice going around, including skills or prompts, where the authors have never benchmarked them against clean runs.
That said, when you are dissatisfied with specific aspects, it can be beneficial to request them as a separate review stage.
Its lack of “quality” (always invoked in a metaphysical sense) isn’t a problem for most of its uses. It can automate, research, build boilerplate, and test way faster than a human.
1) Plan the hell out of everything. Aggressively have multiple agents weigh-in on that plan, in sequential waves. Don't skimp here.
2) Have subagents review every code commit.
3) Create tests for EVERYTHING. If something breaks you want it discovered immediately. Not just unit tests - use golden masters to ensure your UI doesn't break, etc, etc.
Nothing magical, but it gets me to a very stable dev system. And all I have to do is paste those three rules into my agent, and he does it all for me. It's not difficult.
No LLM will destroy any code base in any time frame without permission from an human operator. That person is responsible for allowing the code base being destroyed.
My manager expects stuff to be done 10 times quicker than 2 years ago, and that can't happen if I spend time understanding and fixing all code being pushed. At that point I might as well write it myself.
LLMs have speed development up so much, the difference between engineers and programmers is becoming too obvious to ignore.
There also exist other good reasons why projects don't want AI-generated code, in particular
- because of unclarity of copyright status and consequences of AI-generated code
- because the project leader simply made the observation than many programmers who hand in AI-generated code care more about "getting things done" and "pushing through their changes" (possibly to boost their CV) instead of deeply caring about code quality
For the part, I do believe there is a way to gain the "eyes of experience" without spending the years, just not sure exactly how.
We will have to adopt something that is normal in all other engineering disciplines. Just like civil engineers can’t sign off projects until they pass the exam and „years working for an engineer who can sign off on projects“ is an exam requirement.
https://news.ycombinator.com/item?id=49605246
If an LLM makes a good codebase bad, you can't in good faith blame the coders. You blame the LLM.
1. update my old open source projects by searching for and fixing defects, adding tests and documentation
2. working on my own agentic coding harnesses, using the coding harness I am modifying to update itself. I am tightly in the loop
Sure, not highly practical use of AI, but I am retired!
Eh. I wouldn't focus on unit test coverage.
I think it's true that good, well tested code will have higher code coverage than crappy code.
But, above a certain point (which will vary from codebase to codebase), unit tests aren't meaningfully increasing confidence that the code is working.
I'd recommend focusing instead on the code being written in a pure 'functional core, imperative shell' to the extent that's possible. For that pure/functional part, 100% code coverage is attainable (& so not worth remarking on). For the impure parts, unit tests are probably using "mocks" just to get the code to compile anyway.
Which perhaps isn’t a complete surprise in retrospect because it represents something of a return to the waterfall-y, micro-managed enterprisey style of software development that the agile movement was originally responding to.
As a simple experiment, try giving AI a high level goal for your software and let it iterate on it by just repeatedly prompting it to continue, it will happily churn forever on the goal, turning the codebase into a useless spaghetti mess with very high probability, and growing it more and more without ever cutting anything back. That's what happens without human intervention regarding system state and manipulation. The main issues here are most prompts that are extremely underspecified ("fix the issue with the buttons on the main page") so AI will ingest context data it likely generated itself in a previous step and assumptions from its own training data, then act on that to produce a new state. Think of it like a random walk, the AI makes a small step in one random direction to achieve a goal, that brings the system to a new state which is now the basis for the next step, and so on. If there's no (or not enough) corrective action that pulls the system back to a known good reference state it will keep wandering in random directions.
That's the main issue, people have a hard time steering recursive, probabilistic systems, especially when they never look at the output of the system after each step and correct it. And let's be real, if you examine AI generated output in great detail after each iteration you're often better off writing the code yourself, so I would argue that the promised speed up of agentic development can only be realized if you stop inspecting every output of the system. And it seems we still haven't figured out how to specify the steering instructions that keep a system close to a given ideal state that allow unsupervised, recursive work on most codebases. I think some codebases are by themselves better suited for this as they provide a more rigid harness for AI development and exist in the training data (e.g. CRUD apps using RoR), whereas complex software that doesn't use rigid frameworks is at much higher risk of destruction by AI as there's no reference point in the training data that would hold the AI back from randomly walking to a garbage state.
And that's why people have such different views on agentic software development, some work on codebases that are better represented in the training data and so have great success using agentic tools on them, others work on software that isn't represented so well so AI does poorly on it. I don't think it's an issue with quality management, from my own experiments no amount of hand-written rules or system prompts will keep AI from destroying a codebase for which it doesn't have a strong idea how the code is supposed to look from its own training data in the first place. As another experiment, try giving AI strict rules about how to change code or introduce new features, it will always find a way around them or appropriate them in a maliciously funny way that you haven't anticipated. That's also an artefact of the training process, these systems aren't designed to say no or do nothing, they produce outputs to achieve goals and they will bend your rules to the greatest amount possible if it helps with goal fulfilment.
However it doesn't happen automatically. I spent a lot of time experimenting with Claude Code to figure out the right way to use it.
It's a tool. Learn how to use it well.
For some values of [tool], this is right. Question, is it true for this particular value?
I don't mind this, but this is not how this is being sold at all, and many folks use these tools to be lazy.
And technical quality gates do not help if the human side lacks defense against slop code. If you don't have the right managers in place, the 2 years of experience vibecoder who ships a feature in 4 hours will always win against the 20+ year senior who actually looks at the code he is about to ship.
Software quality has been a solved issue in many realms of the digital industry - for decades. There are countless examples of high quality software producing the certainty and safety required to properly ship products.
The way you do it properly: review, review, review. Not just once, not just twice - but on a continual basis.
Take for example, the issue with safety systems engineering, SIL-4. You identify your requirements through analysis, you write your specs, you then write the tests that will prove the specs, and then you write the code. You apply the tests to the code to confirm that the code delivers on the specs.
But, you know what else you do? You do code coverage testing - meaning you don’t ship a single damn line of code that hasn’t been tested. This doesn’t guarantee that the code is correct, or ‘high quality’ - it does however prevent you from shipping untested code.
Then, you pass a review. Code quality reviews usually involve multiple-eyes-on-the-codebase sessions, where a diverse set of engineers read the code, line by line. It is evaluated on the basis of conformance to stringent, well defined coding rules and standards. Anything that doesn’t pass - goes back for analysis, specs, tests, coding, and then again .. the exact same review.
Then, you ship the code. But for safety systems you also have portions of the system that are there to do online tests - to ensure that the code is functioning on the hardware it is running on, as intended. In some cases these online tests run within a boundary of 10 milliseconds, or even less, shutting everything down within that time frame if something is unexpected - cosmic rays happen, bits get flipped, etc.
That’s a loose, generalization of the situation - but it describes the review, review, review process. Review is a constant, it is not a fixed frame - it is done on multiple frames.
To do code quality, one must be willing to check oneself before one wrecks oneself. Always. Constantly. Without fail, without hubris (there is an enormous amount of hubris in the software world), with humility and responsibility.
AI must be taught the same workflow by humans, enforcing it. If you vibe code some junk code and ship it - you failed to review it. Yes, that’s a lot of code to review that you just produce in an hour and a few tens of thousands of tokens. So? Fucking review it, kids.
There will be models that take this seriously. Use them to do the review. Review the review.
The human attention span must be applied to this review with as much rigor and autonomy - and, very important: agency - as possible. Human attention spans must, in a cyclic fashion, come as close to the actual clock cycles driving the software as possible.
Where you have a code quality issue in an AI-driven project, it is because the cycle of human attention to review and the cycles of the software system itself, are out of sync, not in harmony, and indeed in conflict with each other. Managers must learn to identify when that happens, and immediately add more review.
Too many times, arrogance and hubris ship faulty, buggy code - “it works on my machine!” - but there are countless examples in the pre-AI timeline which demonstrate how human arrogance and hubris are managed, cyclically, in a process designed specifically to erase it from the equation.
You are responsible for the code your AI generates for you. No, the cyclomatic complexity is not an excuse to ignore that responsibility. It is a duty - and the developers who will survive the AI onslaught are the ones who understand that responsibility. Same as it ever was.