How well do agents use test/verification techniques?
Posted by vinhnx 14 hours ago
Comments
Comment by nseskin 5 hours ago
Testing makes this especially clear: “use fuzzing” turns into generating random bytes, and “use formal verification” into proving a property that isn’t particularly useful. Formally, the technique is applied, but its actual purpose is lost.
This may also explain why different testing techniques produce similar results: the problem may be less about knowing a particular technique and more about the agent’s ability to keep its actual goal as the primary constraint.
Comment by getnormality 2 hours ago
I think the limiting factor is often the human's ability to articulate the goal and hold the agent accountable.
The working memory of technical people and the way they communicate often seems to prioritize how they're doing a task and not why they're doing it. So when they talk to you about their problem, they give you the human equivalent of modem noises and stack traces. This is such a common problem on tech support and Q&A forums like StackOverflow that it spawned its own terminology and website: https://xyproblem.info/
The agents don't respond to XY problems any better than a colleague does, and usually much worse.
An effective colleague asks why, and we need a similar relationship between the agent and its human supervisor.
Comment by orbital-decay 3 hours ago
Comment by cyanydeez 3 hours ago
The system prompt should be considered a starting point only if you want reusable intelligence and focus. It should not be a grab bag of tools and PR style guides, etc...
Grab your agent and inspect its prompt.
Comment by grohan 4 hours ago
Comment by siscia 12 hours ago
The way you test code cannot (and should not) be decoupled by the way in which you architect the code itself.
80%+ of effective testing is not in the testing framework but in the code architecture.
The author doesn't mention how the code is being architected and managed.
For what it is worth, I found that forcing agents on DI/hexagonal architecture and forcing a trivial coverage check is quite useful and produce overall good enough code with relative little effort
Comment by kqr 9 hours ago
Comment by sceptic123 7 hours ago
Comment by MrJohz 5 hours ago
> For the first eval, I tried giving agents the zstd RFC (plus errata) and telling them to implement a complete zstd decoder (agents are stuck in a container without internet access). The tests were not given to agents. For something with the surface area of zstd, it's not really reasonable to expect that the tests cover every possible case. For example, even though zstd is a fairly well-tested piece of software, I once found a data corruption bug in zstd. The test suite isn't intended to find extreme corner cases that might be lurking for years and is instead intended to check various cases that can "easily" be derived from the RFC that should work.
So basically, the agents were given the zstd RFC, told to implement a decoder, and also told to use a particular testing strategy (or not, for the control runs).
Comment by themgt 3 hours ago
It's a bold strategy, cotton.
Comment by yorwba 5 hours ago
Comment by sceptic123 2 hours ago
I don't think that's a very effective test of an agents ability.
Comment by yorwba 2 hours ago
Comment by wesselbindt 8 hours ago
I'm not sure I fully understand. In my view, it's pretty evident that tests should test behavior and not structure. Coupling to implementation details is unavoidable but should be managed, and ideally minimized, to keep the test suite robust and maintainable.
Given the choice between a test suite that monkey patches out some dependency (redis, say) on a per-test basis, and one which substitutes a fake for this dependency in some centralized composition root, which one do you prefer, and why?
Comment by siscia 5 hours ago
> tests should test behaviour and not structure.
Of what?
The whole idea of software architecture, and underlying of my messages, is to structure the code so that it is easy, but before easy, possible, to work at the right abstraction level.
Which allows to tests the behaviour of components and not their structure.
Trivial example, say your code read from a socket and manage the bytes with some CPU operations and write them to another socket. (You may recognise it is basically what a compressor do)
The way in which you manage read and write to the sockets will make dramatically simpler or much more difficult to write good tests.
If you adopt strategies like sans-io, you will see that the testing is almost trivial.
If all the logic sprawl up from the read syscall in a loop, you will notice how more challenging testing becomes.
---
To answer your question, the way I let LLMs write code is very DI (dependency injection) based.
A class never instantiates another class - all the dependencies are passed to the constructor. Including time. Including whatever DB iteraction.
The reason why I prefer this is that I can test each component at every level of abstraction. I don't have to. But I can.
My current approach is to force coverage higher than say 80% as default and then when a bug is discovered drill down to the components.
Comment by MrJohz 5 hours ago
That follows largely from what you're saying though, so I think the two of you agree, you're just coming at the same thing from different perspectives. Like you say, you want to avoid/manage coupling tests to implementation details, and typically the best way to do that is to bundle a significant, testable chunk of code into a single module, define a clear public interface to that module, and then test that interface. The internals of that module are then free to change, but the test interface should stay the same.
But the corollary of that is that you can't just design your modules independently of your tests, you need to design them so that they are testable. Which is why the previous comment links testing to code architecture/design.
Regarding dependencies like that, the best situation is where you can either:
(a) design a module so that the dependency is injected with a clear interface (not just "here's an instance of libredis.rs" but "here's a series of callbacks for saving data, those callbacks might save to Redis, but they might be in-memory only, or they might save everything to the blockchain, either way it doesn't matter").
(b) just spin up an instance of Redis and test directly against that — it should be quick enough, and there's plenty of ways to ensure that the tests don't affect each other even while accessing a shared resource.
Comment by crabbone 3 hours ago
Not OP, but I think I know what OP has in mind. There are some trivial ways in which code can be made easier to test as well as more complex (let's call them architectural) ways.
Examples of trivial conveniences include: making tested functionality accessible from the test (i.e. private class methods in Java present a problem for testing), coding against an interface rather than an implementation (this allows substituting a test implementation for a production implementation that can fish out some bugs), having internal to the code correctness checks that don't necessarily influence program execution (simply having runtime type checks, like it's done in eg. Java, compared to eg. C makes testing easier and more productive). Of course, there's more, these are just very common and easy to understand.
The architectural properties that improve the ability to test a program might be: the so-called "observability" programmed into the product, i.e. the product generates metrics data that's not part of the desired output. The product is split into modules with formally defined interfaces, which allows testing modules in isolation and lowers the number of possible combinations to test.
Improving testability isn't necessarily a good thing because it's likely to negatively impact complexity, size or speed of the program. So, depending on what's more important for any given program, the testing strategy will be different...
Comment by gregwebs 4 hours ago
How can I inspect exactly how agents were prompted? How can I reproduce this setup and try out my own AI setup? Am I missing a link to a repo somewhere?
These evals probably match how many people are using AI, but its not the full package of how we know software development needs to be done, and not how I do it with AI. The closest would be "Audit" which scored the highest when using xhigh- that actually incorporates a review cycle- something we know is the most important part of the software development process for code correctness (design/specification is not as much of an issue in this problem since the task is to write code against an existing spec). However, we don't know what instructions they have in their "Audit".
I would love to benchmark my own flow [1] if I can be given their exact problem. What it does is (assuming there is already a solid spec)
* plan with expensive model. Review the plan.
* implement with cheap model. Review for spec compliance and code quality.
* Reviews are done adversarially from the expensive model with a fresh context.
* ensure that verifications (automated or manual) are performed.
For non-trivial changes, the review and verification process almost always catch significant issues.The workflow does use TDD. I do find useless tests being written and I need to dig into this part of the workflow a lot more, so its great to see that aspect of this article. My experience writing software has taught me that code must be written to be easy to test, but not necessarily done TDD style.
Comment by jakevoytko 3 hours ago
- Are you sure that temperature and other nondeterminism isn't affecting your output?
- Are you sure you're not being routed through an A/B test at this moment?
- Are you sure there's not a bug affecting the model at this moment?
- Are you sure that you picked the right model and effort level?
- Are you sure that your result generalizes across providers?
- Are you sure that you set up the correct level of sandboxing and the agent can't e.g. look at a sister directory or git history in the current directory for answers?
- Are you sure that the agent isn't leaking answers in memory or its conversation history?
- Are you sure that tool calls aren't somehow affecting results?
- Are you sure that your results are robust, i.e. you see the same results with mild tweaks to the prompt?
- Are you comfortable keeping your blog post live when your results are invalidated next week with the next model launch?
And that's just a quick list off the top of my head.
I personally decided that it wasn't worth it, I'm glad that Dan decided to publish his. Frankly I think we could use a lot more of these "I ran these 2 techniques side by side and here's what I saw" anecdata, because most people who promote prompt techniques can't produce a single prompt they ran twice because they never actually tested it per se.
Edited to add: formatting + the word "promote"
Comment by gregwebs 2 hours ago
I am just asking for the bare minimum to actually understand what has been tested and for reproducibility of methods- the publication of the prompts. It would take a lot less time than all the guess work analysis write up and be a lot more useful.
Without the prompts the rest of your questions about reproducibility are moot.
Comment by ivanzhaowy123 9 hours ago
For example, when I use the Superpowers skill set, the agent proactively adopts TDD for every new feature. But its understanding of testing often stays superficial: if the user asks for a screen with a “Send” button, it first writes a test checking whether the button exists. The test fails, so it adds the button to make it pass.
As a result, the test suite fills up with low-value cases that check whether a property exists or a string matches exactly. The agent follows the “write a failing test, then implement the feature” workflow, but never really tests the business behavior: When should sending be allowed? What should happen after success or failure? How should duplicate submissions be handled?
The problem isn’t that agents can’t write tests. It’s that they seem prone to reducing TDD to a rigid sequence of steps, struggling to independently derive meaningful test cases from business requirements and use them to drive development.
Comment by andai 5 hours ago
Until... I inspected the code: it was just a bunch of print statements that said "test passed!" and didn't actually test anything.
That was last year, so hopefully it doesn't do that anymore.
Comment by throwatdem12311 5 hours ago
Every once in a while one of those useless tests you mention will sneak in…but I still read the code and either just remove it or tell the agent to get rid of it. It happens so rarely it’s barely an inconvenience.
But I don’t use any skills or anything like that to drive it. I just have a line in my CLAUDE.md to follow TDD best practices.
Personally I find most skills like these “superpowers” are just bullshit and don’t really help at all.
Comment by jiaosdjf 8 hours ago
What actually is the point of TDD? - If its to force you to think about edge cases early before you've started building the feature then that sounds like a human trait - If its to be living documentation then that sounds like a human trait
We're going into weird rabbit holes where we've mismatched the tool that is AI which produces extremely cheap code very quickly - with the processes that we've built for slow and expensive to write human-generated code.
Comment by pydry 5 hours ago
* A way of matching requirement use cases to tests.
* to modulate the number of tests written not only to make sure you have enough coverage but also to make sure you don't pointlessly cover the same edge cases multiple times.
* a way to cheaply provide feedback and validation on code as you are writing it.
With AI the ability to churn out useless tests has exploded (both with TDD done badly and with no TDD at all) and that has actually incurred a new type of cost we didnt have to face before.
Comment by emtel 1 hour ago
But once you got there, Claude had no trouble proving the code satisfied the spec, and I uncovered a few interesting bugs this way.
Comment by anitil 11 hours ago
Comment by ngruhn 10 hours ago
assert(CONSTANT_CONFIG == valueOfConfig)
or tests for keywords in prompts: assert(prompt.includes("repo url"))Comment by ponector 8 hours ago
I've seen in multiple projects things like assertTrue(true).
I'm sure the agent is better in testing than average enterprise developer.
Comment by zaphirplane 7 hours ago
Comment by __alexs 10 hours ago
I think they show that additional prompting for testing approaches have wide differences in error rate (worst has twice the error rate of the best) but actually no extra prompting is pretty fine and most custom prompts are worse than no prompt.
Comment by gmueckl 3 hours ago
Comment by simonwsimonw 5 hours ago
Comment by andai 5 hours ago
Also, had a funny experience where AI implemented an architectural change completely backwards. The implementation was pointless and made things worse rather than better. But I still got "all tests green" lol, because it just proved that the incorrect thing worked properly.
I noted with some amusement that formal verification wouldn't have helped there either, it would just have been an even stronger proof of the "correctness" of the thing that shouldn't exist to begin with.
Comment by Schlagbohrer 4 hours ago
Comment by movpasd 8 hours ago
The agent really, really struggled to bridge the gap between the code and the actual business rules it was meant to be modelling. It also struggled to work out which functions at which layer were appropriate to write tests for. So, its tests tended to be very brittle to changes to domain logic.
Essentially, agents have always seemed to struggle with modularity and problem decomposition. Good testing is about finding the right things to test, which means working out how to subdivide the input state into an appropriate product state, and checking each behaviour independently. IMO, this is the most difficult and complicated thing about programming, so I won't say it struggled _more_ than a human would --- but humans have the advantage of being able to sleep on it?
One minor thing I observed was that it tended to get really hung up on floating point edge cases (NaNs, infinities). Maybe floating point edge cases are over-represented in the property-testing training data, but it's essentially irrelevant for my usecase, at least as far as the business rules go.
Comment by vetronauta 12 hours ago
The article claims that the agents didn’t actually use TDD or mutation testing. While it’s possible for an agent to ignore the TDD procedure even when instructed to follow it, it can’t ignore a build failure.
Comment by lolpython 1 hour ago
Comment by kqr 9 hours ago
But I'm surprised Dan Luu didn't make a remark about this. Surely he must know the difference!
Comment by yorwba 5 hours ago
If you're wondering why he doesn't go into greater detail, several testing methodologies earlier we have: "Since just saying that agents didn't really meaningfully do the thing is repetitive, I'll make these sections short and only highlight particular curiosities."
Comment by coder-pm 5 hours ago
Does anyone else gate at zero rather than a percentage threshold?
Comment by devhunt-org 9 hours ago
Comment by kqr 9 hours ago
Comment by ngruhn 10 hours ago
try to make illegal states unrepresentable
In my experience, agents know how to do it. They just don't if it's not the default style of the language.Comment by ckvibubueu 8 hours ago
The tests that get created also tend to be higher quality than say the slop I see in python.
Though I do suspect my codebases are doing heavy lifting in terms of steering towards quality outcomes.
Comment by salmonellaeater 6 hours ago
Comment by williamdclt 4 hours ago
I'd be interested to hear more about why? because my experience with Go has been the opposite, I found it pretty bad for "making illegal states unrepresentable". In fact, its zero-values system often makes illegal states the _default_
Comment by penguin_booze 5 hours ago
Comment by filmdilemma 1 hour ago
Comment by gz09 12 hours ago
So it's understandable that the agents wont be able to go beyond proving trival things, given how much more difficult it is to write such code.
A (more) interesting experiment (to me) would be to write a high level spec manually for a non-trivial system (liveness etc.) and see if the agent can produce an implementation using guided refinements that satisfies this specification.
Comment by gr_norm 11 hours ago
Throwing tools haphazardly at the LLM and hoping they increase the correctness of its output is expectedly pretty ineffective. Good to see this borne out in the article.
Comment by lmeyerov 3 hours ago
Comment by CharlieDigital 4 hours ago
A big, big part of this space is "traceability" through the SDLC. If something goes wrong, there is a long chain of liability from the final effect (a subject with an adverse event) to the root cause (system recorded the wrong data in a clinical trial) all the ways to a 3rd party vendor. We call software "validated" if it has gone through the GAMP 5 V-model of software development and produced the necessary artifacts that correlate verified behavior to specification.
While this is way too much rigor for many scenarios, I think more folks should be familiar with the GAMP 5 V-model in the agentic era. The left side of the V defines the requirements moving from high level business requirements to low level technical requirements, the right side of the V defines the verification artifacts corresponding to the technical and business requirements (in that order). In this structure, testing covers both the functional requirements as well as the technical requirements.
I think this mental model is extremely useful and reflects the interaction model with agents: humans now focus on the requirements and less on the implementation details. It is useful to separate the business from the technical and therefore, the two types of artifacts that need to be produced to satisfy verification thresholds.
This style of structured software development (pre-agents) is very expensive. With a 2-3 weeks planning/specification phase on the front-end (since the test specification has a dependency on the approved requirements), a 4 week implementation phase (this is iterative with the test team, but informal; changes from the approved plan are documented as deviations), and then a 2-3 week formal test/verification phase. But in a post-agent world, this model feels like it is 1) more reasonable, 2) perhaps more effective, 3) more manageable.
While the teams I've now been on in startups have focused heavily on fast iteration with AI, a big downside I've noted is that it seems that we keep building the wrong things or things that are not useful because we are no longer verifying the business requirements before building. We assume the cost of building is low so we build and the iterate; testing in this reality becomes ad-hoc and full of gaps because the functional behavior of the software is no longer scoped before building; it's "vibes" based on each iteration because the iterations are cheap. The result? The teams I've been on feel like we're going nowhere in many cases, just spinning faster and producing more throwaway code (not a comment on whether this is good or bad; it may vary by domain and nature of the company).
If one is looking to build a software factory or pipeline, I think it behooves an architect to examine the GAMP 5 V-model and consider how to adapt ideas from this model to agent systems.
Comment by kqr 9 hours ago
Comment by tomrod 12 hours ago
Comment by epsteingpt 3 hours ago
Comment by ianjbutler 11 hours ago
The whole premise of the question is hilarious. They change text in an existing one-line comment and the best models in the world think, gee, maybe I'll lint everything AND run 4000 units. Maybe 5000 integration tests too, just to ensure we collide with any other work in progress. So you write the obligatory but often-ignored obvious things into agent memory or project steering markdown or periodic nudges: You must have a hypothesis when you run expensive tests, you must spot check changes first, then start with the most relevant tests only, then move outwards only as necessary to broader labels and only then suites and only then ALL suites in a widening gyre.
But like a falcon ignoring the falconer, the models want to run the everything for anything. So you sigh, you get the model to write a deterministic hook to catch the wrong invocation of the test suite, and you spend weeks refining the rules every time you hit a edge-case, and so it goes. At least you don't have to write the involved regexes by hand, and maybe one day it will be finished..
Comment by crabbone 7 hours ago
> Perhaps the limiting factor is just that knowledge of effective test techniques isn't very widespread
That linked to: https://danluu.com/testing/
But of course... the problem of testing is a lot harder than performance optimization... I'm surprised this comes as a surprise. Performance optimization has plenty of evaluation metrics by its very nature. Testing? -- I wish there was anything tangible at all... Because we have metrics for optimization, we have theories of optimization, i.e. we have a way of explaining how or what optimization should do. With testing? -- we are nowhere close to this point.
Another aspect of this disparity is that we also know how to sell performance optimizations. It's easy to write into an ad pamphlet that the version 2.0 of gobbledygook does 185% more gobbledygook than the 1.0! (The number faithfully copied from my cereal box!) With testing? -- How can you even tell the customer that the product was tested better? Swear on your life and cross your heart (twice, as opposed to the last time when you only did it once?)
In general, in the field, I've only have so far met with extreme pessimism about feasibility of "theory of testing" existing. Even though the need for testing goes without saying, the actual testing task is reserved for the least competent and there's little no no effort made to improve anything in this department as it's perceived to be a black hole in the budget: no matter how much you could spend on testing, the effect is likely to be the same.
Comment by quietraster 9 hours ago
Comment by kqr 9 hours ago
Comment by a2ff6eeb0 5 hours ago
This matches my experience, where the job of the engineer is mostly copy pasting requirements, letting the model do the thinking, and then manually testing the results.
Comment by thiagoc77 1 hour ago
Comment by simonwsimonw 5 hours ago
Comment by lmeyerov 2 hours ago
It's been interesting investing into different classic verification & testing methods for gfql (CPU/GPU graph queries on dataframes) and Louie (agentic investigation harness) over the last couple of years with different model & harness generations. The ones below are fewer but more longitudinal efforts compared to the shallower ones Dan worked through, and seem largely consistent:
Overall, my takeaway on standard software has been LLMs favor smart, guided fuzzing, natural language specifications deployed as iterative parallel adverserial AI review (e.g., context-reset subagents) after the functional prototyping. For formal methods / static analysis, cheaper/faster/lighter ones, though for the next level of quality, I can see this changing. The concolic testing world's heuristic results carried through a lot for me vs most others, which is unsurprising as they went deep on engineering ROI curves for fuzzing large code bases and specializing for different bug classes.
Highlights:
- alloy for gfql fell on its face relative to fuzzing, even with guidance. This was disappointing as it was intended as a cheap experiment to justify doing more expensive formal methods.
- prompts & skills need tuning: auto-memory doesn't transfer across harnesses by default, and one-shot auto-authored skills evals badly. More about iterating. Manual version of iterating would be editing skills when we hit new/repeat classes of bugs, though not guaranteed faithfulness: this is what we do the most. Positive experiments automating here, but not enough to invest deeper when we have to do manual anyways.
- natural language specifications are now a thing. Our significant new features now come with a security.md, policy.md, concurrency.md, etc., and important for them to be close to the code and taught to the review skill. Coding and review agents then can triangulate between tests, specs, code, and their own skills & general knowledge. However, we find we prefer not to do localized invariants in method comments as that gets verbose and drifts, and instead, do those as tests. This gets a bit into the global cost question of better code gen iterations and/or better review phases.
- staging coding vs reviewing. Overtesting early kills progress, so we stage heavier quality engineering at the end, and architectural research at the beginning.
- Lessons from the concolic execution era: staging static LLM analysis early with dynamic LLM testing later. Test amplify findings by area, kind, etc to convergence.
- We love community suites. GFQL builds against known Cypher language standards correctness conformance and benchmark suites, and Louie tools often start with community agent evals/benchmarks before we add our own specializations.
I don't know how big and deep the experiment Dan ran was. Something like formal methods is generally a big & invasive investment, and the target of each kind is often a much higher quality rating for a narrow set of properties. This complicates benchmarking setup. Imagine building entire compilers, and experimenting with different combos of methods & having those methods come in at different times & places.
I expect the AI security vulnerability analysis world to have similar findings. We end up baking that into our review harness as well without changing our overall methods above.
Comment by nathan-34 6 hours ago
Comment by luca_iaconelli 7 hours ago
Comment by sheetlite_74536 1 hour ago
Comment by Yaourt12 5 hours ago
Comment by haukebri 9 hours ago
Comment by rwissinger 4 hours ago
Comment by markking 9 hours ago
Comment by yuxinking 9 hours ago
Comment by zhoujinliang 11 hours ago
Comment by boldaxolotl 7 hours ago
Comment by runtime_lens 11 hours ago
Comment by FirstClassTree 6 hours ago
Comment by kestrelquant 13 hours ago
Comment by lucaszbuilds 8 hours ago