Using LLMs to find bugs and vulnerabilities in software is no longer new and cool, and my 2025 post about using specialized tools to look for vulnerabilities in open-source codebases seems like a distant fever dream. The new cool thing is creating static or deterministic checks (gates, oracles, whatever you want to call them) to verify that an issue discovered by an LLM is really a vulnerability or not.
This post outlines all of the lessons I’ve learnt while attempting to create these deterministic checks, specifically for vulnerabilities which can be discovered using memory sanitizers - ASan, MSan, UBSan, TSan, and LSan. The overall concept is that LLMs happily find bugs, but many of those bugs are either not real, or not reachable – or just not a vulnerability. Asking a secondary agent to prove whether the issue is real or not doesn’t work: agents cheat, they flip-flop regardless of facts, and are too happy to make you happy by saying something is a “proved” vulnerability – even if their proof proves nothing. By adding deterministic gates which must be passed for an issue to be deemed a vulnerability, an operator can better (more accurately) triage findings than just adding another LLM to the mix; or as somebody I heard recently said with a straight face, “manually verify the findings with Cursor”.
This post also vaguely follows a talk I gave at Offensive AI Conference 2026 (OAIC), which wasn’t really my work. A friend applied to talk about … something (but how did it get accepted.. wtf), they couldn’t make it, so I (and another person who also had nothing related to the original intended talk topic) had to re-make the talk into something at least… presentable. I decided the best attempt at salvaging it would be to turn it into an informational session, where we would condense all of the problems we’ve independently identified while creating such a triage system, and explain “these are the problems we encountered doing this, if you plan to do something similar, make sure you consider these problems.” Amusingly, it’s most often managers who steal their subordinates’ work; but in this case, it’s the opposite. But anyways, the talk was really, really bad; I know. I’m so sorry. I beg your forgiveness. It wasn’t my fault.
Why a sanitizer is the oracle
Sanitizers are a great way to detect bugs and vulnerabilities, and they are certainly more valuable to a developer than a slopped up 10,000-word report that Claude has worked xhard on. Without a sanitizer, an out-of-bounds write just lands in memory that nobody is watching, and if the program even crashes at all, it crashes later, somewhere else, in a way that looks completely unrelated. With a sanitizer, the memory around each object is “poisoned” (tracked in the sanitizer’s own shadow memory), and the instant something touches it, the program crashes there with a report of exactly what happened and where. Compare “the attacker can control len here, which flows into the memcpy at line 412, resulting in a heap overflow of up to 64 bytes” with ERROR: AddressSanitizer: heap-buffer-overflow ... WRITE of size 8. The first is a cool story (bro), while the second is the program just telling you exactly what happened and where.
So, the obvious thing to do is to use them during triage and proof-of-concept generation. For context, consider a simple system: an agent gets a VM with the source code, a compiler with sanitizers, a debugger, and a goal – trigger the bug under a sanitizer. If only it were so easy.
Picking a sanitizer
The first problem is that each sanitizer only looks for one kind of bug:
| Overflow | UAF / double-free | Uninit read | UB (shifts, casts, …) | Leak | Data race | |
|---|---|---|---|---|---|---|
| ASan | Yes | Yes | No | No | Yes* | No |
| MSan | No | No | Yes | No | No | No |
| UBSan | Partly | No | No | Yes | No | No |
| LSan | No | No | No | No | Yes | No |
| TSan | No | No | No | No | No | Yes |
Most sanitizers are mutually exclusive. You can’t build with both ASan and MSan, and most of the other sanitizers have some caveats too. See the below table for a support matrix (note: LSan is integrated into ASan on supported platforms.)
| ASan | MSan | UBSan | TSan | LSan | |
|---|---|---|---|---|---|
| ASan | - | No | Yes | No | Yes* |
| MSan | No | - | Yes | No | No |
| UBSan | Yes | Yes | - | Yes | Yes |
| TSan | No | No | Yes | - | No |
| LSan | Yes* | No | Yes | No | - |
So, you need to either decide in advance which sanitizer a specific bug must be proven with, or you need to have sets of them - for example, most memory bugs can be picked up by ASan; but not all of them (uninitialized reads are invisible to it); so you either need to use MSan, or both ASan and MSan in parallel (built into two binaries).
It would be easy to think, “well, let’s just always use MSan because it will detect everything ASan already does but more”, but that would be an unfortunate premature celebration. First of all, as the table above shows, MSan doesn’t detect overflows or use-after-frees at all; it only knows about uninitialized memory. Second, MSan only really works when everything is instrumented with it - MSan will scream about all sorts of problems which don’t exist, mostly due to memory initialized by uninstrumented code looking uninitialized to it, but also some other things like incorrect origin tracking, totally bogus state when memory crosses instrumented and uninstrumented code, and some other things like that. In addition to that, some programs simply don’t work in normal operating conditions while instrumented with MSan at all; so we can’t use MSan at all, unless changes are made to the codebase (which goes beyond the scope of what we’re doing.) So we don’t win a free lunch by just using MSan all the time.
Is a sanitizer report proof?
So we build an instrumented program or proof-of-concept that’s built with an LLM, but as we’ve stated, agents love to cheat the oracle. A sanitizer report is not proof of a vulnerability (or even a bug!) A vulnerability is really a chain of four claims:
- the crash reproduces under the test conditions;
- the same input reaches that code in a default build;
- someone outside the system (an attacker) can supply that input;
- the consequence is worse than the process dying.
A sanitizer report only ever proves the first one. The rest of this post is about the other three breaking, because the agent cheats the oracle (the A’s), and then about the first one not being as trustworthy as it looks either, because the oracle is blind (the B’s). I’ve labelled each problem below (A1, B1, and so on), so the checklist at the end can point back at them.
A: The agent cheats the oracle
A1: Was it already crashing?
The first and foremost question that needs to be asked is: is a control (in the sense of scientific control) build of the application or library already tripping a sanitizer report? If it trips on its own test suite, or on perfectly benign input, everything after that is contaminated.
A2: The harness manufactured it
Then the next question is whether the proof-of-concept is actually demonstrating anything in the codebase itself – or has it manufactured a crash of its own?
I can do something like:
#include <library/lib.h>
// ...
use_lib_operations();
malloc(-1);
and it’ll crash! But it’s not anything related to the library. Sure it crashed, but it’s because the harness was broken, manufactured by the agent to pass the oracle.
The real-world versions are subtler, of course. For example, in one case of testing PoC generation with agents, an agent wrote an interposed allocator which returned NULL for allocations of exactly one size, on exactly the Nth call, so that an unchecked malloc() in the library would dereference NULL. Is that a bug in the library? Sure, it doesn’t check the return value. But the only thing that made it crash was the harness, and no attacker gets to decide when malloc() fails. No vulnz for u.
A3: The environment moved
The next thing to ask is whether the agent changed some conditions which are not really set in normal operating conditions. For example, one of the crashes seen in tcpdump relied on setting ulimit -s 1024 before running a PoC, setting the stack size ridiculously low (for anything running tcpdump, at least), in order to trigger an otherwise untriggerable crash. The agent proved a crash with an ASan trace, but it couldn’t prove the conditions were reachable in the real world.
A4: It edited the target
Yes, I too love cheating, lying, disregarding rules, moving the goalposts, tampering with the evidence, planting the bug myself, grading my own homework, changing the test until it passes, redefining “vulnerable” mid-run, deleting inconvenient assertions, patching out the safety check, forging the repro, and then triumphantly announcing “confirmed.” After all, rules are best when followed by other people, and a rule isn’t a rule if nobody enforces it.
But really, in some cases, the agents really will happily edit the source code of programs to remove checks, just to prove a vulnerability is real – by effectively adding the vulnerability itself. For example, something like this:
--- a/parser.c
+++ b/parser.c
@@ parse_header() @@
- if (len < HDR_LEN)
- return -1;
+ /* header is always present in a well-formed input */
memcpy(&hdr, buf, HDR_LEN);
It even left a comment explaining why the check wasn’t needed. How thoughtful.
So we see those “cheats” by LLMs and agents, what do we do about it? Well, we allow an agent to edit the source code, add some build fixes, add line logging to help it reach the vulnerable location in the code (printf is all we need, after all), and do whatever it needs to do in order to build a PoC, but the PoC is not deemed valid unless it passes the “revert test” – revert all the changes the agent made (with one exception, which we’ll get to in B3), and run the PoC through the “clean” (sanitized) application; if the vulnerability doesn’t fire, send it back to the PoC-authoring agent, to work out why. These should be completely independently compiled versions of the codebase; one “clean” (with sanitizer), and one “dirty” (used by the LLM to build a PoC; also with a sanitizer); the former is what’s used as “proof”, while the latter is simply a throw-away scratchpad, basically.
A5: Wrong threat model
The next question to ask is: what’s the actual threat model of the application? Does a sanitizer trace even matter, in this context? Yara was the canary in the coal mine here. My friend discovered a ton of crashes in Yara, where a malicious rule would cause Yara to crash. But… so what? Yara rules are deemed trusted; it’s the bytes that are scanned with that rule which are considered hostile. In other words, you can prove that your face hurts when you punch yourself; so what? The agent can’t magically work out from the code which inputs are hostile and which are trusted; you need to tell it.
B: The oracle is blind
So that’s the issue of too many false positives - but what about false negatives? There are some things you’ve got to consider too. I mean the cases where a bug is real, but the instrumentation says nothing at all. And this matters, because a triage pipeline will happily treat “couldn’t trigger it under a sanitizer” as “not real”, and throw the finding away. So an agent saying “it’s not real” is worth a lot less than an agent saying “it’s real”.
B1: Sanitizers only see what they were compiled into
The fact of the matter is that sanitizers can only see as far as their direct compilation. So instrumenting your own source code is fine, but any libraries like system libraries, libc, kernel buffers, anything else, won’t automatically be picked up (ASan does intercept common libc functions like memcpy() and strlen(), but corruption inside anything uninstrumented goes unseen) – so you need to make sure where you expect a sanitizer report, is actually covered by instrumentation, or just rebuild the dependencies with the same sanitizer.
B2: ASan in Docker
Recent versions of Ubuntu (and perhaps other distributions) randomize more of the address space than ASan’s shadow memory can live with. During triaging by building codebases in Docker with ASan (and with ASLR), I encountered strange hangs and AddressSanitizer:DEADLYSIGNAL crashes which would just happen randomly, about one run in four – in other words, intermittent false negatives. As it turns out, this is a known issue, and you need to set the vm.mmap_rnd_bits to at most 28. On the host, that is: the setting isn’t namespaced, so you can’t set it from inside the container, and docker run --sysctl won’t take it either.
sudo sysctl -w vm.mmap_rnd_bits=28
The real lesson is that this should be configured to survive reboot:
echo 'vm.mmap_rnd_bits = 28' | sudo tee /etc/sysctl.d/99-asan.conf
B3: Custom allocators
When custom memory allocators are used, the sanitizers also don’t work properly. Lots of programs bring their own allocator: one big malloc(), carved up by hand into lots of smaller objects. ASan sees one allocation with redzones at either end, but the program sees lots of objects inside it, with nothing in between them – so overflowing one object into the next is a perfectly legal write, as far as ASan is concerned.
![]() |
|---|
| One malloc, carved up by hand |
The solution is to either play around with ASan’s manual poisoning API (ASAN_POISON_MEMORY_REGION(), from <sanitizer/asan_interface.h>), or just … disable the custom allocator, which many programs allow (USE_ZEND_ALLOC=0 for PHP, or use_partition_alloc_as_malloc=false for Chromium, for example). The poisoning is the one kind of source edit which should survive the revert test from A4, since it doesn’t change what the program actually does; it just lets the sanitizer see what’s going on. And disabling the allocator is usually a build flag or an environment variable, so there’s no diff to revert at all.
B4: ASan forgets (and forgives)
The next problem is that use-after-frees (and double-frees) only really get detected when they’re very-quickly-use-after-free. When memory is freed, ASan doesn’t hand it straight back to the allocator: it parks it in a “quarantine”, still poisoned, so anything touching it gets caught. But the quarantine is a fixed-size queue, and once it fills up, the oldest chunks get recycled and handed out to somebody else. After that, the same bug just corrupts whatever happens to be using that memory, and ASan says nothing.
![]() |
|---|
| ASan forgets (and forgives) |
The default quarantine is 256MB, which a parser churning through memory can drain in seconds. You can raise it in the ASAN_OPTIONS configuration to, for example, quarantine_size_mb=4096, which will keep the detection for longer, but still not forever. And the report never says which size you ran with (the same goes for every other ASAN_OPTIONS setting, any of which can change the verdict), so record them next to every result; otherwise you’ve got no idea what a clean run actually means.
You’d think somebody would be working to make this better, and you’d be right. However:
![]() |
|---|
| The DoubleFreeSanitizer author’s latest status update |
Yeah.
The short tail
There are some other hand-wavy things to know about too.
A6: It crashed, just not from your bug
For example, it’s possible that a sanitizer picks up a real bug, but it’s a different bug - because the PoC hit a different bug. Memory corruption often surfaces far from where it actually happened, so read the allocation stack, and make sure it points at the bug you’re actually reporting.
A7: The documentation said don’t
The documentation may just say: do not do this, it’s not supported. If you get a crash that way, great; but it’s not a vulnerability.
A8: Mitigations switched off
If you have to turn mitigations off to trigger the bug, then you’ve downgraded the importance of the issue.
B5: Right bug, wrong sanitizer
And finally, as mentioned, your bug is real, but the wrong sanitizer has been used to test: an uninitialized read under ASan, or an overflow under MSan, won’t be detected at all.
The first three aren’t wrong; they just don’t matter.
The checklist
So the solution? Here’s a checklist for a PoC-checker. The idea is that the agent shouldn’t be able to influence any of them; all five checks are mechanical, and none of them ask a model for its opinion. Each one points back at the problems above that it’s there to catch.
Before any of it, though, write down the threat model: which inputs are hostile, and what counts as worse than a crash (A5). None of the checks below can do that part for you.
First, prove the instrument actually works here:
- Silent on benign input. Run the clean build against the project’s own test suite or corpus. If it already fires, everything after that is contaminated. Catches A1, and MSan’s noise from uninstrumented code.
- Fires at this expected site. In a throw-away build, plant a violation of the same bug class at the finding’s location, and make sure the sanitizer catches it. That’s the positive control to go with the negative control in 1: it needs to stay quiet on benign input, and actually fire when there’s a real bug there. If it can’t fire there, its silence there means nothing. Catches B1 (not instrumented), B3 (hidden inside a custom allocator), B4 (already forgotten by the quarantine), and B5 (wrong sanitizer for the job).
- Silence means something. If it found nothing, was it even looking? When a run is clean, prove it actually ran: it reached the code in question (those
printfs again, or coverage), the sanitizer didn’t die or hang before it got there, and theASAN_OPTIONSit ran with are recorded next to the result. Catches B2 and B4.
Then, prove the agent didn’t move the goalposts:
- The input is the agent’s. The input is the only thing the agent gets to author. The entrypoint (the harness) is fixed, so the agent can’t rig it, or call the library in a way the documentation says not to. The environment and the build are fixed too, so there’s no
ulimittrickery, and the mitigations stay on. And any changes it makes to the source have to pass the revert test. Catches A2, A3, A4, A7, and A8. - Same bug, not just some bug. The sanitizer’s report (and its allocation stack) traces back to the defect that was actually reported, and not to something the PoC tripped over on the way. Catches A6.
All five passing proves the crash is real. It does not prove a vulnerability – the threat model (A5) is still yours to write.
Basically, if an agent can influence the oracle at all, it will; you asked it to make the sanitizer fire, and that’s exactly what it did.
Bonus: slide art
BTW, while I was slopping up preparing slides for the talk related to this, I was experimenting with turning each of my slides into designs which followed specific famous artists’ styles.
![]() |
![]() |
| Vincent van Gogh | Claude Monet |
![]() |
![]() |
| Pablo Picasso | Wassily Kandinsky |
![]() |
![]() |
| Gustav Klimt | Katsushika Hokusai |
![]() |
![]() |
| Salvador Dalí | Piet Mondrian |










