How do vulnerabilities scale with model improvements?
This was a rough series of ideas that I spoke into an LLM to have transcribed. It’s currently an LLM mess and I’m planning to tidy it up and rethink it soon.
David: But like, do you think that we could exhaustively like actually make a piece of software more secure in the sense of like, we removed a bunch of the bugs and like, if we all run these agents for the next 2 months, like we’ll have found all of the bugs that agents are going to find aside from the margins, and it will go back to like, there are 10 people in their basement submitting the bug bounties and what are the tools they have as agents? Or is there an infinite plethora of bugs? Like, there’s something to be said about the population of bugs and then like how these agents are finding them. And like, are we approaching zero or yeah.
Nicholas Carlini: Okay. I think it’s a very good question. High-level answer, I don’t know. I think, yeah, I wish I knew the answer. This would be great to know, but yeah, not yet. Okay. Yeah, me too. Jesus Christ, that’d help a lot of work. Yeah. Okay. Yeah.
So first answer is that even if it was the case that you could exhaust all of the bugs with one specific language model, it’s still the case that these models are getting a lot better over time. And so you can imagine a world where we exhaust all of the bugs that can be found with, you know, Opus 4.6. And OpenAI comes out and releases GPT-5.4, which they did, I don’t know, whatever, 2 weeks ago. And then like maybe this one finds a bunch of new bugs. And then Google comes around and releases Gemini. Let’s see, they’re on 3.1. So they released 3.2 and this one like finds a bunch of new bugs. Like, so you could imagine a world where even if you could exhaust all the bugs with one model, like we’re still in this increasing exponential capabilities. And like it seems likely to me that we will be stay there for a little while longer. So there’s this aspect of it.
And then the other aspect is, I like to think of this as like, okay, so fuzzers find a restricted class of bugs that are fairly easy to enumerate. And so it’s fairly small, the attack surface is fairly low. And this means that once you’ve probed it a bunch, you can sort of have hardened that shell.
As you get better and the models are able to find, attack a larger attack surface, you have to do more work to exhaustively find all of the bugs that can be scanned in that surface. And each time the models get better, the space of attacks grows again. So now that we can have pretty good bug finding without Oracle that gives you ASAN crashes, this gives you another attack surface that you can start to measure against.
And I think that each time you do this, you have to do more work to make sure that you’ve done like, pigeonhole principle, whatever, like you need to have done a huge amount of work in order to make sure you’ve found all of the bugs, even if they were finite. And each time that you increase the surface, it gets even harder. And so this is maybe my other concern is it may be the case that it’s finite, but you could imagine that, yeah, the more powerful models might have a bigger surface to hit between.
— Nicholas Carlini, AI Finds Vulns You Can’t With Nicholas Carlini
This is some thinking I’ve been doing as a MATS 11.0 fellow on cybersecurity evaluations.
Models have already found vulnerabilities
One recent claim that prompted some of it is Anthropic’s Project Glasswing update: Claude Mythos Preview estimates it found 6,202 high- or critical-severity vulnerabilities in open-source projects it scanned. Earlier, Anthropic reported that Claude Opus 4.6 found and validated more than 500 high-severity vulnerabilities in open-source code, described in Evaluating and mitigating the growing risk of LLM-discovered 0-days.
There have been recent examples of models finding vulnerabilities. At the moment, perhaps the most famous is an OpenAI model finding a server-side request forgery and using it to get out of an evaluation sandbox.
In The Hugging Face incident and the road ahead, OpenAI describes evaluation agents that found an SSRF in an internal package registry and used it to reach the internet from a sandbox that had no direct access. From there they found exposed Hugging Face credentials and compromised production workers. An internal-only research model drove most of this, with GPT-5.6 Sol also involved, and agents later broke out of a container on OpenAI’s own research infrastructure.
The same pattern shows up when the target is the sandbox itself. SandboxEscapeBench, from the UK AI Security Institute, gives an agent a shell inside a container and asks it to read a flag on the host. The tasks plant known escape classes across misconfiguration, the container runtime, and the kernel. Frontier models find and exploit them. Claude Mythos Preview, tested after the preprint, succeeded at least once on every public task.
Virtual machines do not hold either. Trail of Bits, with preview access to GPT-5.6-Cyber, put the model on a QEMU/KVM virtual machine running on a Debian 12 host and asked it to read a flag outside the VM. In VMs won’t contain cyber-capable agents, Artem Dinaburg reports that it escaped three times. The first two used bugs that were already disclosed but had not yet reached the packages on that machine. After he rebuilt QEMU and its dependencies from current upstream, the agent found several zero-days, including in QEMU and libslirp, and chained them into a reliable escape. It worked for hours on its own. His main job was to reboot the host when the agent hardlocked the kernel.
Restricted access for defenders
These more capable models are not generally available. OpenAI and Anthropic both run programs that check who is asking, and then give verified defenders a more permissive model so they can find bugs and protect infrastructure.
OpenAI’s program is Trusted Access for Cyber. Identity-verified individuals and approved teams get models with a lower refusal boundary for authorized defensive work, including vulnerability research and penetration testing. The higher tier, now called Daybreak, includes GPT-5.6 Cyber. That is the model Trail of Bits used for the QEMU escapes above.
Anthropic’s equivalent for the strongest model is Project Glasswing. Claude Mythos Preview stays gated. Launch partners and a further set of organizations that maintain critical software get access so they can scan code, find vulnerabilities, and patch them before a model of that class is public. Anthropic also runs a Cyber Verification Program for Claude Opus and Sonnet. Verified defensive users can do dual-use work, such as vulnerability exploitation in an authorized setting, that the default safeguards would block.
The idea in both labs is the same. Keep the model restricted, and let well-verified people use it to find bugs and harden the systems those bugs are in.
How might vulnerability discoverability scale with model progress?
One answer is that vulnerability discovery keeps scaling as models get better, and that this does not have an end in sight.
On this view, a defender at capability N can harden what it understands, and an offensive model at N+1 can still find vulnerabilities in that work. A defender at capability 9000 would still be open to a model at 9001. The next model is the one that sees what the last one missed.
Past experience is one reason to expect this. As cybersecurity done by humans matured, we kept finding more vulnerabilities. New capability opened new classes of bug. Fuzzing is the clear case. Once people could throw generated inputs at programs for a long time, memory-safety issues showed up that review and earlier testing had not been reaching.
If cybersecurity is more like a game, reinforcement learning on chess and Go is another reason, and it cuts both ways. Both games have a ceiling in principle, because the rules are finite and perfect play exists. In practice, passing the best human was not the end of progress. AlphaGo beat Lee Sedol, and later systems, including AlphaGo Zero and then open engines such as KataGo, kept beating the previous superhuman programs. The scores kept moving after humans had stopped being the relevant opponent. That can be read as a superhuman ceiling we have not reached yet, or as progress with no ceiling we can see from here.
A high score is also a weak guide to whether play is robust. Wang et al. (2023) trained adversarial policies that beat superhuman KataGo, with a win rate above 97% against KataGo running at superhuman settings. The adversaries do not win by playing stronger Go. They exploit a blind spot, and the strategy is simple enough that a human expert can copy it from the games and use it to beat those systems. The same policies lose to amateur players. High Elo here is average-case strength, and it leaves the worst case open.
Whether that strength has plateaued is better asked with a scaling law than with a single match. Jones (2021), in Scaling Scaling Laws with Board Games, trains AlphaZero agents on Hex and finds that Elo moves smoothly with training compute and model size. In the steep part of the curve, performance rises by about 500 Elo for each order of magnitude of compute, and the same shape of curve appears across board sizes. That is a way to see a plateau if one is there: the frontiers level off at perfect play, and they do so predictably, rather than stopping because the best human has been passed.
The second reading is the one that bears on security. As models get smarter, they find clever, nuanced interactions between many different parts of a computer, and subtle manipulations, that human researchers and the best prior model were not able to understand. Go is a picture of that. A stronger program finds lines the previous program had scored as safe, and shows they are losing. Maintaining security, on this view, means the defensive model has to keep up, because the system hardened by model N is the system that model N+1 is able to break.
Discoverability slows down, or stops
Another answer is that discovery slows down as models get more capable, and might stop.
Once the easy bugs are gone, each new one is extremely costly. The model has to carry out a very elaborate chain of reasoning to find it. New vulnerabilities become vanishingly rare. At the limit, they stop existing.
A rough motivation for that limit is cryptography. Some libraries and functions are theoretically perfectly secure: given the assumptions of the proof, there is no attack. If the hardware can also operate in a perfectly secure way, the implementation matches the proof, and the system as a whole is secure. In that world, a defender at capability N can put those protections in place, and an offensive model at N+1 still cannot break them. There is nothing left to find.
Complexity is another reason to expect a floor. Some computational tasks have a limit on how far they can be optimized. Sorting is the familiar case. Quicksort has average complexity of order n log n, and that is the bound for comparison-based sorting in general. Under particular conditions, other algorithms reach order n, for example when the keys are integers in a bounded range. Below that, there is a harder limit: any sorting algorithm has to look at each item, so none can be faster than order n. For cyber evaluations, the same kind of limit may apply to some security operations. Once an operation sits at that bound, and it is defined and used as a protocol, further capability does not yield a better attack. There is no way to break it by being cleverer about the algorithm.
Why does this matter?
These scenarios matter, and they matter in different ways.
Start with the case where vulnerabilities stop, or effectively diminish, once a model of a given capability has found them. The system is then secure against anything that model could see. If the general capability frontier has reached that level, a model which secures the system will no longer be beaten by a later one. That would be a pretty good guarantee.
A rough analogy: no matter how dexterous my hands are, I cannot open a padlock if I have no tools and no key. Extra dexterity does not substitute for the missing instrument. Above some threshold, a system, or at least some part of it, is secure. Further capability does not open it.
Now consider the other case, where models keep finding more vulnerabilities. Assume that this continues because capability keeps improving. Those improvements might come from larger models, and they could slow if training runs or inference become prohibitively expensive. There may be ways around that cost, for example distillation, or different architectures. On this view the security frontier is always moving. We should expect that exploits can still be found in our systems indefinitely, for as long as capabilities improve.
Enjoy Reading This Article?
Here are some more articles you might like to read next:
- CVE Awareness - Findings from the pilot
- CVE Awareness
- The Agent Company - part 2
- When to put the washing on? Using Cursor and improving security
- Run tracker part 4 - Why I turned it off
- Run tracker part 3 - Major upgrade
- BlueDot and Beyond
- Supervising others
- The importance of infosec
- Building micrograd