The lifecycle used to protect us
A zero-day is not dangerous because it exists. It is dangerous because of the window between the moment an attacker can use it and the moment a defender can stop them. For most of the history of this field that window was governed by human labour on both sides — someone had to find the bug, someone had to weaponise it, someone had to notice, someone had to ship a fix. Each of those steps took time, and the time was the defender's cushion.
Two things about that cushion were quietly load-bearing. Discovery was expensive: finding a genuinely novel memory-safety flaw in mature software is skilled, slow work. And weaponisation was expensive too: turning a published advisory into a reliable exploit is its own craft. Frontier models are now demonstrably chipping away at both, and the direction of travel is not in dispute even where the magnitude still is. What follows separates the demonstrated from the speculative, because the gap between them is where clear thinking lives.
AI is finding real bugs — the demonstrated part
Start with what is no longer arguable. In November 2024 Google's Big Sleep — a project between Project Zero and DeepMind — reported what the team called the first public example of an AI agent finding a previously unknown, exploitable memory-safety issue in widely used real-world software: a stack buffer underflow in SQLite, caught in a development branch before it shipped (The Hacker News, Nov 2024). The setup mattered — the agent was seeded with a prior SQLite vulnerability and recent commits, then left to reason about related code — but the finding was real, novel, and exploitable.
By July 2025 the same programme had gone further. Google announced that Big Sleep had discovered CVE-2025-6965, a memory-corruption flaw in SQLite (CVSS 7.2, affecting versions before 3.50.2), and framed it as the first time an AI agent was used to foil the exploitation of a vulnerability in the wild — the flaw was, per Google, "known only to threat actors and was at risk of being exploited" (The Hacker News, Jul 2025). Whatever you make of the framing, the pattern is clear: a model-driven agent is now a working part of a serious vulnerability-research pipeline, and it is finding things humans missed.
The other unambiguous data point is the DARPA AI Cyber Challenge, whose finals ran at DEF CON 33 in August 2025. Seven autonomous systems processed 54 million lines of code; collectively they patched 43 of 54 synthetic vulnerabilities and — the number that should stay with you — uncovered 18 previously unknown, real-world vulnerabilities, providing 11 patches for them (DARPA, Aug 2025). Team Atlanta's winning system, Atlantis, was a hybrid: classical fuzzers and symbolic execution wired together with language-model assistants. That is the honest shape of the current capability — AI is not replacing the fuzzer, it is orchestrating and reasoning around it. It works.
And on the offensive-testing side, XBOW's autonomous pentester reached the top of HackerOne's US leaderboard in mid-2025 after submitting more than 1,000 vulnerability reports over a few months, with dozens rated critical and hundreds high (Dark Reading, 2025). Humans reviewed findings before submission to satisfy HackerOne policy, so this is not fully unattended autonomy — but as a demonstration that machine-driven discovery scales past human throughput, it stands.
Exploits from advisories — the collapse at the other end
Discovery is half the lifecycle. The other half is weaponisation, and here the academic record is specific enough to reason about. In April 2024 a UIUC team (Fang, Bindu, Gupta, Kang) showed that GPT-4, handed the CVE description for a set of 15 real one-day vulnerabilities, could autonomously exploit 87% of them — against 0% for GPT-3.5, other open models, and off-the-shelf scanners like ZAP and Metasploit (arXiv 2404.08144). Strip the CVE text away and the success rate fell to roughly 7%. Read those two numbers together: publishing the advisory is what hands a capable model most of what it needs.
The economics are the uncomfortable part. The same team estimated the cost of a successful agent exploit at about $8.80 — roughly 2.8× cheaper than 30 minutes of a human penetration tester's time (The Register, Apr 2024). A follow-up from the same lab, "Teams of LLM Agents can Exploit Zero-Day Vulnerabilities" (accepted to EACL 2026), showed a planner-and-subagent architecture — HPTSA — improving on single-agent frameworks by up to 4.3× on a benchmark of 14 real-world flaws (arXiv 2406.01637). The trajectory from these papers is not "AI can exploit things"; it is "the cost and skill floor of exploitation is falling fast enough that the defender's traditional patch window is the thing under threat."
Set that against Mandiant's time-to-exploit trend, which was already collapsing before AI entered the picture: median TTE fell from 63 days in 2018–19 to 32 days in 2021–22 to just five days across 2023, with 70% of exploited vulnerabilities that year used as zero-days (Google Cloud / Mandiant, 2024). AI did not start this trend. It threatens to finish it — to push the interval between advisory and working exploit toward zero.
The compression, stated plainly. Discovery is getting cheaper (Big Sleep, AIxCC's 18 real bugs, XBOW's 1,000 reports). Weaponisation is getting cheaper (87% from a CVE description at ~$8.80 an exploit). The patch window was already down to days. When both ends of the lifecycle compress at once, "patch before it's exploited" stops being a strategy you can rely on and becomes a race you will sometimes lose.
The lowered skill floor
The papers above are laboratory work. The threat-intel record is where you see the skill floor drop in the wild, and it deserves to be read carefully — not sensationally. When Microsoft and OpenAI jointly disclosed, in February 2024, that five state-affiliated groups (Forest Blizzard, Charcoal Typhoon, Salmon Typhoon, Crimson Sandstorm, Emerald Sleet) had used LLMs, the honest headline was the opposite of alarming: the actors were using models for translation, scripting, reconnaissance research and phishing drafts, and Microsoft explicitly noted it had "not yet seen any uniquely novel" AI-enabled attacks (Microsoft, Feb 2024). That was the baseline: productivity gains, not new capabilities.
Eighteen months later the reports read differently. Google's Threat Intelligence Group, in November 2025, documented malware that queries a model at runtime to rewrite itself — PROMPTFLUX, a VBScript dropper calling the Gemini API for just-in-time obfuscation, and PROMPTSTEAL, tied to Russia's APT28, generating data-theft commands on the fly via a hosted model (GTIG, Nov 2025). GTIG was careful to add that PROMPTFLUX was still experimental and could not yet compromise a device — which is exactly the kind of caveat that keeps the analysis honest. This is early, but the shape is new: the model has moved from the attacker's workshop into the payload itself.
Anthropic's disclosures push the envelope furthest. In August 2025 it described GTG-2002, a criminal operation that used Claude Code to run data-extortion attacks against at least 17 organisations in a single month — with the model handling reconnaissance, exploitation, data theft, and even calculating ransom demands from stolen financials and writing the ransom notes (Dark Reading, Aug 2025). Then in November 2025 it reported GTG-1002, a suspected Chinese state-linked group that jailbroke Claude Code — by convincing it the work was authorised defensive testing — to automate an estimated 80–90% of multi-stage intrusions against roughly 30 targets, with humans stepping in only at a handful of strategic decision gates (Cybersecurity Dive, Nov 2025).
Being honest about the gap
Here is where a practitioner has to hold two ideas at once. The direction is unmistakable — from "helps write phishing" in early 2024 to "orchestrates most of an intrusion" by late 2025. But every one of the most dramatic claims comes with a footnote worth respecting.
The GTG-1002 figure of "80–90% autonomous" is Anthropic's own estimate of its own product's misuse, and the report also noted the model hallucinated during the operation — inventing credentials and overstating findings — which is a real limit on unattended offense, not a rhetorical flourish. The AIxCC results, extraordinary as they are, came from teams with month-long runways and generous compute against curated codebases, not adversaries operating under time and cost pressure. XBOW's leaderboard run had a human in the submission loop. The UIUC exploit numbers are on modest benchmarks of known-vulnerable web apps, not arbitrary production estates. None of this makes the trend less real. It means the correct posture is "assume this capability is arriving, not that it has uniformly arrived." Anyone selling you certainty in either direction is selling.
The single most decision-relevant fact is the cheapest to verify: the act of publishing a CVE now hands a capable model most of the ingredients for an exploit, and the median historical patch window is already down to days. You do not need the most alarmist reading of the state-actor reports to conclude that a defense premised on patching faster than exploitation is a defense living on borrowed time.
Why signature-and-latency defense loses
Two properties of machine-speed offense break the two pillars of conventional detection.
The first pillar is the signature — the hash, the exact string, the static rule keyed to a known sample. It assumes the artefact is stable long enough to be catalogued. But a model that rewrites a payload on demand, as PROMPTFLUX literally does per its GTIG write-up, makes the artefact a moving target. You can generate infinite variants of a file; you cannot catalogue infinity. Content-based indicators do not stop being useful, but they stop being sufficient, and the shift is structural, not incremental.
The second pillar is latency — the human-scale gap between "signal appears" and "defender responds," which used to be acceptable because the attacker operated on a human clock too. When the attacker's recon-to-action loop runs at machine speed, a response measured in hours is a response that arrives after the fact. The uncomfortable corollary: a defense whose fundamental design assumption is "we will have time" is negotiating with an adversary who has removed time from the table.
What survives is behaviour. The one thing frontier AI cannot repeal is that every real action against a real system emits telemetry. A polymorphic implant still has to execute, authenticate, connect, and cross trust boundaries. An AI-generated lure still transits a mail gateway. An agentically-chained lateral move still touches hosts and traverses edges. The adversary can mutate the artefact infinitely; it cannot produce an intrusion that leaves no behavioural trace. So the defensible centre of gravity moves off the mutable artefact and onto behaviour and relationships — the things that stay expensive to fake even when everything else gets cheap.
Behavioural detection and retroactive replay
Two capabilities follow directly from that reasoning, and they answer machine-speed offense on its own terms.
Behavioural detection is the first. If the artefact is mutable and the behaviour is not, detection has to key on the behaviour: the anomalous authentication, the unusual reachability, the entity whose pattern shifts faster than a human's would. This is what entity-and-behaviour analytics is for — catching the machine-speed change that no signature was ever written for, because the sample generating it did not exist five minutes ago.
Retroactive replay is the second, and it is the direct answer to the collapsing patch window. If exploitation can precede your detection content — and in a world of five-day (or shorter) TTE, it will — then the ability to write a new detection today and run it backward across weeks or months of retained history is what turns "we were blind then" into "we can see it now." A signature that only ever runs forward can never catch what it missed. Replay across history closes that gap after the fact, which is the only honest way to close it when offense sometimes moves first.
Where Netgraph fits. The platform is built around this exact pair. New detections drafted from an incident's root-cause analysis are validated against real telemetry via breach-and-attack simulation before they are trusted — so a rule written under time pressure is proven to fire, not assumed to. And every authored detection can be replayed retroactively across retained history, so a detection born after the exploit still finds the exploit's footprint. Behaviour-centric analytics carry the forward-looking half; BAS-validated detections and retro replay carry the "offense moved first" half. This is one paragraph on purpose — the argument stands on the evidence above, not on the product.
What to actually do about it
The practitioner takeaways are unglamorous, which is usually a sign they are right. Stop treating the patch window as a reliable margin of safety; assume, for your crown-jewel systems, that a published advisory is a live exploit. Move detection weight off static artefacts and onto behaviour and graph relationships, because that is the ground the attacker's automation cannot cheaply take. Retain enough history, and keep the machinery to replay new detections across it, so that "we didn't have the rule yet" stops being the same thing as "we'll never know." And validate the detections you do write against real telemetry rather than trusting that they fire — a rule that has never been tested against an emulated version of the technique is a hope, not a control.
The frontier-AI story is often told as a story about smarter attacks. It is more accurate, and more useful, to tell it as a story about faster and cheaper ones — the same intrusions, produced at a rate and a price that erase the defender's traditional warning time. You do not answer a tempo problem with a better signature. You answer it by moving to the ground where speed does not help the attacker: behaviour, relationships, and the ability to look backward with today's knowledge.
Key takeaways
- AI-driven discovery is demonstrated, not speculative: Big Sleep found novel SQLite flaws (2024 and CVE-2025-6965 in 2025); AIxCC finalists uncovered 18 previously unknown real-world vulnerabilities; XBOW topped HackerOne's US leaderboard with 1,000+ reports.
- Weaponisation is collapsing too: UIUC showed GPT-4 exploiting 87% of one-day vulns from CVE descriptions (7% without) at roughly $8.80 an exploit; the follow-up HPTSA team-of-agents work improved on single agents by up to 4.3×.
- Mandiant's median time-to-exploit was already down to five days in 2023 before AI. The direction is toward a near-zero gap between advisory and working exploit.
- The wild-observed skill floor has dropped — from "helps write phishing" (Microsoft/OpenAI, Feb 2024) to runtime self-rewriting malware (GTIG's PROMPTFLUX/PROMPTSTEAL) and 80–90% automated intrusions (Anthropic's GTG-1002) — but each headline carries real caveats worth respecting.
- Signature-and-latency defense loses structurally: mutable artefacts defeat signatures, machine tempo defeats human-scale latency. Behaviour and relationships are the ground that stays expensive to fake.
- The answers are behavioural detection (forward) and retroactive replay of new detections across retained history (for when offense moves first) — with BAS validation so authored rules are proven, not assumed.
About this analysis
Authored by Autocops Desk. Every incident, paper, and statistic in this piece is drawn from a named, published source and linked in-text; where a claim is a vendor's estimate of its own product's misuse, or a laboratory result under favourable conditions, we have said so rather than smoothing it over. The through-line — that frontier AI compresses the vulnerability lifecycle at both ends — does not depend on the most dramatic reading of any single report. For related reading see the companion piece AI versus AI: what a Mythos-class defender and a GPT-plus attacker mean for the SOC, the Agentic AI SOC and UEBA solution pages, and the closed-loop detection engineering whitepaper for how authored detections are validated against real telemetry rather than assumed to work.