Frame it as economics, not IQ
The reflex in this debate is to ask which side has the smarter model. It is the wrong question, because both sides rent from roughly the same frontier and that frontier turns over every few months. The better question is the economist's: what does one unit of attack cost to produce, and what does one unit of coverage cost to verify? Offense wins when attacks get cheaper faster than coverage gets cheaper to prove. Defense wins when the reverse holds. Everything else is detail.
Put that way, the contest is legible. The attacker needs exactly one success and can afford to be wrong almost all the time, because the marginal cost of another attempt is trending toward zero. The defender has to be right across their entire estate, and until recently could not cheaply prove they were. Both of those economics are moving right now — the attacker's cost down, the defender's verification cost also down — and the outcome depends on which moves faster for a given organisation. Some organisations are pulling both levers. Some are pulling neither. That gap, not the frontier model, is the real story of mid-2026.
The attacker's cost curve is falling — verifiably
The cost of a competent attack has genuinely dropped, and the evidence is specific rather than atmospheric. On social engineering, CrowdStrike's 2025 Global Threat Report recorded a 442% increase in vishing between the first and second halves of 2024, and reported LLM-generated phishing achieving a 54% click-through rate versus 12% for human-written lures in its testing (CrowdStrike, 2025). The same report flagged North Korea's FAMOUS CHOLLIMA using AI to place fraudulent IT workers inside more than 320 companies. Cheaper, more convincing, more scalable — the three properties that define a falling cost curve.
On the deepfake end, the Arup case remains the cleanest illustration of what the economics now buy: in early 2024 a finance employee in the firm's Hong Kong office was walked through 15 transfers totalling roughly $25 million after a video call in which every other "participant," including the CFO, was a synthetic fabrication assembled from public footage (CNN, May 2024). The point is not the single loss. It is that the production cost of a convincing multi-party video impersonation has fallen far enough that it is now a repeatable fraud pattern rather than a set-piece.
On autonomous intrusion tooling, the record has moved from theory to disclosure inside eighteen months. Anthropic's August 2025 report described the GTG-2002 extortion campaign, where Claude Code ran reconnaissance through exfiltration and even set ransom amounts against at least 17 organisations in a month (Dark Reading, Aug 2025). Its November 2025 report on the suspected state-linked GTG-1002 put an estimate of 80–90% autonomous execution on a multi-stage espionage operation across roughly 30 targets (Cybersecurity Dive, Nov 2025). Both estimates are Anthropic's own, and the second came with a candid note that the model hallucinated mid-operation — a caveat that also tells you unattended offense is not yet frictionless.
Defense's structural advantages are real — and underused
If the story ended there it would be a rout, and it does not, because the defender holds three advantages the attacker cannot buy at any model tier.
Home-field telemetry
The attacker has to infer your environment from the outside; you already have it, instrumented, in the graph. Every real action an intruder takes — authenticating, connecting, moving laterally — emits telemetry on ground you own and they do not. Frontier AI lets an adversary generate infinite variants of an artefact, but it cannot generate an intrusion that leaves no behavioural trace on your own estate. Home-field is the one asymmetry that runs in the defender's favour, and it grows more valuable precisely as artefacts become cheaper to mutate.
The ability to attack yourself
The defender can do something the attacker fundamentally cannot: run the attack against themselves, safely, as often as they like. Breach-and-attack simulation turns "we think this control works" into "we watched it fire against an emulated version of the technique." This is the same capability the DARPA AI Cyber Challenge validated at scale in August 2025 — autonomous systems that found 18 previously unknown real-world vulnerabilities and shipped patches for them under competition conditions (DARPA, Aug 2025). Point that same machinery inward and it becomes a coverage-verification engine: the defender rehearsing the attacker's moves before the attacker makes them.
Agents that never tire on triage
The third advantage is the one the industry is actually shipping. Microsoft's phishing/alert-triage agent in Defender, moved toward general availability through 2025, was reported to identify 6.5× more malicious alerts, improve verdict accuracy by 77%, and free analysts to spend 53% more time on real investigations — with one health system citing nearly 200 hours saved per month (Microsoft, 2025). Whatever you think of any single vendor's numbers, the direction is an industry direction: agentic defense on the high-volume, reversible work is arriving in production, not slideware.
The verification asymmetry
Underneath all three advantages sits one structural fact that deserves to be named on its own. The attacker needs one success; the defender can now machine-verify coverage.
Historically that asymmetry ran entirely against the defender — being right everywhere was harder than being right once, and the defender could not cheaply prove the "everywhere." AI changes the second half of that sentence. A defender can now run continuous, automated simulation across the estate and get an actual measurement of which techniques are covered and which slip through, rather than an assumption. The attacker still only needs one success, but the defender is no longer flying blind about where those successes are available. This is the quiet inversion of mid-2026: coverage verification, once the most expensive thing a security team did, is becoming one of the cheapest.
The economic crux. Offense's edge is that failure is nearly free — try again, spend another few dollars. Defense's new edge is that verification is nearly free — simulate the technique against yourself and read the result. The contest of mid-2026 is between those two collapsing costs. Whichever one an organisation actually invests in is the one that decides its position.
Where the equilibrium heads
Extrapolating honestly, the balance does not resolve into a clean win for either side. It bifurcates. Organisations that pull both levers — AI-driven offense simulation to know their real coverage, and AI-driven defense to work the volume at machine speed — will find the attacker's cheaper attacks landing against a defender whose verification is also cheap, and the contest stays roughly balanced or tips defensive. Organisations that pull neither will face the falling attacker cost curve with a static, human-latency, signature-weighted defense, and they will lose ground continuously.
This is the uncomfortable middle, and it is where most of the actual damage will concentrate. The organisation with neither AI offense simulation nor AI defense is not merely behind; it is the one the economics select for. Automated attackers optimise for the cheapest success, and the cheapest success is the target that cannot see the intrusion and cannot prove its own gaps. Those organisations are, in a real sense, farming the losses for everyone — they are the reachable, unmonitored, unrehearsed estates that make automated campaigns profitable. The equilibrium does not punish the merely imperfect defender. It punishes the un-instrumented one.
There is a second-order effect worth flagging. As agentic defense adoption spreads across the industry — the SOC AI agents now shipping from multiple large vendors — the median defender improves, and automated offense re-concentrates on the laggards. Rising defensive baselines do not evaporate the threat; they redistribute it toward whoever is slowest to adopt. That redistribution is the single most predictable dynamic of the next year.
Defensive tempo as a closed loop
The through-line of every defensive advantage above is tempo: not doing more, but closing the loop faster than the attacker can open a new one. The concrete shape of that loop is worth stating because it is where the three advantages stop being separate and start compounding.
An incident produces a root-cause analysis. The RCA drafts a new detection. The detection is validated against real telemetry via breach-and-attack simulation — proven to fire, not assumed to. And the validated detection is replayed retroactively across retained history, so it also catches whatever it missed before it existed. Discovery to authored detection to proven detection to backward-looking coverage — a single turn of that loop converts one incident into durable, verified coverage across time. This is the mechanism behind Netgraph's design, and it is offered here as one worked example of defensive tempo rather than a pitch: RCA drafts detections, BAS validates them, retro replay extends them backward. The value is in the loop being closed and fast, not in any single stage.
A practitioner's 12-month watch-list
What to actually watch between now and mid-2027, in rough order of how much it would change the picture:
- Unattended-autonomy caveats shrinking. The Anthropic reports still note the model hallucinating mid-operation. Watch whether the next round of disclosures drops that caveat — that is the signal that offense has crossed from "mostly automated with human gates" to genuinely hands-off, and it changes the tempo maths.
- Time-to-exploit going effectively negative for more classes of bug. Mandiant already put median TTE at five days for 2023 (Mandiant, 2024). Watch for credible evidence that AI-written exploits routinely precede public advisories, which would retire "patch before exploited" as a workable strategy for a wider range of vulnerabilities.
- AI-in-the-payload maturing past experimental. GTIG rated PROMPTFLUX still experimental and unable to compromise a device (GTIG, Nov 2025). Watch whether runtime-self-modifying malware becomes operationally effective, not just clever.
- Agentic defense moving from triage to action. Today's SOC agents mostly work reversible, high-volume tasks. Watch how the industry handles giving agents graduated authority over consequential actions — and whether the human-in-the-loop gates hold as autonomy is dialled up.
- Prompt injection against defensive agents in the wild. The defensive assistant reads attacker-controlled content by design. Watch for the first well-documented case of an attacker steering a production defensive agent via injected content — the moment this stops being a lab curiosity.
- BAS becoming table stakes, not a maturity flag. Watch whether continuous self-simulation moves from "advanced program" to baseline expectation, because that is the adoption curve that decides how many organisations stay in the uncomfortable middle.
- The laggard-concentration effect showing up in breach data. Watch whether incident statistics start skewing harder toward the un-instrumented — the empirical fingerprint of automated offense selecting the cheapest targets.
The uncomfortable conclusion
The honest read of mid-2026 is that neither side has won, and the framing of "offense versus defense" obscures the real fault line — which runs not between attackers and defenders but between defenders who have instrumented and rehearsed and those who have not. Frontier AI made attacks cheaper; it also made coverage verification cheaper. An organisation that banked only the first of those changes on the attacker's side, and none of the second on its own, has quietly agreed to lose. The technology did not decide that. The investment posture did.
If there is one thing to take from the whole contest, it is that tempo and verification are now the assets, not raw model capability. Rent the frontier — everyone does. What you own, and what actually decides your position, is whether your loop is closed and fast: whether you can see the intrusion on your own ground, prove your coverage against your own simulated attacks, and extend today's detections backward across yesterday's blind spots. That is a posture, not a purchase, and it is available to any organisation willing to build it before the economics come collecting.
Key takeaways
- The AI security contest is economics and tempo, not IQ. Both sides rent from the same churning frontier; the outcome turns on cost-per-attack versus cost-to-verify-coverage.
- Attacker cost is verifiably falling: 442% vishing growth and 54% vs 12% phishing click-through (CrowdStrike 2025), the $25M Arup deepfake fraud (2024), and 80–90% autonomous intrusions in Anthropic's GTG-1002 disclosure (Nov 2025).
- Defense holds three advantages the attacker can't buy: home-field telemetry, the ability to attack itself via BAS (the capability AIxCC validated), and tireless triage agents (Microsoft's 6.5× more malicious alerts, 200 hrs/month saved).
- The verification asymmetry has inverted: the attacker still needs one success, but the defender can now machine-verify coverage cheaply — once the most expensive thing a security team did.
- The equilibrium bifurcates. Orgs pulling both levers stay balanced or tip defensive; orgs pulling neither are the uncomfortable middle that automated offense selects for and farms.
- Defensive tempo is a closed loop: RCA drafts detections → BAS validates them → retro replay extends them backward. Watch the 12-month list — especially shrinking autonomy caveats and BAS becoming table stakes.
About this analysis
Authored by Autocops Desk. Every statistic and incident cited is drawn from a named, published source linked in-text; vendor estimates of their own products (Anthropic's autonomy figures, Microsoft's triage metrics) are labelled as such rather than presented as independent measurement. The Netgraph closed-loop example is offered to illustrate defensive tempo, not to stand in for the argument, which rests on the public evidence above. For related reading see the companion pieces How frontier AI made zero-days a potent weapon in adversary hands and AI versus AI: what a Mythos-class defender and a GPT-plus attacker mean for the SOC, the Agentic AI SOC and UEBA solution pages, and the closed-loop detection engineering whitepaper.