Disclosure: These views are my own and do not represent my current or any former employers. This paper uses only public sources and does not describe non-public product information.
Executive summary
Most defensive engineering carries an assumption that is rarely written down: an attacker who fails will have to spend scarce human effort before trying again. Blocking a payload forces retooling. Rotating a credential forces another intrusion. Changing a detection lengthens the attacker's development cycle. Obscure configuration slows down anyone who has to read it. Each works partly because somebody on the other side has to do work, and that person has limited time and is being paid.
Anthropic's threat intelligence report of 10 September 2026 is the clearest public evidence so far that this assumption is weakening in specific places. It covers activity disrupted between December 2025 and August 2026, and describes suspected state-sponsored groups, financially motivated criminals and politically motivated individuals using Claude inside agent harnesses.[1] Its central cyber finding concerns economics rather than technique. The methods were largely familiar: stolen credentials, exposed services, unpatched software, phishing, injection flaws. What changed was how many attempts could be made, how many targets could be handled at once, and how quickly a failed attempt could be turned into a different one.[1] In one case, monitoring agents checked whether deployed malware was being detected by known security products, and when it was, other agents modified and rebuilt it until detection stopped: a retooling loop that historically consumed a developer's day and now consumes scheduled compute.[1]
The question this paper asks is narrow: which controls still work when attempts are cheap, repeated, parallel and informed by fast feedback?
The argument has four parts.
- AI is lowering particular costs, and only particular costs. Research, reconnaissance labour, translation, scripting, environment interpretation, campaign coordination, variant production, data triage and the delay between failure and retry are all falling. Compute, infrastructure, initial access, scarce credentials, specialist exploits, legal exposure, detection risk and monetisation remain real costs.[1][2][15]
- Familiar attacks therefore reach more targets, get retried more often and are adapted faster. UK NCSC, Google Threat Intelligence Group, OpenAI, Microsoft, ENISA, Europol and CISA broadly agree, and most still assess that AI enhances existing tradecraft rather than creating a new class of exploitation.[2][4][7][8][9][10][11]
- Controls should therefore be assessed by their behaviour under repeated adaptive attempts: attempt limitation, feedback restriction, durability under adaptation, authority ceilings, elimination of reusable attacker state, and trusted recovery. Controls that work mainly by imposing human labour are the most exposed.
- None of this establishes offensive dominance. Defenders hold environmental knowledge, telemetry and control points, and RAND has argued that advanced AI could shift economics towards defence given the right people, processes and incentives.[13]
The contribution is a practitioner operating model, grounded in the 2026 incident record, connecting attempt budgets, empirical retry-sensitivity curves, feedback-loop speed, control durability, authority ceilings, state elimination and trusted recovery into a single per-attack-path assessment. The component ideas have extensive prior art in security economics, computational pricing, guessing budgets, NIST guidance on attempt ceilings, non-persistence and recovery integrity, and the 2011 kill chain paper's treatment of adversary feedback.[16][17][19][20][21][22][23][24][25][26] Google SRE uses retry budgets to constrain amplification, best-of-N safety research measures success under repeated sampling, and Microsoft's Security Durability Model focuses on controls that do not regress after remediation.[29][30][31] Garg and Dev have already written directly about AI and the economics of cyberattacks.[15] This paper claims the cross-domain attack-path assessment and its measurement, not the parts.
1. Familiar attacks, different economics
The Anthropic report runs to 154 pages and covers eight months of disrupted activity.[1] The observation that matters most for defensive engineering is not about any single intrusion. It is that the methods were mainly conventional while the economics around them had changed: reconnaissance, exploitation, tool development and data processing were increasingly delegated to models running in harnesses, in parallel, at machine speed, with individual operators handling dozens of victims concurrently.[1]
Anthropic states the qualifiers itself. It uses internal Generative Threat Group identifiers rather than public actor names, and defines uplift as the additional harm attributable to AI, assessed through speed, scale and depth, and presented as an analytical judgement rather than a measured counterfactual.[1] It also says humans retained target selection, monetisation and review of high-value results, and describes several of the most severe compromises as involving humans directing every step.[1] The report supports a claim about reduced operational labour. It does not support a claim that operations run themselves.
The often-quoted figure of breaches completed in two to three hours is a report-level synthesis rather than a measured average.[1] The closest named example is an intrusion associated with ShinyHunters affiliates that moved from one stolen developer token to full cloud administration in roughly three hours.[1] Read it as evidence that some intrusions now compress into a few hours, which is a meaningful change in the time available to a defender, and not as the duration of every case in the report.
With those qualifiers in place, the economic reading is straightforward. What becomes cheaper is the part of an operation that used to require a person to sit down and think about one specific target: reading an unfamiliar environment, writing the glue code for it, translating a lure into the right idiom, deciding which of the exfiltrated files matter, and then doing all of that again for the next target. What stays expensive is everything that consumes a resource the attacker cannot manufacture: a genuine credential to a hardened system, an exploit against a well-maintained product, infrastructure that survives takedown, compute that somebody has to pay for, and a route to turn access into money without being arrested.
That split has a direct consequence for control design. A control that charges the attacker in the first currency is losing value. A control that charges in the second currency, or that limits what a single success is worth, is not.
This is a change in degree rather than a new phenomenon. Schechter and Smith drew the distinction between serial and parallel theft in 2003: once tooling can be reused across similar targets, the marginal cost of each additional target falls sharply, so defences should impose a cost that cannot be amortised, or cap the value available from each success.[17] Kanich and colleagues showed in 2008 that a spam campaign could remain viable at a conversion rate below one in ten million because delivery was cheap and massively distributed, a historical figure whose lesson generalises: very low per-attempt success rates can sustain an operation when attempts are cheap enough.[20] Anderson's observation that attackers externalise their costs holds more literally when the compute is billed to the victim's API key,[19] and Becker's framing of offending as responsive to expected returns, detection and punishment remains the background model, allowing that state and ideological actors are not maximising financial return.[16]
What the 2026 record adds is that the reusable part of an operation now includes target-specific adaptation, which previously resisted reuse because it required judgement.
2. What Anthropic observed
The report's cyber section is organised around named threat groups. Five matter here. The matrix summarises them; the paragraphs that follow pick out the detail that bears on control assessment.
| Case | Actor assessment | Role of AI | Reported scale and effect | Evidence caveat |
|---|---|---|---|---|
| GTG-20006 | Russian-speaking espionage operation using the handle JackPoterz; tradecraft assessed as consistent with Russian state-nexus activity and with public reporting linking the actor to Midnight Blizzard | Infrastructure acquisition, phishing, persistence, command and control, lateral movement and exfiltration; monitoring agents checked whether malware was detected and rebuilt it until it evaded known products; scheduled jobs renewed stolen tokens and harvested cloud storage unattended | More than 20 organisations in planning, reconnaissance or live operations; more than two dozen Ukrainian government organisations scanned; at least three hotel Wi-Fi vendors compromised as indirect routes; more than 300,000 identity records from one government compromise | Attribution is an assessment rather than an independent determination; victim-side effects rest partly on open-source and partner reporting |
| GTG-50014 | Financially motivated operators suspected to be ShinyHunters affiliates, disparate and differently tooled, assessed as parts of one operation | Credential pipeline across ten cloud workers; 1.8 million Android packages decompiled and scanned for secrets; agents reported as performing nearly all of the work in one large intrusion | More than a terabyte exfiltrated from one provider; a SaaS compromise exposing roughly 200 downstream customers; more than 2,100 Azure AD token sets from more than 40 tenants in about 34 hours; one intrusion from stolen developer token to full cloud administration in roughly three hours | Volumes and victim counts are as reported, and some derive from actor claims or artifacts rather than victim confirmation |
| GTG-10007 | Sustained espionage by Chinese-speaking operators likely located in Changsha, Hunan; two identified as undergraduates, one with a security-vendor internship | Lead agent decomposing work to subagents; persistent project memory holding targets, credentials and campaign state between sessions; 13 standing collection agents on a schedule; an exploit foundry driving decompilers, forming hypotheses, generating and testing exploit code and iterating | Roughly 50 organisations targeted; more than a dozen possible zero-day findings from one continuous network-appliance workflow in a month; previously unknown vulnerabilities in a major security product reported as validated in the actors' laboratory | Possible zero-day findings are possible findings, not confirmed vulnerabilities, assigned identifiers or demonstrated exploits; no reproducible detail is published |
| GTG-50020 | Russian-speaking financially motivated actor previously focused on hotel booking and financial-technology platforms | Instructions injected into an AI vendor's evaluation sandbox, which exposed production API keys from several providers; per-target scope files launched parallel agents that ran injection, cross-site scripting and server-side request forgery testing against production applications unsupervised | Roughly 26 GB exfiltrated in one intrusion with a demand of USD 1.5 million to USD 2.5 million; a campaign against roughly 30 AI companies in about four days, repeating one successful path with minor adaptations; more than a dozen attempted routes to a pre-release model, all unsuccessful | Extortion figures are demands rather than payments; the pre-release attempts failed, which is evidence that some boundaries held |
| GTG-50029 | Single French-speaking operator targeting European political parties, media organisations, think-tanks and their SaaS providers | Coordinating framework managing subagents for reconnaissance, code review and cross-model checking; scanner for exposed API keys; local proxy rotating stolen keys to blend with legitimate traffic; model assistance developing a previously undocumented WordPress reinstallation race condition | 42 tracked targets with at least 14 internally accessed; estimated 12 GB to 26 GB exfiltrated including membership, donor and payment data; about 140,000 records containing political opinions from an exposed search endpoint; the race condition reported as successful against at least four sites | Exfiltration volume is an estimate; the race condition is described as previously undocumented within the report and is not independently catalogued |
Four details in that table do most of the analytical work.
The rebuild loop in GTG-20006 is the cleanest example of automated retooling. Agents deployed malware, monitored whether known security products detected it, rebuilt it when they did, and iterated until the tooling was undetected before staging it on disposable hosting.[1] The defensive value of a signature includes the retooling work it forces, and when that work is done by a scheduled agent it is paid in compute rather than in operator hours.
The exploit foundry in GTG-10007 is the clearest example of iteration applied to vulnerability research. Firmware and binaries were loaded into decompilers through tool servers; agents traversed decompilation and cross-reference chains, formed hypotheses, generated exploit code, tested it against laboratory systems and iterated.[1] The reported output was more than a dozen possible zero-day findings from one workflow in a month, which are candidate results from an automated pipeline rather than confirmed weaponised zero-days.
The two-to-three-hour synthesis matters for what it does to response time rather than for its precision. If an intrusion can move from a stolen token to cloud administration in about three hours, a control that depends on a human noticing, deciding and acting within that window has almost no margin.[1]
The parallel-victim pattern appears in every case: dozens of targets handled concurrently by one operator, one successful path replayed with minor adaptations, standing agents collecting on a schedule.[1] A single operator occupies a position that previously required a team. A stolen-key thread also runs through GTG-50014, GTG-50020 and GTG-50029, with keys taken from customer environments, used for weeks, and rotated through a proxy to blend with legitimate traffic; Anthropic states that these were customer keys and that its own systems were not compromised.[1] Section 10 returns to this.
3. What the report does and does not prove
The report is primary vendor telemetry about selected operations. It is unusually detailed, and it is not a prevalence study. Six limitations should travel with every use of it.
Selection and detection bias. The cases are ones Anthropic detected, investigated and chose to publish, and there is no denominator for total Claude activity, total malicious AI-assisted activity, or total intrusions in the period.[1] Operations that are noisy, clumsy or unsuccessful are easier to find, and actors who concealed model use, ran local models or used other providers are outside the sample by construction.
Visibility boundary. Anthropic says its visibility can end once activity leaves its platform, and that it then relies on open-source research, partner data and public reporting.[1] Victim-side outcomes rest on evidence of varying strength, and some figures originate in actor claims.
Counterfactual weakness. Uplift asks how much additional harm occurred because of AI, and the same actors running the same campaigns without AI were not observed. This is an analytical judgement, and Anthropic presents it as one.[1]
Attribution language. Suspected, consistent with, likely and assessed are doing real work in the report and should not be flattened. The GTG-20006 attribution is described as consistent with public reporting linking the actor to Midnight Blizzard rather than as an independent determination.[1]
Outcome uncertainty. Possible zero-day findings, estimated data volumes and extortion demands are not confirmed vulnerabilities, verified losses or received payments.
Vendor incentives. Anthropic benefits from demonstrating both that its models are capable and that its safeguards work. That does not invalidate the evidence, and it does shape which cases appear and how they are presented. The statement that no malicious activity was found on particular safeguarded models means that none was detected in this investigation. One practical consequence: use the dated PDF and keep a copy, because live pages are revised after publication.[1]
What the report does support is narrower and still substantial: that multi-agent harnesses were used against real victims; that a detection-driven rebuild loop ran without human involvement; that vulnerability research and exploit development ran continuously under agent control; that one operator handled dozens of victims concurrently; and that stolen AI credentials were used as working attack infrastructure.[1] That justifies re-examining controls which assume attacker retooling is expensive. It does not establish how common any of it is.
4. Independent evidence
A single vendor report should not carry an engineering argument on its own. Government and independent sources published between early 2025 and late 2026 converge on one conclusion: AI currently produces its largest demonstrated cyber effect by making familiar tactics faster, cheaper, more parallel, easier to localise, and available to less-skilled operators.
UK NCSC assessed in May 2025 that AI would almost certainly increase the volume and impact of attacks through the evolution of existing tactics, techniques and procedures rather than through novel threat vectors, and forecast that AI-assisted vulnerability research would shorten the interval between disclosure and exploitation.[2] That is an intelligence forecast rather than an incident census, and NCSC says technical surprise remains possible.
Google Threat Intelligence Group reported in January 2025 that threat actors were obtaining productivity gains across research, reconnaissance, scripting, payload development and localisation, without observing breakthrough capability.[4] OpenAI's October 2025 disruption report described actors attaching AI to existing playbooks to move faster rather than acquiring capabilities they did not previously have.[7] Microsoft's 2025 Digital Defense Report characterised most AI augmentation as automation of previously time-intensive activity.[8] ENISA described a new level of malicious scalability while characterising most observed use as augmentation.[10] Europol's 2026 assessment identified speed, efficiency, reach, personalisation and lower entry barriers as the dominant present effects.[11] CISA's review of fiscal years 2024 and 2025 notes increasing AI automation of exploitation steps against a baseline in which most compromises still involve known, relatively simple flaws in exposed systems.[9] Each source has its own visibility limits: providers observe their own platforms, ENISA aggregates reporting that is partly derivative, Europol combines member-state input with forward assessment, and CISA does not quantify AI prevalence. The agreement between them is about direction rather than magnitude.
There is nevertheless a real change between early 2025 and late 2026, and it is where the older assessments start to strain. GTIG's November 2025 tracker described PROMPTSTEAL, used by APT28, querying a model during live operations to generate discovery and collection commands, and PROMPTFLUX, which attempted to regenerate or obfuscate its own code through model calls.[5] PROMPTFLUX was incomplete, and the provenance of the stolen token it used was assessed rather than established. In February 2026, GTIG reported HONESTCUE obtaining stage-two download-and-execute code from a model, while still finding no revolutionary shift.[6] Anthropic then reported autonomous detection-and-rebuild loops, multi-agent exploitation against production targets, continuous exploit research, a previously undocumented WordPress race condition used against victims, and possible zero-day findings at volume.[1]
The right distinction is between four claims of decreasing support.
- Novel workflow and operational architecture: increasingly supported by the 2026 record.
- Novel vulnerabilities and individual exploits produced with AI assistance: emerging, incompletely validated.
- A fundamentally new class of exploitation deployed at scale: not established.
- Fully autonomous strategy, target selection and monetisation: not established.
This creates some tension with NCSC's judgement that fully automated end-to-end advanced attacks were unlikely by 2027.[2] It does not clearly falsify it, because the actors in Anthropic's cases retained target selection, review, monetisation and in some cases hands-on control of the most severe intrusions.[1] A defensible reading is that operational execution has automated considerably faster than strategic direction. Deployment-side guidance points the same way: Ireland's NCSC risk assessment for public-sector AI notes that agentic systems can act at scale with elevated credentials and that AI expands the attack surface in ways requiring continuous rather than point-in-time controls.[3]
5. What becomes cheaper
Precision matters more than emphasis here, because the argument depends on which costs move and which do not.
Costs that AI reduces:
- Research and reconnaissance labour, including reading unfamiliar environments and summarising what was found.
- Translation and localisation, which previously constrained who could be phished convincingly and in what language.[11]
- Scripting, glue code and target-specific interpretation: working out what this particular network, application or tenancy is and where the interesting things are.
- Campaign coordination across many concurrent targets.[1]
- Variation in malware and phishing artifacts, including rebuilds driven by observed detection.[1][5]
- Data triage after exfiltration, which converts volume into usable intelligence or saleable records.[1]
- The delay between a failed attempt, the decision about what to change, and the next attempt.
Costs that remain:
- Compute. Agent loops, decompilation pipelines and exploit testing consume real resources, which is why stolen keys and cloud credentials are attractive.[27][28]
- Infrastructure. Domains, hosting, proxies and command channels cost money and are subject to takedown.
- Initial access. Something still has to obtain a first foothold, and against a hardened target that remains the expensive step.
- Scarce credentials. Hardware-bound, phishing-resistant authenticators are not produced by iteration, and specialist exploits against well-maintained products remain costly.
- Legal and detection risk, which rises with volume and noise.
- Monetisation, which still involves buyers, laundering, negotiation and exposure.
Two structural points follow.
Some remaining costs can be transferred rather than paid. Stolen cloud credentials provide compute, stolen API keys provide model access billed to someone else, and compromised hosts provide infrastructure.[19] A control that charges a scarce resource only works if the attacker cannot charge it to a third party.
The fall in coordination and interpretation cost also changes who is a target. When per-target human effort dominates, attackers concentrate on organisations worth the effort; when it falls, the population of economically viable targets widens.[17] Smaller organisations whose main protection was obscurity or low value are the most affected.
6. Retry economics without false precision
It is tempting to reach for a formula, and the usual formula is wrong in exactly the cases that matter.
Let q_i be the conditional probability that attempt i succeeds, given that all earlier attempts failed. Then:
P(success by N attempts) = 1 - product_i(1 - q_i)
That is a chain-rule identity, correct and assuming nothing about independence or identical trials. Its practical weakness is that the conditional probabilities are hard to estimate, and estimating them is the whole problem.
The familiar simplification is:
P(success by N attempts) = 1 - (1 - p)^N
This holds only when the conditional success probability stays constant at p along the failure path. In cyber operations that condition is routinely violated, in both directions:
- Failed password guesses eliminate candidates, so later guesses are drawn from a smaller pool.
- Error messages, timing differences and response variance teach the attacker something, so later attempts are better informed.
- Variants produced by a rebuild loop are related, so their detection outcomes are correlated, and controls adapt after detection, so the defence facing attempt 50 is not the one that faced attempt 1.[1]
- Lockouts and quotas reduce later opportunities, while distributed identities, sources and tenants bypass per-account and per-source counters, so the true attempt count exceeds what any single counter observes.
- Exploitability is correlated across similar systems, which is why one working path can be replayed against 30 companies with minor adaptation.[1]
- Attackers stop after success, and defenders patch or isolate during the campaign.
One correction matters for guess budgets in particular. For a single fixed secret drawn uniformly from M possibilities, attacked with N distinct guesses, P(success) = N / M, which is not the independent-trials expression, and real password distributions are not uniform. Bonneau's ordered guessing metrics are the better foundation there.[22]
The practical alternative is an empirical retry-sensitivity curve over a defined population of test cases:
S(N) = proportion of test cases compromised
when each case is allowed at most N adaptive attempts
S(N) is attacker success, so lower is better. It is only meaningful under stated conditions: the attacker capability assumed, the target population and number of test cases, the control configuration and version, the time window, the identity and source distribution, the feedback channel available to the attacker, and the stopping rule. Two teams measuring S(N) under different assumptions are measuring different things.
Most organisations will not have enough incident data to estimate S(N) from production. It can still come from adversarial testing, purple-team exercises and repeated test cases that each receive an attempt budget, provided the conditions are recorded alongside the numbers. A single ordered series of 200 variants against one target does not estimate a success proportion by itself; for that case, record the attempt at which first success occurred and the outcome sequence instead of labelling it S(N). Best-of-N safety research offers a close methodological parallel, while operating in a different threat domain.[30]
The question the curve answers is whether the control still performs acceptably at the attempt counts and adaptation speeds now economically available to the actors in scope. A control that blocks the first five attempts and fails at the fiftieth was adequate when fifty adaptive attempts meant a fortnight of skilled work. Its value is different when they mean an afternoon of scheduled compute. Two traps come with the measurement: aggregate attempt counts understate the budget of an attacker who distributes across identities and sources, and comparing curves across control versions without re-testing older variants hides regression.
7. A control taxonomy under cheap retry
The following classification is analytically useful and is not an established standard. It groups controls by the mechanism through which they work, because that mechanism determines how each behaves when attempts get cheaper. The classes are overlapping lenses rather than a partition: a spend cap is both attempt-limiting and a scarce-resource control, while a clean rebuild both eliminates state and supports recovery.
| Control class | Mechanism | Effect of cheaper attempts | Examples |
|---|---|---|---|
| Labour-friction | Requires interpretation, manual adaptation or target-specific work from the attacker | Most directly weakened, because the work can be delegated to a model or outsourced | Obscure configuration, signature-driven retooling cost, puzzles requiring manual solving, analyst-only triage |
| Scarce-resource | Charges a resource that remains scarce for each attempt | Useful while the resource cannot be cheaply stolen, parallelised or billed to somebody else | Money, trusted identities, hardware-backed keys, proof-of-work |
| Attempt-limiting | Caps or delays admissible attempts | Strong where enforcement covers the attacker's real distributed campaign rather than one counter | Rate limits, lockouts, quotas, concurrency caps, spend ceilings |
| Feedback-limiting | Reduces the quality or speed of information available for adaptation | Rises in importance as automated feedback loops get faster | Generic errors, uniform timing, delayed disclosure, deception and canaries |
| Authority and blast-radius | Limits what one success can reach, change or remove | Does not weaken directly when attacker labour falls, but erodes through exceptions or principals able to change the boundary | Least privilege, segmentation, scoped tokens, egress restriction, transaction approval |
| State-eliminating | Invalidates or destroys reusable attacker state | Useful in proportion to how completely the relevant state is removed | Session and token revocation, credential rotation, ephemeral rebuilds, reimaging |
| Adaptive detection and containment | Learns from observed activity and changes the response | Necessary, and creates an arms race that can leak feedback to the attacker | Behavioural analytics, variant clustering, automated isolation |
| Recovery | Restores trusted operation and bounds the duration of a compromise | Essential; cheaper attacks do not make restoration faster | Isolated backups, tested rebuilds, dependency recovery, clean-room restoration |
Microsoft's Security Durability Model uses durability to describe fixes that remain enforced, resist drift and are tested for regression over time.[31] This paper uses control durability more narrowly: whether a control continues to limit an attack path as attempts become repeated and adaptive. A durable control needs both properties.
The classes most exposed to cheap attempts share a property: their cost to the attacker is denominated in human attention. Static signatures used as the only malware control are exposed because the retooling they force can be automated, as GTG-20006 demonstrated.[1] Security through obscurity and any friction assuming a human inspects each target are exposed because reading an unfamiliar system is something models are demonstrably good at, and manual-only triage is exposed because the attacker's iteration rate is no longer bounded by a person while the defender's still is.[4][7] Per-source rate limits without identity, device, tenant or campaign correlation are exposed because distribution is cheap, and one-time blocking that leaves sessions, credentials and persistence intact is exposed because the attacker keeps reusable state. Puzzles whose solution can be outsourced more cheaply than the resulting criminal return are exposed by arithmetic, and Dwork and Naor's pricing argument depends on a solver cost that is not stable when compute is stolen and hardware is unequal.[21]
The classes that hold up best are the ones whose cost to the attacker does not fall when labour does:
- Hard attempt and concurrency limits enforced across the whole identity estate.[23]
- Phishing-resistant authentication, short-lived scoped credentials, and deterministic authorisation boundaries evaluated per action.
- Segmentation, egress restriction and transaction-level approval for high-impact actions.
- Non-persistent or rapidly replaceable workloads, which NIST SP 800-160 Volume 2 treats as a cyber-resiliency technique in its own right.[24]
- Complete session and token invalidation, spend caps and usage quotas.
- Isolated, tested recovery capable of restoring trusted state.[26]
No class is sufficient alone. Attempt limits do not help once an attempt succeeds, authority ceilings do not stop the first compromise, state elimination is a response rather than a preventive measure, feedback limitation slows adaptation without preventing it, recovery bounds consequence without reducing likelihood, and detection catches what the current behavioural models cover. Classifying controls this way shows which parts of the portfolio depend on an assumption that is weakening. Libicki, Ablon and Webb made a related point about the difference between making a system hard to enter and limiting what follows once somebody is inside.[18]
8. Static detection under an automated rebuild loop
The GTG-20006 rebuild loop deserves careful handling, because it is the most quoted detail in the report and the easiest to over-read. Monitoring agents checked whether deployed malware was detected by known security products; when detection occurred, agents modified and rebuilt it; iteration continued until the tooling was undetected, after which it was staged on disposable hosting.[1] That supports one claim, that an automated loop closed the gap between a detection event and a modified artifact without a human developer in the path. It does not support the claims that signatures are useless, that behavioural detection was defeated, or that this loop is now standard practice. One case in one report, selected for publication, is not a prevalence finding.
The useful conclusion concerns which part of a signature's value is affected. A static detection stops the specific artifact it matches, and it imposes a retooling cost. The first is unchanged. The second shrinks when the retooling is automated, and shrinks rather than disappears, because compute, testing and staging still cost something and each rebuild risks producing a variant that fails or trips another control.
Three practical implications follow.
Detection should be weighted towards behaviour and effect rather than artifact form. What an implant does to obtain persistence, reach a command channel, enumerate an environment and move data changes more slowly than how its bytes are arranged. This is the observation at the centre of intelligence-driven defence: indicators derived from adversary methods and objectives survive longer than indicators derived from one file.[25] Google's reporting on model-assisted malware points the same way, since PROMPTFLUX regenerated its own code while its purpose stayed constant.[5]
Variant clustering should be a first-class capability, because a rebuild loop generates a family with shared structure across build artifacts, configuration, infrastructure and staging pattern. Detecting the family is the durable objective; detecting each member is a treadmill whose speed the defender does not set. Isolation limits the value of any single successful variant in the same way: an endpoint running an undetected implant while holding long-lived cloud tokens and standing administrative rights makes the evasion worth a great deal, and the same evasion against a host holding a short-lived scoped credential is worth far less.
Detection feedback should be treated as an information disclosure. A rebuild loop needs a verdict, and anything that gives an adversary a fast, cheap oracle for whether an artifact is currently detected accelerates the loop: public multi-engine scanning of unreleased samples, verbose block messages that name the detection, distinctive rejection behaviour, consistent timing differences. Some of this is outside the defender's control, and some trades against operational transparency for legitimate users. The design question is how much precision a failure response needs to carry, and the answer is usually less than it currently carries.
Signatures remain worth having. They contribute less deterrent cost than they used to once the adversary has automated the rebuild loop, and they should not carry weight that a variant generator can remove overnight.
9. Applying the model across common attack paths
The questions are the same for every path: what counts as one attempt, what a failed attempt returns to the attacker, what authority one success obtains, and what reusable state remains afterwards. The answers differ enough to be worth writing down separately.
9.1 Credentials and recovery channels
Attempt unit: one authentication attempt against one identity, or one recovery request against one account. Recovery channels are frequently weaker than the primary authenticator and are counted separately, if at all.
Feedback: differentiation between unknown user, wrong password, correct password awaiting a second factor, and lockout, with timing carrying the same information more quietly. Anything that separates valid from invalid identifiers turns guessing into enumeration followed by a smaller guessing problem.
Authority ceiling: whether the credential can mint further credentials, create service principals, alter logging, change recovery details or approve its own escalation. A stolen developer token that reached full cloud administration in roughly three hours is the illustration.[1]
State elimination: every session and token derived from the credential, including refresh tokens, device registrations, application passwords, keys created during the compromise, and downstream sessions established through federation. NIST SP 800-63B gives the baseline for attempt ceilings and delays, and those limits apply per authenticator rather than to breadth-first guessing across an estate or to weaker recovery channels.[23]
9.2 Exposed services and vulnerability exploitation
Attempt unit: one exploitation attempt against one exposed asset, with the campaign unit being the sweep across a discovered population.
Feedback: response codes, error text, version banners, differential behaviour and timing. An exploit-development loop of the kind reported for GTG-10007 needs a testable oracle, which in that case was a laboratory copy of the product.[1]
Authority ceiling: service accounts holding database credentials, instance metadata and unrestricted egress turn one application flaw into a tenancy compromise.
State elimination: rebuilding from a trusted image, because patching in place leaves web shells, scheduled tasks, modified configuration and harvested credentials.[24] CISA reports that most compromises still involve known, relatively simple flaws in exposed systems, so exposure inventory remains a dominant variable however the exploitation step is automated.[9] UK NCSC separately forecasts that AI-assisted vulnerability research will shorten the interval from disclosure to exploitation, increasing the value of patch speed.[2]
9.3 Phishing
Attempt unit: one message to one recipient, with the campaign unit being the domains, templates, languages and payloads used in a period.
Feedback: delivery telemetry, bounce behaviour, click and submission rates, and whether a lure domain gets blocked. Fast feedback plus cheap variation produces a tuning loop, which is what localisation at scale means in practice.[8][11]
Authority ceiling: a session token unlocking mail, files, chat and a dozen federated applications is worth a great deal. An origin-bound hardware authenticator prevents the phishing message itself from yielding a replayable primary credential. Attackers can still target the session issued afterwards, authenticator enrolment, fallback methods and recovery channels.
State elimination: the captured session everywhere it has been used, including federated applications and any credentials created afterwards. Blocking the sending domain closes one channel and does nothing about access already obtained.
9.4 Malware and endpoints
Attempt unit: one variant delivered to one host. Feedback: the detection verdict itself, which is what the GTG-20006 loop consumed.[1]
Authority ceiling: cached cloud tokens, browser sessions, stored keys, mapped shares and lateral reachability set the value of one successful execution.
State elimination: isolating a host is partial if tokens taken from it stay valid, lateral sessions remain open and persistence exists elsewhere. Endpoint isolation and identity revocation belong in one automated response rather than in sequence across two teams.
9.5 Cloud control planes and AI APIs
Attempt unit: one API call, or one attempt against one control-plane identity.
Feedback: authorisation errors distinguishing nonexistent from forbidden, policy evaluation differences, and quota responses. Control planes are conveniently uniform for an attacker, which is what allows one successful path to be replayed across roughly 30 organisations with minor adaptation.[1]
Authority ceiling: whether a principal can create identities, change policies, disable or redirect logging, raise quotas, or reach data directly. A principal that can alter its own observability has a higher effective ceiling than its permission set suggests.
State elimination: keys, sessions, refresh tokens, service principals, grants and any identities created during the compromise, followed by verification that the mechanism which created them is gone. Uniformity cuts both ways: it also lets a defender enforce spend caps, egress restriction and workload identity everywhere at once.
9.6 Incident recovery
Attempt unit: one re-entry attempt during or after recovery, using credentials, persistence or backdoors that survived. Feedback: the visible progress of the recovery, which an adversary retaining any access can watch.
Authority ceiling: what the restored environment grants. Restoring from a backup taken after the compromise restores the attacker's position along with the data.
State elimination: this is the core of recovery rather than a step within it, and the order matters, because rotating credentials while the harvesting mechanism still runs produces a fresh set of stolen credentials. NIST SP 800-61 Revision 3 treats restoration integrity as part of incident response and cautions against reading too much into simple response-time metrics.[26] Measure time to trusted restoration separately from time to service availability; the gap is informative.
10. AI API credentials as attack infrastructure
This section is deliberately subordinate to the main argument. It belongs here because AI credentials are a control gap created by the same economics, and because they appear repeatedly in the 2026 incident record. An AI API key now functions as several things at once: a production credential, a metered compute entitlement, potential access to sensitive prompts and data, operational attack infrastructure, and a resaleable commodity. Many organisations protect cloud administrative credentials carefully and treat model-provider keys as ordinary application secrets. That gap is the point.
Model access. A stolen key gives an attacker models, quotas, regions and features they may not otherwise hold. Anthropic reported a hacktivist campaign operating on stolen keys for around a month, and ShinyHunters-associated operators moving their attack workloads onto victims' keys after obtaining them.[1]
Stolen compute and spend. The victim pays. Sysdig documented stolen cloud credentials being validated automatically against ten AI services, and estimated that full abuse of the applicable Claude quotas could exceed USD 46,000 per day.[27] Unit 42 has reported stolen credentials integrated into a transfer station within minutes, with one case approaching USD 1 million in charges before containment.[28] Both are vendor incident-response reports: the Sysdig figure is a modelled worst case, and Unit 42 withholds victim identities and timelines. They are indicative rather than a loss rate for any industry.
Resale and access routing. Keys can be pooled behind reverse proxies and resold without exposing the underlying credentials, with the proxy handling rotation, routing and billing.[27][28] Anthropic reported a group combining fraudulent resale with malware that harvested customer credentials and replacement sessions.[1] Shared subscriptions, compromised accounts and transfer-station proxies can also route around geographic or provider controls; the joint NSA, CISA and FBI advisory describes millions of requests and billions of tokens moving through distributed pathways, while making clear that not every pathway involved a stolen key.[12]
Attribution ambiguity. Anthropic says stolen keys provide cover, because activity is initially associated with the legitimate customer's credential.[1] The accurate claim is that they complicate initial attribution and misdirect account-level logs, rather than that they guarantee durable concealment: providers can correlate network, device, payment, behavioural and usage telemetry over time.
Persistence. AI credentials support persistence when keys are long-lived, when a compromised principal can mint replacements, when billing alerts can be disabled, when stolen refresh tokens renew sessions, when hooks intercept newly configured keys, or when malware keeps harvesting replacements. Rotation alone is insufficient wherever the mechanism that steals or mints keys is still present. This is state elimination applied to a credential class often excluded from identity governance entirely.
The controls follow from the functions: short lifetimes, scoped permissions, workload-bound identity, egress restriction, hard spend caps, usage anomaly detection, revocation that reaches derived sessions, monitoring of key creation, and alerting on changes to logging and quota configuration.
11. Defender advantages and automation
An argument about cheaper attempts can drift into an assumption that offence wins. The evidence does not support that. RAND's Heitzenrater argues that advanced AI could shift the economics of cybersecurity towards defenders, because defenders possess what attackers have to work to acquire: knowledge of their own environment, its data, its normal behaviour and its control points.[13] That is a strategic perspective rather than an empirical estimate, and the logic is sound. An attacker running an agent against an unfamiliar network is inferring structure from outside, while a defender running an equivalent capability inside has the inventory, the identity graph, the configuration history and the telemetry.
RAND's 2026 controlled study cuts the other way and should be reported honestly: agentic systems allowed novices to complete offensive cyber challenges previously beyond them.[14] Those were controlled challenges without production constraints, detection, legal exposure or the friction of a real environment. The finding is about capability floors rising, which is a real effect and a different claim from dominance.
Defenders have three automation opportunities that match the pressures described here.
Containment speed. If an intrusion can move from stolen token to cloud administration in about three hours, containment that waits for a human decision is the bottleneck.[1] Pre-authorised automated containment for well-understood conditions closes that gap: revoke sessions on confirmed credential theft, isolate a host on confirmed execution of a known family, disable a key on confirmed anomalous usage. The prerequisites are demanding: precise triggers, bounded scope, a tested reversal path and a named owner.
Correlation. Cheap distributed attempts defeat counters that each see a fraction of the campaign, and correlating attempts across identities, devices, tenants, source networks and time scales badly with analysts and well with automation.
Recovery. Rebuilding from trusted sources, verifying that restored state is clean, and re-establishing identity and key dependencies in the right order are all mechanisable. Automation removes serial human steps without proving any universal recovery time.
Defender automation carries its own risks. Automated actions have blast radius, agents holding production authority need authority ceilings of their own, and Ireland's NCSC assessment notes the risk of agentic systems acting at scale with elevated credentials.[3] The honest summary is that automation is now necessary on both sides and confers no automatic advantage on either.
12. A practical operating model
The purpose of this section is to make the preceding argument usable. The unit of assessment is an attack path rather than a control, because the same control performs differently depending on what is being attempted through it.
For each attack path in scope, record the following fields.
Target population. The assets, identities or endpoints the path applies to, and how many there are, since reusable tooling is what makes a large population attractive.[17]
Attempt unit. The smallest action that either succeeds or fails. Without this, none of the other fields can be counted.
Admission boundary. The component that decides whether an attempt is admitted, and the identity, device, tenant, network or resource that owns the counter.
Attempt budget. How many attempts the actor can make before hitting a hard cap, a scarce cost, detection or a time constraint, whether that cap is aggregate or per source, and whether it can be reset or distributed around.
Google SRE uses retry budgets to prevent cooperating clients from amplifying failures in distributed systems.[29] The accounting idea is useful here, while the trust model differs: a hostile actor will not self-throttle, so the admission boundary has to enforce the budget across the identities and sources the actor can distribute over.
Scarce per-attempt cost. What each attempt costs the attacker, and whether it can be stolen, outsourced, amortised or billed to a third party.[19]
Feedback exposed. What a failed attempt tells the attacker through error content, status differentiation, timing, side effects or observable defensive response.
Attacker adaptation time and defender response time. How long the attacker needs to turn a failure into a materially different attempt, against how long the defender needs to detect, decide and deploy a change. If adaptation is faster, the control is outpaced and needs support from a class that does not depend on response speed.[1]
Control result after N attempts. The retry-sensitivity result from section 6 with its conditions recorded, choosing N from what is economically available to the actors in scope.
Authority obtained on success, and reusable state created. What one success grants, including whether it can mint credentials, alter logging, change policy or raise quotas; and the sessions, tokens, keys, registrations, persistence mechanisms, implanted code and modified configuration it leaves behind.
State invalidation method. The mechanism that removes each item above, who runs it, how long it takes, and how completion is verified.
Trusted recovery method. How the service is restored to a state that does not contain the attacker, including backup isolation, rebuild sources, dependency ordering and verification.
Evidence confidence. For each field group, whether the value is a production observation, a test result, a vendor assessment, an intelligence forecast, a model assumption or an estimate.
Two cautions apply. Do not collapse the record into a single risk score, because several fields are ordinal, correlated or uncertain. And do not fill it in once: attacker adaptation time and control result after N attempts both change without anything in the defender's environment changing at all.
13. What to measure and test
Indicators are only useful when somebody owns them and they change a decision. Alongside the path-specific lists below, three general measures matter everywhere: attacker success S(N) after 1, 10, 100 and 1,000 adaptive attempts with conditions recorded; the proportion of incidents in which all reusable attacker state was invalidated; and time to trusted restoration measured separately from time to service availability.
Credentials:
- Guesses admitted per account and across the whole identity estate, including breadth-first patterns.
- Coverage of phishing-resistant authentication, by population rather than by licence count.
- Token and session lifetimes, time to revoke every derived session, and the proportion of credentials that can mint further credentials.
- Rate limits and identity-proofing strength on recovery channels.[23]
Exposed services:
- Completeness of the internet-exposed asset inventory, measured against independent discovery.
- Patch latency from disclosure and from first observed exploitation.[2][9]
- Residual vulnerable assets after remediation is declared complete, and the proportion of workloads that can be recreated cleanly rather than patched in place.[24]
Phishing:
- Unique domains, templates, languages and payload variants per campaign.
- Time from first user report to domain, token and session containment.
- Whether a capture yields broadly reusable credentials, and whether revocation reaches every connected application.
Malware and endpoints:
- Observed variant generation interval within a family.
- Detection rate against chronologically ordered, previously unseen variants, and regression rate against previously detected ones.
- Whether host isolation also invalidates cloud tokens and lateral sessions.
Cloud and AI APIs:
- Number, age and scope of standing keys, including keys outside central inventory.
- Quota and spend ceiling per key, and usage deviation from baseline.
- Whether a principal can disable logging or raise its own quota, and time from suspected theft to key and derived-session invalidation.
Recovery:
- Restore tests completed successfully, by service rather than by platform.
- Time to restore identity, key and directory dependencies.
- Evidence that restored data and images are free of attacker-introduced state, and re-compromise events during or shortly after restoration.[26]
Conformance tests turn those indicators into evidence. Each should produce a pass, a fail or a documented gap.
- Distributed bypass. Run the attempt volume a per-source limit is meant to block, spread across many sources, identities and tenants, and see whether any aggregate control notices.
- Adaptive variants. Measure detection across a series of generated variants of a known family rather than at a point.
- Stale detections. Re-test previously blocked variants after each update, recording regression explicitly.
- Token and session revocation. Revoke a credential, then attempt to use every session, refresh token, registration and downstream session derived from it.
- Compromised credential minting. Establish whether a realistic compromised principal in a test tenancy can create credentials, alter logging or raise quotas.
- Feedback leakage. Compare error content, status codes and timings across valid and invalid identifiers, detected and undetected artifacts, and permitted and forbidden actions.
- Control regression. Confirm that the current configuration still blocks what last year's blocked, and that no exception has quietly become permanent.
- Clean rebuild. Rebuild a representative service from trusted sources and verify that no configuration, credential or artifact carries over.
- Re-compromise during recovery. Model an actor retaining one credential and one persistence mechanism through the recovery sequence, and check whether the plan removes them and in what order.
For each result, record what kind of evidence it is. A controlled test, a production observation, a vendor assessment and an intelligence forecast are different things, and mixing them is how a measurement programme ends up with confident numbers that nobody can defend.
14. The strongest counterargument
The strongest objection is that this paper takes one vendor report, generalises from it, and recommends expensive changes on the strength of a trend that may not hold. It has several parts, and they deserve separate answers.
Attacker costs remain, and some are rising. Infrastructure, access, scarce credentials, legal exposure and monetisation are unchanged or harder, and Anthropic's own record includes an actor who attempted more than a dozen routes to a pre-release model and failed every time.[1] That is why this paper argues about which costs moved rather than about cost in general.
Automation is noisy. High-volume machine-driven activity generates telemetry, and noisy actors are easier to find, which is part of why they appear in vendor reports at all. That is a detection bias in the evidence as well as a real defensive advantage, and it has limits: noise only helps when somebody is looking at the right signal within the attacker's operating window, and a three-hour intrusion does not leave much of one.[1]
Defenders get the same tools. Section 11 accepts this, and RAND's argument that defenders hold the environmental advantage is serious.[13] The claim here is conditional: these controls matter most if the defender's environment does not also automate, and they remain sound engineering if it does.
Scarce chokepoints still exist. Hardware-bound authenticators, deterministic authorisation boundaries, spend caps and segmentation do not weaken when attacker labour gets cheaper, which is why the recommendation is to move weight towards those classes rather than to abandon everything else. Static signatures also remain a useful layer: they stop known artifacts cheaply, carry low false-positive rates when written well, and filter commodity activity before anything expensive has to look at it.
Tighter controls have costs. Hard attempt limits lock out real users, aggressive revocation breaks integrations, generic errors slow legitimate debugging, and automated containment creates its own outages. Organisations that ignore that price roll the control back after the first bad week, which is worse than never deploying it. The mitigation is proportionality: apply the tightest settings to the highest-authority paths, and measure the legitimate-user cost as deliberately as the security benefit.
The evidence base is thin, and this is the most serious part of the objection. There is no industry-wide incident rate for AI-enabled operations, no denominator in the Anthropic report, and no controlled counterfactual in the literature.[1] The independent corroboration is directionally consistent and mostly assessment rather than measurement.[2][4][7][8][9][10][11] The proportionate response is the one taken here: argue for measurement and testing rather than for a magnitude, and treat the retry-sensitivity curve as something an organisation measures for itself.
A final version of the objection is that all of these recommendations were good practice before 2026. They were, and they appear in long-standing guidance on attempt ceilings, non-persistence and recovery integrity.[23][24][26] That is why the contribution here is a prioritisation and assessment method rather than a new control. What the 2026 record changes is the relative weighting, and the case for spending the next increment of effort on classes that do not depend on the attacker running out of patience.
15. What this paper does not claim
It does not claim that attack cost has become zero. Compute, infrastructure, access, scarce credentials, legal risk and monetisation all remain, and some are the binding constraint in most operations.[1][15]
It does not claim that every actor is autonomous. Anthropic reports humans retaining target selection, monetisation and review of high-value results, and describes several of the most severe compromises as involving humans directing every step.[1]
It does not claim that retries are independent. Section 6 sets out why the independent-trials formula is usually wrong here, and why an empirical curve under stated conditions is more defensible.[22]
It does not claim that static detection is useless. The narrower claim is that the retooling cost a signature imposes shrinks when the retooling loop is automated.[1]
It does not claim that friction controls are universally obsolete. Friction that charges a resource which stays scarce still works, and Dwork and Naor's pricing argument holds where the solver's cost is real and cannot be stolen or outsourced.[21]
It does not claim that Anthropic's cases establish prevalence. They are selected, vendor-detected operations without a denominator, drawn from one provider's visibility.[1]
It does not claim that possible zero-day findings are confirmed vulnerabilities. They are candidate outputs of an automated workflow, without published technical detail, independent validation or assigned identifiers.[1]
It does not claim that all AI API key abuse is key theft. Shared subscriptions, compromised accounts, commercial proxies and unauthorised aggregation produce similar traffic through different mechanisms, and the joint advisory on distillation describes pathways that require no stolen key at all.[12]
It does not claim that defenders lose. RAND argues the opposite case seriously, and defender automation may produce equal or greater advantage.[13]
It does not claim any industry-wide incident rate for AI-enabled operations. No source used here supports one.
It does not claim that the control taxonomy in section 7 is an established standard. It is an analytical grouping intended to make the assessment questions easier to ask.
Conclusion
The change documented in the 2026 record is narrower than the headlines suggest and more consequential than a list of techniques implies. The techniques were familiar. What moved was the cost of trying again, of trying against one more target, and of changing the attempt after it failed.
That matters because a meaningful part of the defensive portfolio was built on the assumption that repeated, target-specific adaptation consumes scarce human effort. Signatures work partly by forcing retooling. Obscurity works by forcing interpretation. Manual triage works when the adversary is also constrained by a working day. Each has become weaker in proportion to how much of its value came from the attacker's labour cost.
The controls that retain more of their value are the ones whose cost to the attacker does not fall when labour does: hard attempt limits across the identity estate, credentials that cannot be guessed or minted freely, authorisation evaluated per action, segmentation and egress restriction that cap what one success reaches, revocation that removes every piece of reusable state, and recovery that restores a trusted system rather than a working one.
None of that is new advice. What is new is a reason to reweight it, and a way to check. Take an attack path, write down what one attempt is, count how many attempts the path admits, work out what a failure tells the attacker, test the control against a hundred adaptive attempts instead of one, and then look at what a single success obtains and what has to be invalidated afterwards. The method is designed to find paths whose apparent strength depends on nobody having tried often enough.
That is the assumption worth retiring. It was never a property of the control. It was a property of the attacker's budget.
Appendix A: Retry-sensitive control assessment template
# Retry-sensitive control assessment
Attack path:
Assessed by: Date:
Review due:
Owner (control):
Owner (response):
## Assumptions
Actor class assumed: [commodity | organised criminal | state-nexus | insider]
Automation assumed: [none | scripted | agentic single | agentic multi]
Source distribution assumed: [single | small set | wide distribution]
Target population and size:
Evidence basis for assumptions: [production observation | test | vendor assessment |
intelligence forecast | estimate]
## Attempt budget
Attempt unit (one success-or-fail action):
Admission boundary (deciding component):
Counter owner: [identity | device | tenant | source network | resource | none]
Hard aggregate cap: [yes | no] Value: Period:
Per-source throttle: [yes | no] Value: Period:
Can the counter be reset by the attacker: [yes | no | unknown]
Can the counter be distributed around: [yes | no | unknown]
Scarce resource charged per attempt:
Can that resource be stolen, outsourced or billed to a third party: [yes | no | partly]
Evidence class: [observation | test | vendor assessment |
forecast | assumption | estimate]
## Feedback
Error differentiation exposed:
Timing differentiation exposed:
Observable side effects of a failed attempt:
Observable defensive response:
Offline testing oracle available to attacker: [yes | no | partly]
Attacker adaptation time (failure to changed attempt):
Defender response time (detection to deployed mitigation):
Which is faster: [attacker | defender | unknown]
Evidence class: [observation | test | vendor assessment |
forecast | assumption | estimate]
## Control state
Controls in scope (by class from section 7):
Retry-sensitivity result S(N), attacker success proportion:
N = 1: [value | not measured]
N = 10: [value | not measured]
N = 100: [value | not measured]
N = 1,000: [value | not measured]
Conditions under which S(N) was measured:
Capability assumed:
Target population and test-case count:
Control configuration and version:
Time window:
Identity and source distribution:
Feedback available to the attacker:
Stopping rule:
First-success attempt and outcome sequence (for a single ordered series):
Regression check against earlier variants: [pass | fail | not run]
Legitimate-user cost at current setting:
Evidence class: [observation | test | vendor assessment |
forecast | assumption | estimate]
## Authority
Authority obtained by one success:
Can it mint further credentials: [yes | no]
Can it alter or disable logging: [yes | no]
Can it change policy or raise quotas: [yes | no]
Can it reach data directly: [yes | no]
Blast radius (systems, identities, data):
Evidence class: [observation | test | vendor assessment |
forecast | assumption | estimate]
## State invalidation
Reusable state created by one success:
- Sessions and tokens:
- Keys and credentials:
- Device or application registrations:
- Persistence mechanisms:
- Modified configuration:
Invalidation mechanism for each:
Owner and expected duration:
Verification method:
Last exercised: Result:
Evidence class: [observation | test | vendor assessment |
forecast | assumption | estimate]
## Recovery
Rebuild source: [trusted image | backup | in-place repair]
Backup isolation (separate credentials and network): [yes | no]
Dependency restoration order defined: [yes | no]
Poisoned-state check on restore: [yes | no]
Last restore test: Date: Result:
Time to service availability (measured):
Time to trusted restoration (measured):
Evidence class: [observation | test | vendor assessment |
forecast | assumption | estimate]
## Evidence confidence
Highest-confidence field and basis:
Lowest-confidence field and basis:
Fields requiring measurement before the next review:
Appendix B: Control-durability and recovery checklist
Attempt limits
- Is there a hard aggregate cap as well as a per-source throttle?
- Does any control correlate attempts across identities, devices, tenants and source networks?
- Can the attacker reset a counter by rotating identity or source?
- Are recovery and support channels rate-limited to the same standard as primary authentication?
- Are limits tested at the volume an automated actor could actually generate?
Scarce resources
- What scarce resource does each attempt consume?
- Can it be stolen, outsourced, amortised across targets, or billed to the victim?
- Are spend and usage ceilings set and enforced, including for AI and cloud APIs?
Feedback
- Do error responses distinguish valid from invalid identifiers?
- Are response timings uniform across those cases?
- Does a block message reveal which control fired or why?
- Can an attacker test artifacts against your controls without deploying them?
- Is verbose diagnostic output limited to authenticated support paths?
Durability under adaptation
- Has each significant detection been tested against a series of adaptive variants?
- Are previously blocked variants re-tested after every update?
- Is regression recorded and owned rather than discovered during an incident?
- Is detection weighted towards behaviour and effect as well as artifact form?
- Is variant clustering a maintained capability?
Authority ceilings
- What does one successful credential, process or host actually reach?
- Can any compromised principal mint credentials, alter logging or raise quotas?
- Are cloud and AI API credentials scoped, short-lived and bound to a workload?
- Is egress restricted by default from workloads that do not need it?
- Do high-impact actions require an approval evaluated outside the compromised path?
State elimination
- Is there an inventory of the reusable state a compromise creates on each path?
- Does revocation reach refresh tokens, device registrations, application passwords and federated sessions?
- Are credentials created during an intrusion identified and removed?
- Is the mechanism that steals or mints credentials removed before rotation begins?
- Is completion of invalidation verified rather than assumed?
Response speed
- Is the defender response time measured and compared with attacker adaptation time?
- Is any containment pre-authorised for well-understood conditions?
- Does host isolation also invalidate identity state held on that host?
- Is there a named owner able to act without further approval?
Recovery
- Are backups isolated by separate credentials and network path?
- Are restores tested per service, with a recorded date and result?
- Is restored data checked for attacker-introduced state?
- Is dependency order for identity, keys and directories defined and rehearsed?
- Is time to trusted restoration measured separately from time to service availability?
- Has re-compromise during recovery been exercised at least once?
Evidence discipline
- Is each recorded value labelled as observation, test result, vendor assessment, forecast or estimate?
- Are attribution qualifiers preserved when threat reporting is used internally?
- Are vendor figures reported with their caveats attached?
- Is the assessment re-run when the actor assumptions change rather than only when the system does?
About the author
Jason Doyle writes about reliable software, observability, incident leadership, applied AI, and practical controls for systems that influence human and organisational decisions. He publishes at jasondoyle.ie and can be contacted at [email protected].
References
- Anthropic, Detecting and countering misuse of AI: September 2026, 10 September 2026, https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a597377d3b5/Anthropic-Detecting-and-countering-091026.pdf. Selected vendor-detected cases with no prevalence denominator and no independent counterfactual.
- UK National Cyber Security Centre, Impact of AI on cyber threat from now to 2027, 7 May 2025, https://www.ncsc.gov.uk/report/impact-ai-cyber-threat-now-2027. An intelligence forecast rather than an incident census; technical surprise is expressly possible.
- National Cyber Security Centre Ireland, 2026 NCSC AI Cyber Security Risk Assessment: Public Sector Deployment, 2026, https://www.ncsc.gov.ie/pdfs/NCSC_AI_risk_assessment_2026.pdf. Deployment-risk guidance rather than independent incident evidence.
- Google Threat Intelligence Group, Adversarial Misuse of Generative AI, 29 January 2025, https://cloud.google.com/blog/topics/threat-intelligence/adversarial-misuse-generative-ai. Google service telemetry, with limited visibility into local models and other providers.
- Google Threat Intelligence Group, AI Threat Tracker: Advances in Threat Actor Usage of AI Tools, 5 November 2025, https://cloud.google.com/blog/topics/threat-intelligence/threat-actor-usage-of-ai-tools/. PROMPTFLUX was incomplete, and stolen token provenance was assessed rather than conclusively established.
- Google Threat Intelligence Group, Distillation, Experimentation, and (Continued) Integration of AI for Adversarial Use, 12 February 2026, https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use. A novel malware architecture does not by itself imply a novel offensive objective or strategic effect.
- OpenAI, Disrupting malicious uses of AI: October 2025, 7 October 2025, https://cdn.openai.com/threat-intelligence-reports/7d662b68-952f-4dfd-a2f2-fe55b041cc4a/disrupting-malicious-uses-of-ai-october-2025.pdf. OpenAI observes model interactions and not always what happens after the activity leaves its platform.
- Microsoft, Microsoft Digital Defense Report 2025: Lighting the Path to a Secure Future, October 2025, https://www.microsoft.com/en-us/security/security-insider/threat-landscape/microsoft-digital-defense-report-2025. Mixes Microsoft telemetry, external research and strategic interpretation.
- Cybersecurity and Infrastructure Security Agency, CISA Vulnerability Review: Fiscal Years 2024 and 2025, August 2026, https://www.cisa.gov/sites/default/files/2026-08/cisa-vulnerability-review-fy-2024-2025-508.pdf. Does not quantify AI prevalence or provide named AI-enabled cases.
- European Union Agency for Cybersecurity, ENISA Threat Landscape 2025, 1 October 2025, revised 9 January 2026, https://www.enisa.europa.eu/publications/enisa-threat-landscape-2025. Aggregates vendor and open-source reporting, so some figures are derivative.
- Europol, IOCTA 2026: The Evolving Threat Landscape, 28 April 2026, https://www.europol.europa.eu/cms/sites/default/files/documents/IOCTA-2026.pdf. Combines member-state reporting, industry intelligence and forward-looking assessment.
- National Security Agency, Cybersecurity and Infrastructure Security Agency and Federal Bureau of Investigation, China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies, advisory AA26-251A, 8 September 2026, https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a. Unauthorised pathways include shared subscriptions and proxies as well as stolen API keys.
- Chad Heitzenrater, The Winning Economics of Cybersecurity in an Age of Advanced Artificial Intelligence, RAND Corporation, August 2025, https://www.rand.org/pubs/perspectives/PEA3691-11.html. A strategic perspective rather than an empirical estimate of attacker or defender costs.
- Benjamin Sperisen, Jair Aguirre, Henri van Soest and Zylex Lopez, AI Agents Put Offensive Cyber Within Reach of Novices, RAND Corporation, June 2026, https://www.rand.org/pubs/research_reports/RRA3892-2.html. Controlled challenges do not reproduce production intrusion, detection or legal constraints.
- Vaibhav Garg and Jayati Dev, Artificial Intelligence and the New Economics of Cyberattacks, USENIX ;login:, 29 August 2024, https://www.usenix.org/publications/loginonline/artificial-intelligence-and-new-economics-cyberattacks. Conceptual analysis predating the September 2026 incident evidence.
- Gary S. Becker, Crime and Punishment: An Economic Approach, Journal of Political Economy, March 1968, https://doi.org/10.1086/259394. General economic theory; state, ideological and destructive actors may not maximise financial return.
- Stuart E. Schechter and Michael D. Smith, How Much Security Is Enough to Stop a Thief?, Financial Cryptography, 2003, https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/fc03.pdf. A simplified model of a financially motivated outsider.
- Martin C. Libicki, Lillian Ablon and Tim Webb, The Defender's Dilemma: Charting a Course Toward Cybersecurity, RAND Corporation, June 2015, https://www.rand.org/content/dam/rand/pubs/research_reports/RR1000/RR1024/RAND_RR1024.pdf. Heuristic analysis based partly on a small set of interviews.
- Ross Anderson, Why Information Security Is Hard: An Economic Perspective, ACSAC, December 2001, https://www.acsac.org/2001/papers/110.pdf. Foundational theory rather than a contemporary intervention study.
- Chris Kanich, Christian Kreibich, Kirill Levchenko, Brandon Enright, Geoffrey M. Voelker, Vern Paxson and Stefan Savage, Spamalytics: An Empirical Analysis of Spam Marketing Conversion, ACM CCS, 27 October 2008, https://doi.org/10.1145/1455770.1455774. One historical botnet with only 28 observed purchases.
- Cynthia Dwork and Moni Naor, Pricing via Processing or Combatting Junk Mail, CRYPTO 1992, https://www.wisdom.weizmann.ac.il/~naor/PAPERS/pvp.pdf. Hardware inequality, stolen compute and legitimate-user burden can defeat the economics.
- Joseph Bonneau, The Science of Guessing: Analyzing an Anonymized Corpus of 70 Million Passwords, IEEE Symposium on Security and Privacy, May 2012, https://jbonneau.com/doc/B12-IEEESP-analyzing_70M_anonymized_passwords.pdf. Password-specific; distributions change by population and authentication design.
- National Institute of Standards and Technology, SP 800-63B-4: Authentication and Authenticator Management, July 2025, https://pages.nist.gov/800-63-4/sp800-63b.html. Per-authenticator limits do not by themselves stop distributed or recovery-channel attacks.
- Ron Ross, Victoria Pillitteri, Richard Graubart, Deborah Bodeau and Rosalie McQuaid, SP 800-160 Volume 2 Revision 1: Developing Cyber-Resilient Systems, National Institute of Standards and Technology, 8 December 2021, https://doi.org/10.6028/NIST.SP.800-160v2r1. Techniques can conflict and carry operational and architectural costs.
- Eric M. Hutchins, Michael J. Cloppert and Rohan M. Amin, Intelligence-Driven Computer Network Defense Informed by Analysis of Adversary Campaigns and Intrusion Kill Chains, Lockheed Martin, 2011, https://www.lockheedmartin.com/content/dam/lockheed-martin/rms/documents/cyber/LM-White-Paper-Intel-Driven-Defense.pdf. A vendor-authored linear model that underrepresents concurrent and looping operations.
- National Institute of Standards and Technology, SP 800-61 Revision 3: Incident Response Recommendations and Considerations for Cybersecurity Risk Management, 3 April 2025, https://doi.org/10.6028/NIST.SP.800-61r3. General guidance that does not prescribe universal detection or recovery thresholds.
- Alessandro Brucato and the Sysdig Threat Research Team, LLMjacking: Stolen Cloud Credentials Used in New AI Attack, 6 May 2024, https://www.sysdig.com/blog/llmjacking-stolen-cloud-credentials-used-in-new-ai-attack. The cost figure is an estimated worst case, and some illustrative proxy infrastructure was not attributed to the observed attacker.
- Unit 42, Palo Alto Networks, Token Jacking: Cybercriminals Could Be Stealing Your AI Resources, 6 August 2026, https://unit42.paloaltonetworks.com/ai-token-jacking/. Vendor incident-response reporting with victim identities and full forensic timelines withheld.
- Mike Ulrich, Addressing Cascading Failures, and Alejandro Forero Cuervo, Handling Overload, in Google Site Reliability Engineering, 2016, https://sre.google/sre-book/addressing-cascading-failures/ and https://sre.google/sre-book/handling-overload/. Retry budgets and throttling address cooperative distributed-system amplification rather than an adversary attempting to evade the counter.
- John Hughes et al., Best-of-N Jailbreaking, arXiv:2412.03556, December 2024, https://arxiv.org/abs/2412.03556. Repeated prompt sampling is a model-safety experiment and does not estimate success rates for production cyber intrusions.
- Mark Russinovich, Building security that lasts: Microsoft's journey towards durability at scale, Microsoft Security Blog, 26 June 2025, https://www.microsoft.com/en-us/security/blog/2025/06/26/building-security-that-lasts-microsofts-journey-towards-durability-at-scale/. Microsoft's model focuses on security-baseline regression and organisational enforcement rather than retry sensitivity under adversarial adaptation.