Sponsored

The UK’s AI Security Institute has put a number on a risk that finance chiefs have mostly discussed in the abstract. In late 2023, the best frontier models could complete apprentice-level offensive-cyber tasks about 9 per cent of the time. Today, according to the institute’s first Frontier AI Trends Report, that figure is 50 per cent. In 2025, AISI tested the first model able to complete cyber tasks designed for experts with more than ten years’ experience. The capability is not plateauing. It is compounding.

The detail behind the headline is what should concentrate minds in bank and insurer control functions. On a 32-step simulated attack against an enterprise network, a task AISI estimates would take a human expert around 14 hours to run end to end, the best-performing system it tested, Anthropic’s Claude Opus 4.6 released in February 2026, averaged 15.6 steps, roughly six of those fourteen hours of expert work. The National Cyber Security Centre, drawing on the same evaluations, notes that the leading model in early 2026 completed nearly six times as many attack steps as the best model eighteen months earlier, and that a full run at that simulated intrusion now costs about £65.

That £65 figure is the one that collapses an assumption UK finance has quietly relied on. Sophisticated, multi-step intrusions have historically been rationed by cost, scarce expertise and time. When a chunk of that workflow can be automated for the price of a team lunch, the economics of who can mount a credible attack, and how often, change. AISI is careful to note the models it tested still generate detectable noise and would, in a well-monitored environment, likely be spotted and disrupted before achieving significant progress. That caveat is the point: the defensive margin now depends on monitoring and response that actually works, not on the attacker lacking capability.

The preparation window is also shrinking. In a separate evaluation published on 17 July, AISI found that recent open-weight models trail closed frontier systems on cyber capability by four to seven months, narrowing from the six to ten months it measured through most of 2025. Once a capability is baked into a downloadable open-weight model, it is effectively ungoverned and available at scale. The lag between what sits behind a lab’s safety controls and what anyone can run on rented hardware is closing.

For UK financial services, this is where capability research meets the supervisory rulebook, and the fit is uncomfortably direct. The operational-resilience framework the Bank of England, PRA and FCA finalised in 2021 required firms to identify their important business services, set impact tolerances for how long disruption can be tolerated, and prove they can stay within those tolerances through severe but plausible scenarios. Full compliance became mandatory in March 2025. A “severe but plausible” cyber scenario is not a static benchmark. If frontier models are compressing a fourteen-hour expert intrusion into six, and doing it for £65, the plausible scenario a board must plan against is faster, cheaper and more repeatable than the one many resilience self-assessments were built on.

The same logic runs through threat-led penetration testing. The CBEST regime and its intelligence-led equivalents were designed to test firms against realistic adversary behaviour. Realistic now includes AI-assisted reconnaissance and exploitation that moves at machine speed. Supervisors will reasonably ask whether a firm’s testing, patch cadence and detection tooling have been recalibrated for a threat that AISI’s own numbers show doubling back on itself every few evaluation cycles, rather than the slower curve assumed when many of these controls were designed.

None of this requires a new AI rulebook, and that is the pragmatic edge of the UK approach. The expectations already exist. What has changed is the threat model they are measured against. A bank that has deferred technology remediation, tolerated slow patching or lost visibility of its software inventory has not broken a specific rule. It has simply made itself the cheapest target on a curve that is bending the wrong way.

The NCSC’s guidance to defenders is not to abandon AI but to match it. Its “cyber defenders need to be ready for frontier AI” work sets out three uses that map neatly onto a finance operating model: hardening systems before they are probed, accelerating threat detection and investigation, and automating parts of mitigation and response so defenders can operate at comparable speed to an AI-driven attacker. That is attractive to boards and awkward for supervisors, because AI-enabled defence adds its own model risk, false confidence and dependency on tools few directors fully understand. The honest position is that automated defence is becoming necessary without being sufficient. Basic hygiene, asset inventory, patching, monitoring and tested recovery, still decides whether the detectable attacker is actually detected.

For investors, the read-across is prosaic but material. AI cyber readiness is quietly becoming part of franchise quality. A lender that has spent years postponing core remediation may find frontier AI turns technical debt into a supervisory and capital-markets story. An insurer underwriting cyber risk faces the same moving baseline in its own book and its clients’. A market-infrastructure provider that cannot describe its detection and response with confidence will struggle to call itself resilient, whatever its impact-tolerance documents say.

The value of AISI publishing the numbers is that it removes the excuse of abstraction. The question for UK finance is no longer whether frontier AI will eventually matter to cyber risk. It is whether the institution can keep operating when a fourteen-hour expert attack costs £65 and arrives at machine speed, and whether the board can prove it, before a supervisor asks.

Finance & Markets Correspondent
Covers: Finance, capital markets, technology investing

David Whitmore covers the intersection of capital and code — the funding rounds, market structures and policy moves that shape how money flows through the technology economy.