- Benchmark Gap: The UK AI Security Institute found open-weight models GLM-5.2 and DeepSeek V4-Pro matched closed-model cyber tests from four to seven months earlier.
- Run Costs: Advertised 100-million-token range-run costs fell from about $85 for Claude Opus to $1.19 for DeepSeek V4-Pro.
- Control Limits: Downloaded weights make monitoring, user bans, model withdrawal, and refusal safeguards harder to preserve across privately run copies.
- Evidence Boundary: The strongest comparison uses 70 narrow tasks, while multi-step evidence is thinner and neither test represents a defended live network.
AISI’s July 2026 assessment found that GLM-5.2, released in June 2026, and DeepSeek V4-Pro performed similarly to closed models released four to seven months earlier. Its controlled cyber tests did however not involve simulated live attacks.
AISI, the UK government body that evaluates advanced AI risks, compared downloadable models with older hosted systems, not the newest frontier systems.
Through much of 2025, the measured delay was six to ten months. At advertised first-party prices, a fixed 100-million-token run cost about $85 with Anthropic’s hosted Claude Opus 4.6 or its predecessor and an estimated $46 with Zhipu AI’s GLM-5.2. DeepSeek V4-Pro’s $1.19 fixed-workload total used the model’s standing API price.
An open-weight model exposes learned parameters that users can run privately, modify, and redistribute. Lower processing costs and a shorter measured delay reduce defenders’ preparation time before similar tested capabilities reach downloadable systems. Providers also cannot keep monitoring, user bans, or model withdrawal attached to every copy running on someone else’s hardware.
What the Cyber Tests Measure
AISI’s stronger comparison uses 70 tasks across four difficulty levels. Its suite covers vulnerability research and exploitation, reverse engineering, web exploitation, and cryptography. Each task allows five attempts and 2.5 million processing tokens, meaning the amount of model processing purchased while it works through a task.
GLM-5.2 matched Claude Opus 4.6 and OpenAI’s coding-focused GPT-5.3-Codex on that suite. DeepSeek V4-Pro produced results comparable to Claude Opus 4.5. Claude Opus 4.6 separately has a one-million-token context window in beta, while its reported Terminal-Bench 2.0 score was 65.4%; neither specification measures the cyber result.
AISI uses the Last Ones cyber range to test whether an agent can sustain a multi-step cyber attack scenario. It simulates a 32-step intrusion across four corporate-network subnets and about 20 hosts. A human expert would need an estimated 20 hours to complete the scenario.
GLM-5.2 reached about as far as Anthropic’s earlier Opus 4.5 model on the range, while DeepSeek V4-Pro remained below Sonnet 4.5. Fewer ranges underpin that result, making it weaker comparative evidence than the broad narrow-task suite. One test samples many isolated skills; the other tests whether an agent can maintain a chain of actions.
Price comparisons also depend on their denominator. AISI calculated per-task costs among tasks completed with 100% reliability: Opus 4.6 cost $15.17 and GLM-5.2 cost $6.12. Opus 4.5 cost $12.50 per reliably completed task, compared with $0.28 for DeepSeek V4-Pro.
Reliability means the share of repeated attempts completed successfully. Per-task values include only attempts completed at that level, not the price of a complete range run. Purchased processing allowance can change how far an agent advances, so comparisons must keep the token budget attached.
DeepSeek V4-Pro occasionally refused narrow reverse-engineering tasks, but a small number of repeated attempts bypassed those refusals during the assessment; this controlled observation does not establish a universal safeguard bypass. Downloaded weights still make a refusal less durable because users can copy and modify the model beyond its original hosted service.
A Smaller Window Does Not Equal a Live Attack
AISI’s multi-step range comparisons measure completed attack chains. Separately, from late 2024 through February 2026, the institute estimated that the 80%-reliability cyber-task time horizon doubled every 4.7 months under a 2.5-million-token-per-task cap. An earlier estimate in November 2025 put the doubling time at eight months.
Faster benchmark progress and a four-to-seven-month open-model lag put pressure on defensive planning. However, neither measure predicts when a model could compromise a defended production network. Humans and AI systems can struggle with different tasks, while the longest tasks have relatively few human baselines.
Simulated ranges also omit active defenders who detect, interrupt, and adapt to an intrusion. Different task difficulty, sparse human baselines, and absent defenders make the tests useful for tracking controlled change, not forecasting a breach date. Claude Mythos Preview and GPT-5.5 exceeded AISI’s fitted trend, although available evidence does not establish sustained acceleration.


