StandardsAboutContact
The Weights
UK's AI Security Institute Finds Open-Weight Models Are

UK's AI Security Institute Finds Open-Weight Models Are

New AISI benchmarking puts open-weight models like GLM-5.2 within four to seven months of frontier closed systems on cyber tasks, at a fraction of the cost — landing days before Demis Hassabis floated a pre-release review body that, as proposed, wouldn't cover open-weight releases at all.

UK government testing confirms the open-weight/closed-weight cyber-capability gap has narrowed to four-to-seven months, down from six-to-ten months in 2025: GLM-5.2 and DeepSeek V4-Pro now approach frontier models like Claude Opus 4.6 on narrow tasks, though closed models still lead on complex, chained attacks, and open models cost up to 45 times less per task.

The Weights Desk · 4 min read

The UK's AI Security Institute (AISI) has published benchmark results showing that open-weight AI models have closed most of the cyber-capability gap with closed, frontier systems: models like GLM-5.2 and DeepSeek V4-Pro now trail top closed models by only four to seven months, down from six to ten months measured in 2025 internal testing. The finding lands in the same news cycle as Google DeepMind CEO Demis Hassabis's proposal for a FINRA-style federal body to review frontier models before release — a proposal that, as posted, would not cover open-weight releases.

The Numbers Behind the Gap

AISI tested open-weight GLM-5.2 (June 2026) and DeepSeek V4-Pro against closed comparators Opus 4.6, Opus 4.5, GPT-5.3-Codex, and Sonnet 4.5 across 70 narrow cyber tasks spanning four difficulty tiers, plus end-to-end 'cyber range' simulations including a 32-step corporate-network scenario called The Last Ones. On the narrow tasks, GLM-5.2 matched Opus 4.6, a model released just 4.3 months earlier — the tightest gap AISI has measured to date.

Where the Gap Still Holds

The convergence is not uniform. On cyber-range tasks — which chain multiple steps into an autonomous attack rather than testing an isolated skill — GLM-5.2 reached only the level of Opus 4.5, a model roughly seven months older than the comparison point used for narrow tasks. AISI's data suggests open-weight models are catching up fastest on discrete technical skills, while sustained, multi-step operational tradecraft — the kind an actual intrusion requires — remains harder to replicate outside a closed lab.

Cost Changes the Threat Model

AISI also measured cost per task: DeepSeek V4-Pro completed evaluation tasks at roughly $0.28 each, against $12.50 for Opus 4.5 — a roughly 45-fold difference. A four-to-seven-month capability lag matters less to a threat model if the lagging model is also nearly free to run at scale; cheap, downloadable weights change who can attempt an attack, not just what the single best available model can do.

The Governance Response Already on the Table

Days after AISI's data circulated, Hassabis proposed a federally overseen Standards Body, modeled on FINRA, under which frontier labs would voluntarily submit models for national-security review up to 30 days before release. As posted, the plan is voluntary and aimed at closed frontier labs; it does not address open-weight releases like GLM-5.2 or DeepSeek V4-Pro, which is precisely where AISI's data shows the fastest capability convergence and the least premarket oversight today.

The Verdict

AISI's benchmark is the confirmed part of this story: a measured, narrowing capability gap and a real cost asymmetry, from a government lab whose job is exactly this kind of testing. Hassabis's Standards Body is not confirmed policy — it is one executive's public, non-binding proposal, and it is silent on open-weight models. Read together, the finding outpaces the response: the pre-release-review architecture being floated does not yet reach the release channel AISI's own data flags as closing fastest.

What exactly did the UK AI Security Institute measure?
AISI ran 70 narrow cyber tasks across four difficulty levels plus autonomous 'cyber range' attack simulations — including a 32-step scenario called The Last Ones — on open-weight models GLM-5.2 and DeepSeek V4-Pro, benchmarked against closed models Opus 4.6, Opus 4.5, GPT-5.3-Codex, and Sonnet 4.5.
Does Demis Hassabis's proposed Standards Body apply to open-weight models like GLM-5.2?
No. As posted, the FINRA-style review body is voluntary and framed around frontier labs submitting models before release; it does not name or bind open-weight releases, which is exactly where AISI's own data shows the capability gap closing fastest.
Is the open/closed capability gap closing uniformly across all task types?
No. It is closing faster on narrow, discrete technical tasks than on complex, multi-step cyber-range attacks, where closed frontier models retain roughly a seven-month lead.
  1. Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan — Import AI
  2. How far behind the frontier are leading open-weight models on cyber? — UK AI Security Institute (AISI)
  3. Proposal for a FINRA-style Standards Body for frontier AI review — X (Demis Hassabis)