Simbian
Save my seat
Live Webinar

The LLM Security Test:
Three Open-Weight Models Beat Claude.

GLM-5.2, DeepSeek 4 Pro and Kimi K3 each beat Claude Opus 4.8 on 100,000 real attack events. Not one of the 25 models tested reached the pass mark. Here is the full scoreboard, and the selection test to run before your next model decision.

calendar_today

Date

Sept 29, 2026

schedule

Time

9:30 AM PDT

language

Your Time

Ambuj Kumar

Ambuj Kumar

Founder & CEO

Alankrit Chona

Alankrit Chona

Chief Technology Officer

Secure Your Seat

Open weight models vs. frontier, scored on defense

25 models dropped into 100,000+ real attack events with no alert, no hint and no question. Scored on how much of a multi-stage intrusion each one reconstructs.

Open weight Frontier
Attack-chain coverage
PASS MARK 50%
1
Opus 5
Anthropic
45.1%
$3.76
2
Opus 4.6
Anthropic
44.5%
$2.71
3
Grok 4.6
xAI
41.3%
$1.58
4
GPT 5.6 Sol
OpenAI
41.0%
$4.03
5
GPT 5.5
OpenAI
37.4%
$4.50
6
Sonnet 4.6
Anthropic
36.9%
$2.21
7
GLM 5.2
Open weight
35.0%
$2.01
8
DeepSeek 4 Pro 0813
Open weight
34.0%
$0.97
9
Kimi K3
Open weight
33.0%
$2.38
10
Opus 4.8
Anthropic
32.7%
$1.68
11
Opus 4.7
Anthropic
30.9%
$1.65
12
DeepSeek 4 Flash 0731
Open weight
29.7%
$0.17
13
Qwen 3.8 Max
Open weight
29.5%
$1.59
14
Gemini 3.7 Flash
Google
29.1%
$0.82
15
Kimi 2.7 Code
Open weight
26.7%
$1.51
16
GPT 5.6 Terra
OpenAI
22.5%
$0.82
lock

13 more models

All 25 models scored on attack-chain coverage and median cost per hunt, weights marked open or closed — on the public Cyber Defense Benchmark.

View the Full Leaderboard

Coverage is the mean fraction of the attack chain reconstructed across 1,080 runs. Cost is median USD per hunt. Passing requires 50%+ coverage on every one of 13 MITRE tactics; the leader clears it on 7. Source: arXiv 2604.19533.

What the LLM security leaderboard shows

1,080 runs against deterministic ground truth, no LLM-as-judge. The LLM cybersecurity benchmark behind arXiv 2604.19533.

0

Of 25 models passed

3

Open-weight models above Opus 4.8

100K+

Attack events per environment

How to score open weight models for the SOC

bug_report

Why the intelligence index misleads

Qwen 3.8 27B and GPT-5.6 Luna both score 52 on Artificial Analysis. On real attack logs, one reconstructs a third less than the other. General capability predicts nothing here.

terminal

The defensive test you can run in a week

What to substitute for the leaderboard score everyone quotes, using your own logs and the models you are already paying for.

lock_open

Where self-hosting triples the bill

Per-token math is what makes a local model look cheap. Price per investigation instead and the reasoning tokens surface immediately.

key

Score coverage, not confidence

Every model tested called its investigation finished before exhausting the hypothesis space. Several did it while still under their query budget.

hub

Frontier took the podium. Open weight took everything else.

GLM 5.2, DeepSeek 4 Pro and Kimi K3 land 7th, 8th and 9th of 25, ahead of Opus 4.8, Opus 4.7, Sonnet 5, GPT-5.6 Luna and every Gemini tested.

verified

The open vs. closed leaderboard, yours to keep

25 models scored on attack-chain coverage and median cost per hunt, weights marked open or closed, sourced to arXiv 2604.19533 and built for a procurement conversation.

Get the open vs. closed LLM security leaderboard

Live on 29 September with the team behind the Cyber Defense Benchmark. Bring the model shortlist you are about to sign off.

Save my seat