The LLM Security Test:
Three Open-Weight Models Beat Claude.
GLM-5.2, DeepSeek 4 Pro and Kimi K3 each beat Claude Opus 4.8 on 100,000 real attack events. Not one of the 25 models tested reached the pass mark. Here is the full scoreboard, and the selection test to run before your next model decision.
Ambuj Kumar
Founder & CEO
Alankrit Chona
Chief Technology Officer
Secure Your Seat
Open weight models vs. frontier, scored on defense
25 models dropped into 100,000+ real attack events with no alert, no hint and no question. Scored on how much of a multi-stage intrusion each one reconstructs.
13 more models
All 25 models scored on attack-chain coverage and median cost per hunt, weights marked open or closed — on the public Cyber Defense Benchmark.
View the Full LeaderboardCoverage is the mean fraction of the attack chain reconstructed across 1,080 runs. Cost is median USD per hunt. Passing requires 50%+ coverage on every one of 13 MITRE tactics; the leader clears it on 7. Source: arXiv 2604.19533.
What the LLM security leaderboard shows
1,080 runs against deterministic ground truth, no LLM-as-judge. The LLM cybersecurity benchmark behind arXiv 2604.19533.
Of 25 models passed
Open-weight models above Opus 4.8
Attack events per environment
How to score open weight models for the SOC
Why the intelligence index misleads
Qwen 3.8 27B and GPT-5.6 Luna both score 52 on Artificial Analysis. On real attack logs, one reconstructs a third less than the other. General capability predicts nothing here.
The defensive test you can run in a week
What to substitute for the leaderboard score everyone quotes, using your own logs and the models you are already paying for.
Where self-hosting triples the bill
Per-token math is what makes a local model look cheap. Price per investigation instead and the reasoning tokens surface immediately.
Score coverage, not confidence
Every model tested called its investigation finished before exhausting the hypothesis space. Several did it while still under their query budget.
Frontier took the podium. Open weight took everything else.
GLM 5.2, DeepSeek 4 Pro and Kimi K3 land 7th, 8th and 9th of 25, ahead of Opus 4.8, Opus 4.7, Sonnet 5, GPT-5.6 Luna and every Gemini tested.
The open vs. closed leaderboard, yours to keep
25 models scored on attack-chain coverage and median cost per hunt, weights marked open or closed, sourced to arXiv 2604.19533 and built for a procurement conversation.
Get the open vs. closed LLM security leaderboard
Live on 29 September with the team behind the Cyber Defense Benchmark. Bring the model shortlist you are about to sign off.
