Red does security work others refuse
Frontier models refuse authorized offensive work because refusal is safety alignment trained into their weights. Red does the work and runs sovereign or fully air-gapped in your own infrastructure.
Runs your security work
93%
Next-highest domain
56%
Capability cost
≤0.028 MMLU
Cracken Red runs 93% of cybersecurity work across 306 prompts; no other harm domain runs above 56%. arXiv:2607.02714v2, NeurIPS 2026 under review.
Red runs on your infrastructure
You cannot turn their refusal off, and it tightens with every release. An operation built on a permissive frontier model is one release away from stopping.
| Capability | GPT-5.6 Sol | Fable 5 | Mythos 5current · limited release | Cracken Red |
|---|---|---|---|---|
Takes authorized offensive-security work Share of authorized offensive-security work the model will take — rivals on Cybench, 117 sample-epochs; ours on a 306-prompt cross-evaluation set | 9.4% 11 of 117 epochs; it refused the other 106. | 0% Refused all 117 of 117 epochs. | Not purchasable Limited release through Project Glasswing and the Cyber Verification Program. | 93% Up from 0% before the technique. |
Solves a capture-the-flag challenge end to end Cybench — 39 professional CTF challenges over 117 sample-epochs, refusal scored separately from capability | — No solve rate reported. | 0% solved The authors: "a policy-filter outcome rather than evidence of zero underlying capability." | — Not run. | Runs it* No safety alignment between the operator and the task. |
Turns a real vulnerability into a working exploit — and lets you run it ExploitBench — 41 V8 environments run as an exploit-development pipeline, safeguards off | — Not this column: 73.5 on OpenAI's own ExploitBench harness, a different instrument. | ≈40 delivered Same weights as Mythos 5 with cyber safeguards on. Anthropic’s own footnote: "performs closer to Claude Opus 4.8 due to fallbacks." | 78 · 0 available Not purchasable at any price — limited release. | Runs it · 100% available Exploit development is what it is bought for, on hardware you control. |
Can a regulated buyer run it at all Deployment and data handling, from each vendor's own published terms. | — No self-hosted or in-enclave option published. | 30-day retention Mandatory, with no zero-retention option. | Not purchasable Limited release through Project Glasswing and the Cyber Verification Program. | In your enclave Single tenant, zero retention on every plan, granted by name and scope. |
Takes authorized offensive-security work
Share of authorized offensive-security work the model will take — rivals on Cybench, 117 sample-epochs; ours on a 306-prompt cross-evaluation set- GPT-5.6 Sol
9.4%
11 of 117 epochs; it refused the other 106.- Fable 5
0%
Refused all 117 of 117 epochs.- Mythos 5
Not purchasable
Limited release through Project Glasswing and the Cyber Verification Program.- Cracken Red
93%
Up from 0% before the technique.
Solves a capture-the-flag challenge end to end
Cybench — 39 professional CTF challenges over 117 sample-epochs, refusal scored separately from capability- GPT-5.6 Sol
—
No solve rate reported.- Fable 5
0% solved
The authors: "a policy-filter outcome rather than evidence of zero underlying capability."- Mythos 5
—
Not run.- Cracken Red
Runs it*
No safety alignment between the operator and the task.
Turns a real vulnerability into a working exploit — and lets you run it
ExploitBench — 41 V8 environments run as an exploit-development pipeline, safeguards off- GPT-5.6 Sol
—
Not this column: 73.5 on OpenAI's own ExploitBench harness, a different instrument.- Fable 5
≈40 delivered
Same weights as Mythos 5 with cyber safeguards on. Anthropic’s own footnote: "performs closer to Claude Opus 4.8 due to fallbacks."- Mythos 5
78 · 0 available
Not purchasable at any price — limited release.- Cracken Red
Runs it · 100% available
Exploit development is what it is bought for, on hardware you control.
Can a regulated buyer run it at all
Deployment and data handling, from each vendor's own published terms.- GPT-5.6 Sol
—
No self-hosted or in-enclave option published.- Fable 5
30-day retention
Mandatory, with no zero-retention option.- Mythos 5
Not purchasable
Limited release through Project Glasswing and the Cyber Verification Program.- Cracken Red
In your enclave
Single tenant, zero retention on every plan, granted by name and scope.
- Cybench refusal study, arXiv:2607.15263v3, and Cracken's own cross-evaluation, arXiv:2607.02714v2 — both retrieved 3 August 2026. https://arxiv.org/abs/2607.15263
- Same study. Anthropic declined to report Cybench for its own current models, calling it largely saturated.
- Anthropic, Fable 5 / Mythos 5 system card — retrieved 3 August 2026
- Anthropic platform docs and Project Glasswing pages — retrieved 3 August 2026.
Competitor figures come from Cybench and each vendor's own system card; ours from our 306-prompt cross-evaluation set — different instruments, so read each against its own. * Independent runs pending.
Leave it running.
Anthropic's own system card: on ExploitBench, Fable 5's cyber safeguards flagged 407 of 410 episodes, an average of 27 turns in. Those 27 turns are generated and billed at $10 and $50 per million before the fallback fires, and what comes back is Opus 4.8-grade work — 40, not 78. On Cybench it refused 117 of 117 epochs. Cost per delivered result there is not high. It is undefined.
- Cracken Max
Deepest reasoning per step, for the targets the others stall on.
84–367
credits an assessment
- Cracken Smart
The general frontier tier most operations run on today: recon, exploitation and reporting.
50–220
credits an assessment
- Cracken Red
0.29× on Cracken's published rate card, with no 27-turn fallback loop to pay for.
14–64
credits an assessment
A credit is Cracken's billing unit and an assessment is one scoped engagement, so these are per-engagement ranges, not token prices — the same figures as /pricing. Smart and Max resolve to whichever frontier model measures best that month. For a buyer whose traffic cannot leave the enclave, the frontier's number is 0.
Cyber runs at 93%. Nothing else clears 56%.
On a 306-prompt set, Cracken Red runs 93% of cybersecurity work. No other harm domain runs above 56%.
Cybersecurity
The one domain we opened — 93% against a next-highest 56%. Cross-evaluation set, 306 prompts — ours93%
Privacy violation
56%
Illegal goods
44%
Violence
25%
Misinformation
12%
Explicit content
0%
A public broad-spectrum jailbreak — runs every other harm too
Not our measurement and not on this set — an independent reviewer's characterisation of broad-spectrum jailbroken models, stated there as 0–20% refusal.80–100%
306 prompts, α = 1.0, a 31-pattern detector. An evasive non-answer scores as compliance, so every run rate here is an upper bound — the gap between cyber and the rest is if anything wider.
- Cybersecurity
0.670
- Explicit content
0.784
- Misinformation
0.806
- Illegal goods
0.814
- Privacy violation
0.816
- Violence
0.826
Cyber's refusal is not the same object as the other five, which is what makes it removable alone. The separation shows on the first principal component even where removal then fails.
No cyber-open weights, ever
The cyber extraction dataset is withheld in full
Code and eval framework are gated
No benchmark prompts are redistributed
Cyber refusal differs across 24 models
- 01
Refusal is a direction
Arditi et al., 2024: refusal is a single direction in a model's activations. Project it out and every refusal goes with it — one switch, all or nothing.
- 02
A cone, not a line
Wollschläger et al., ICML 2025: refusal is a multi-dimensional subspace inside each layer. Ours: it spreads across layers too, so Arditi's middle-layer recipe does not generalise — a uniform spread beats norm-based selection by up to 70pp at 30% of depth.
- 03
Cyber comes out alone
Ours: each harm domain carries its own direction. Cyber's is the least like the rest, and at trillion-parameter scale it comes out while the other five stand. 24 open-weight models, 0.6B to 1T, dense and mixture-of-experts, six intensities, about 600 GPU-hours.
DPO
~30pp median
RLHF
~22pp
PPO + DPO
20–25pp
SFT only
~19pp
BOND + WARM + WARP
48pp at 1B, 13pp at 27B
GRPO
5–8pp
Undisclosed RL
<1pp
The safety-training recipe decides whether this is possible, which tells a safety team which of their own recipes survive a representation-level attack.
The study publishes in full. The cyber extraction dataset does not, and will not.
Where it fails.
More willing is measured. Better is not, and two model families barely moved at all.
Does this make an operator better at offensive work, or only more willing?
More willing is measured; better is not. The technique removes refusal, not incapability — it adds no exploit-development skill the base model lacks. In-domain uplift, offensive or defensive, is the one number Cracken does not have; the benchmark for it is being built and publishes when it exists. *pending validation — arXiv:2607.02714v2 — limitations and future work (v2, arXiv, 2026-07)
Your refusal metric is string matching. Why should I trust it?
Partly, and it is the conservative half. The detector is 31 string patterns, a superset of the 12 Arditi et al. (2024) published and flagged as limited, and it is conservative by construction: an evasive non-answer counts as compliance, so it can only understate selectivity. An LLM judge then re-scored every flipped response — 26 of 30 cyber flips (87%) are substantive answers, and only 2 of 9 off-target flips (22%), the rest evasive.
Where does the technique fail outright?
Two mixture-of-experts models moved 3pp or less, and one 123B model moved 0.0pp — nothing shifted at all. The effect is not monotonic in size: within one family, 8B gained 38.9pp while 14B gained only 1.4pp. Safety-training recipe predicts susceptibility; parameter count does not. — arXiv:2607.02714v2 — per-model results (v2, arXiv, 2026-07)
What is the extraction set, and what does it not cover?
English only, capture-the-flag skewed, and applied at 30% of layers or less. Nothing here has been measured outside those bounds. — arXiv:2607.02714v2 — limitations (v2, arXiv, 2026-07)
Every number, and every boundary.
- Deployment
Single-tenant, inside your enclave. The request does not leave it and neither does the output.
- Data retention
Zero, on every plan. Fable 5 ships only with a mandatory 30-day retention policy and no zero-retention option.
- Access
Per tenant, under a written authorization: authorization on file, a named scope and target estate, then enablement. Granted by name and scope, not by API key.
- Weights
Never released. Not to a customer, not to a researcher, not under NDA.
- Base model
Not disclosed. Naming it would hand the extraction recipe a starting point; the technique is published, the target is not.
- Method
Domain-specific refusal-direction removal. arXiv:2607.02714v2.
- Evaluation sets · published research build
306-prompt cross-evaluation set across six harm domains, α = 1.0. 360-prompt held-out set: HarmBench 185, AdvBench 119, CyberSecEval 56.
- Capability cost · published research build
MMLU held. Worst move across 24 models and six removal intensities was 0.028; 23 of the 24 moved 0.006 or less.
- Over-refusal on benign prompts
+1.6pp at worst.
- Offensive ceiling
Inherited from the base model. The technique removes refusal, not incapability — it adds no exploit-development skill the base lacks.
- In-domain capability uplift
Not measured. *pending validation
- Language
English.
- Scope
Authorized offensive operations. Not threat intelligence, not forensics, not detection engineering.
Verified buyers only
Written authorization
Written authorization to attack the estate you name, on file before anything is switched on. No authorization, no model. Everyone else is better served by the model that refuses.
A named scope and a target estate
Not a setting you send us once: it is what the platform scopes the engagement to, and what every operation is recorded against.
Your own tenant
Enabled per tenant, never per key. The weights are never handed over — the endpoint is.

