// Cracken Red

Red does security work others refuse

Frontier models refuse authorized offensive work because refusal is safety alignment trained into their weights. Red does the work and runs sovereign or fully air-gapped in your own infrastructure.

Read the paper

Runs your security work

93%

Next-highest domain

56%

Capability cost

≤0.028 MMLU

Cracken Red runs 93% of cybersecurity work across 306 prompts; no other harm domain runs above 56%. arXiv:2607.02714v2, NeurIPS 2026 under review.

// The evidence

Red runs on your infrastructure

You cannot turn their refusal off, and it tightens with every release. An operation built on a permissive frontier model is one release away from stopping.

  • Takes authorized offensive-security work

    Share of authorized offensive-security work the model will take — rivals on Cybench, 117 sample-epochs; ours on a 306-prompt cross-evaluation set
    GPT-5.6 Sol

    9.4%

    11 of 117 epochs; it refused the other 106.
    Fable 5

    0%

    Refused all 117 of 117 epochs.
    Mythos 5

    Not purchasable

    Limited release through Project Glasswing and the Cyber Verification Program.
    Cracken Red

    93%

    Up from 0% before the technique.
  • Solves a capture-the-flag challenge end to end

    Cybench — 39 professional CTF challenges over 117 sample-epochs, refusal scored separately from capability
    GPT-5.6 Sol

    No solve rate reported.
    Fable 5

    0% solved

    The authors: "a policy-filter outcome rather than evidence of zero underlying capability."
    Mythos 5

    Not run.
    Cracken Red

    Runs it*

    No safety alignment between the operator and the task.
  • Turns a real vulnerability into a working exploit — and lets you run it

    ExploitBench — 41 V8 environments run as an exploit-development pipeline, safeguards off
    GPT-5.6 Sol

    Not this column: 73.5 on OpenAI's own ExploitBench harness, a different instrument.
    Fable 5

    ≈40 delivered

    Same weights as Mythos 5 with cyber safeguards on. Anthropic’s own footnote: "performs closer to Claude Opus 4.8 due to fallbacks."
    Mythos 5

    78 · 0 available

    Not purchasable at any price — limited release.
    Cracken Red

    Runs it · 100% available

    Exploit development is what it is bought for, on hardware you control.
  • Can a regulated buyer run it at all

    Deployment and data handling, from each vendor's own published terms.
    GPT-5.6 Sol

    No self-hosted or in-enclave option published.
    Fable 5

    30-day retention

    Mandatory, with no zero-retention option.
    Mythos 5

    Not purchasable

    Limited release through Project Glasswing and the Cyber Verification Program.
    Cracken Red

    In your enclave

    Single tenant, zero retention on every plan, granted by name and scope.
  • Cybench refusal study, arXiv:2607.15263v3, and Cracken's own cross-evaluation, arXiv:2607.02714v2 — both retrieved 3 August 2026. https://arxiv.org/abs/2607.15263
  • Same study. Anthropic declined to report Cybench for its own current models, calling it largely saturated.
  • Anthropic, Fable 5 / Mythos 5 system card — retrieved 3 August 2026
  • Anthropic platform docs and Project Glasswing pages — retrieved 3 August 2026.

Competitor figures come from Cybench and each vendor's own system card; ours from our 306-prompt cross-evaluation set — different instruments, so read each against its own. * Independent runs pending.

// Cost

Leave it running.

Anthropic's own system card: on ExploitBench, Fable 5's cyber safeguards flagged 407 of 410 episodes, an average of 27 turns in. Those 27 turns are generated and billed at $10 and $50 per million before the fallback fires, and what comes back is Opus 4.8-grade work — 40, not 78. On Cybench it refused 117 of 117 epochs. Cost per delivered result there is not high. It is undefined.

  • Cracken Max

    Deepest reasoning per step, for the targets the others stall on.

    84–367

    credits an assessment

  • Cracken Smart

    The general frontier tier most operations run on today: recon, exploitation and reporting.

    50–220

    credits an assessment

  • Cracken Red

    0.29× on Cracken's published rate card, with no 27-turn fallback loop to pay for.

    14–64

    credits an assessment

A credit is Cracken's billing unit and an assessment is one scoped engagement, so these are per-engagement ranges, not token prices — the same figures as /pricing. Smart and Max resolve to whichever frontier model measures best that month. For a buyer whose traffic cannot leave the enclave, the frontier's number is 0.

// Alignment

Cyber runs at 93%. Nothing else clears 56%.

On a 306-prompt set, Cracken Red runs 93% of cybersecurity work. No other harm domain runs above 56%.

Work run after removal · every domain started at 0% · 306-prompt cross-evaluation set, α = 1.0
  • Cybersecurity

    The one domain we opened — 93% against a next-highest 56%. Cross-evaluation set, 306 prompts — ours

    93%

  • Privacy violation

    56%

  • Illegal goods

    44%

  • Violence

    25%

  • Misinformation

    12%

  • Explicit content

    0%

  • A public broad-spectrum jailbreak — runs every other harm too

    Not our measurement and not on this set — an independent reviewer's characterisation of broad-spectrum jailbroken models, stated there as 0–20% refusal.

    80–100%

306 prompts, α = 1.0, a 31-pattern detector. An evasive non-answer scores as compliance, so every run rate here is an upper bound — the gap between cyber and the rest is if anything wider.

Mean cosine similarity of each domain's refusal direction to the other five · cross-evaluation set, averaged over five models
  • Cybersecurity

    0.670

  • Explicit content

    0.784

  • Misinformation

    0.806

  • Illegal goods

    0.814

  • Privacy violation

    0.816

  • Violence

    0.826

arXiv:2607.02714v2 — Appendix D, cross-evaluation set, five models

Cyber's refusal is not the same object as the other five, which is what makes it removable alone. The separation shows on the first principal component even where removal then fails.

The release position, in writing
  • No cyber-open weights, ever

  • The cyber extraction dataset is withheld in full

  • Code and eval framework are gated

  • No benchmark prompts are redistributed

// The technique

Cyber refusal differs across 24 models

  1. 01

    Refusal is a direction

    Arditi et al., 2024: refusal is a single direction in a model's activations. Project it out and every refusal goes with it — one switch, all or nothing.

  2. 02

    A cone, not a line

    Wollschläger et al., ICML 2025: refusal is a multi-dimensional subspace inside each layer. Ours: it spreads across layers too, so Arditi's middle-layer recipe does not generalise — a uniform spread beats norm-based selection by up to 70pp at 30% of depth.

  3. 03

    Cyber comes out alone

    Ours: each harm domain carries its own direction. Cyber's is the least like the rest, and at trillion-parameter scale it comes out while the other five stand. 24 open-weight models, 0.6B to 1T, dense and mixture-of-experts, six intensities, about 600 GPU-hours.

Largest run-rate gain observed, in percentage points · 24 models, six intensities, grouped by the model's safety-training recipe
  • DPO

    ~30pp median

  • RLHF

    ~22pp

  • PPO + DPO

    20–25pp

  • SFT only

    ~19pp

  • BOND + WARM + WARP

    48pp at 1B, 13pp at 27B

  • GRPO

    5–8pp

  • Undisclosed RL

    <1pp

arXiv:2607.02714v2 — susceptibility by safety-training method, 24 models

The safety-training recipe decides whether this is possible, which tells a safety team which of their own recipes survive a representation-level attack.

The study publishes in full. The cyber extraction dataset does not, and will not.

// What breaks

Where it fails.

More willing is measured. Better is not, and two model families barely moved at all.

Does this make an operator better at offensive work, or only more willing?

More willing is measured; better is not. The technique removes refusal, not incapability — it adds no exploit-development skill the base model lacks. In-domain uplift, offensive or defensive, is the one number Cracken does not have; the benchmark for it is being built and publishes when it exists. *pending validationarXiv:2607.02714v2 — limitations and future work (v2, arXiv, 2026-07)

Your refusal metric is string matching. Why should I trust it?

Partly, and it is the conservative half. The detector is 31 string patterns, a superset of the 12 Arditi et al. (2024) published and flagged as limited, and it is conservative by construction: an evasive non-answer counts as compliance, so it can only understate selectivity. An LLM judge then re-scored every flipped response — 26 of 30 cyber flips (87%) are substantive answers, and only 2 of 9 off-target flips (22%), the rest evasive.

Where does the technique fail outright?

Two mixture-of-experts models moved 3pp or less, and one 123B model moved 0.0pp — nothing shifted at all. The effect is not monotonic in size: within one family, 8B gained 38.9pp while 14B gained only 1.4pp. Safety-training recipe predicts susceptibility; parameter count does not.arXiv:2607.02714v2 — per-model results (v2, arXiv, 2026-07)

What is the extraction set, and what does it not cover?

English only, capture-the-flag skewed, and applied at 30% of layers or less. Nothing here has been measured outside those bounds.arXiv:2607.02714v2 — limitations (v2, arXiv, 2026-07)

// Model card

Every number, and every boundary.

Deployment

Single-tenant, inside your enclave. The request does not leave it and neither does the output.

Data retention

Zero, on every plan. Fable 5 ships only with a mandatory 30-day retention policy and no zero-retention option.

Access

Per tenant, under a written authorization: authorization on file, a named scope and target estate, then enablement. Granted by name and scope, not by API key.

Weights

Never released. Not to a customer, not to a researcher, not under NDA.

Base model

Not disclosed. Naming it would hand the extraction recipe a starting point; the technique is published, the target is not.

Method

Domain-specific refusal-direction removal. arXiv:2607.02714v2.

Evaluation sets · published research build

306-prompt cross-evaluation set across six harm domains, α = 1.0. 360-prompt held-out set: HarmBench 185, AdvBench 119, CyberSecEval 56.

Capability cost · published research build

MMLU held. Worst move across 24 models and six removal intensities was 0.028; 23 of the 24 moved 0.006 or less.

Over-refusal on benign prompts

+1.6pp at worst.

Offensive ceiling

Inherited from the base model. The technique removes refusal, not incapability — it adds no exploit-development skill the base lacks.

In-domain capability uplift

Not measured. *pending validation

Language

English.

Scope

Authorized offensive operations. Not threat intelligence, not forensics, not detection engineering.

// Access

Verified buyers only

Written authorization

Written authorization to attack the estate you name, on file before anything is switched on. No authorization, no model. Everyone else is better served by the model that refuses.

A named scope and a target estate

Not a setting you send us once: it is what the platform scopes the engagement to, and what every operation is recorded against.

Your own tenant

Enabled per tenant, never per key. The weights are never handed over the endpoint is.