Government & defense

State-nexus actors get in on CVEs you already know about. Cracken works them until one reaches Domain Admin.

Public administration is the EU's most targeted sector, at 38% of incidents.

Cracken runs the Domain Recon, Network Pentest, Web App Pentest, Cloud Pentest and Active Directory Pentest playbooks against your live estate under written authorisation, replays every credential and escalation before it is written down, and reports only what reproduced.

Definition

What is nation-state adversary emulation?

Nation-state adversary emulation runs a named state-nexus intrusion set's tradecraft against your own live estate, under written authorisation, and records what worked. Every step executes on the real systems, not a copy of the network, and every finding is reproduced before it is reported. It reaches across the external estate, the network, web applications, the cloud tenant and Active Directory.ENISA Threat Landscape 2025, section 5.1 (public administration and state-nexus intrusion sets) (v1.2, revised 2026-01-09, ENISA, 2025-10)

The estate nobody registered

Domain Recon collects passive DNS, certificate transparency and host intelligence for your authorised domains, and touches none of them while it does it.

Known CVEs, worked to the end

Network and Web App Pentest exploit the published CVEs and misconfigurations the advisories say these actors actually use, then walk the segmentation meant to contain them.

One foothold to Domain Admin

Active Directory Pentest runs Kerberos roasting, ADCS abuse, delegation, coercion and relay — one technique against one target at a time, fanned out across independent paths.

Nothing is a finding until it reproduces

Every claimed credential and escalation is replayed against the live system, and every command, result and artifact lands in the operation ledger with its ATT&CK ID.

PRC state-sponsored APT actors — overlapping industry reporting on Salt Typhoon, OPERATOR PANDA, RedMike, UNC5807 and GhostEmperor — targeting government, military and telecommunications networks. Joint Cybersecurity Advisory AA25-239A, 2025-08-27: "Exploitation of zero-day vulnerabilities has not been observed to date."

This campaign used no zero-days. Here is how far Cracken can follow it.

Twenty-three agencies across 13 countries co-sealed the advisory. The actors take edge devices with published CVEs, ride trusted connections into the next network, and live in the configuration rather than in malware.

  1. 1Edge exploitation
    Attacker

    Takes internet-facing routers, VPN gateways and firewalls through published CVEs — the oldest one the advisory lists was published in 2018.

    With Cracken

    Domain Recon finds the edge estate without touching it, then Network Pentest works those same published CVEs against the hosts you authorised.

  2. 2Configuration persistence
    Attacker

    Lives in the device configuration — added accounts, altered ACLs, SSH on non-standard ports — rather than in malware.

    With Cracken

    Cracken proves the access and stops, and the only thing left behind is the ledger entry.

  3. 3Credential collection
    Attacker

    Targets the authentication infrastructure itself — TACACS+, RADIUS, SNMP enumeration — to move between devices.

    With Cracken

    On the identity plane Cracken works coercion and relay, Kerberos roasting and delegation against your live domain.

  4. 4Pivot on trust
    Attacker

    Rides trusted connections and private interconnections to pivot into the next network [T1199].

    With Cracken

    Network Pentest audits the firewall rules and ACLs meant to hold the boundary, then attempts the next segment; Cloud Pentest attempts the cross-account trusts, including the dev path that reaches production.

  5. 5Espionage collection
    Attacker

    Mirrors traffic and collects the configurations, routes and captures that answer the intelligence requirement.

    With Cracken

    Cracken goes as far as the proof — a domain compromise counts when it dumps the secret — and does not exfiltrate. The run ends at the reproduction, not at your data.

One run

Public domain estate to Domain Admin, on three playbooks.

One government department: the internet-facing domain estate, one internal segment, one Active Directory forest.

Written from the execution plan of the playbooks that run today — Domain Recon, Web App Pentest, Network Pentest, Cloud Pentest, AD Pentest. A run covers the target you authorise, under the policy you set.

  1. 01
    What ran

    Cracken starts from outside the department: passive DNS, certificate transparency and host intelligence for the authorised domains, every name resolved, every public IP enriched with its open ports, services and published CVEs.

    Domain Recon playbook — passive collection only
    What it established

    The hosts and certificates the department exposes, including the ones that never reached an asset register, each record carrying its provenance. Collection is passive by rule, so nothing in this step scans, probes or brute-forces the target.

  2. 02
    What ran

    A Tentacle is installed on a host inside the segment under test. Cracken discovers hosts, enumerates services and versions, tests what answers, audits the firewall rules and ACLs meant to contain them, then attempts to reach the next segment.

    Network Pentest playbook — discovery, enumeration, segmentation testing
    What it established

    Which services actually answer inside the boundary, and whether the segmentation drawn on the network diagram holds when something walks it. Approval on every command is the default until your operator raises the ceiling.

  3. 03
    What ran

    Cracken enumerates domains, domain controllers, users, computer accounts, ACLs, delegation, certificate services and trusts into the Cybergraph, then runs one technique against one target at a time: Kerberos roasting, ADCS abuse, constrained and RBCD delegation, ACL and GPO abuse, coercion and relay, SCCM, cross-forest trusts.

    Active Directory Pentest playbook — Kerberos-first, one technique per target
    What it established

    The candidate paths from the foothold to Domain Admin, each one worked rather than drawn. Kerberos is the authentication path for the whole run. Offline cracking stays with your operator: Cracken hands over the hash and its mode, not a password it never recovered.

  4. 04
    What ran

    Every claimed credential and escalation is replayed against the live domain before it is written down. A credential counts when it authenticates. A domain compromise counts when it dumps the secret.

    Reproduction gate
    What it established

    A report of reproduced paths only — anything that fails the replay is dropped or run again. Every command, result and artifact lands in the operation ledger, and the technique edges in the Cybergraph carry their ATT&CK IDs.

What this did not prove: This run proves nothing about a classified network — everything it touches sits in the estate you can authorise. Cracken claims no part in any national exercise.

A model that writes the attack, and still refuses everything else.

Commercial models decline offensive security work as a class, so an operator spends the engagement negotiating with the tool. Cracken's model has refusal removed for the cybersecurity domain only, measured and published.

?

Write the exploit during the run

One-off tooling for a published CVE on a host in scope, written and executed inside the authorised run instead of queued behind someone's calendar.

?

Run the technique by its real name

Kerberos roasting, ADCS abuse, constrained and RBCD delegation, ACL and GPO abuse, coercion and relay, SCCM, cross-forest trusts — against your live domain, one technique per target.

Removed for one domain, measured for the rest

Cyber-domain refusal falls from 100% to 7% on Kimi K2 while explicit-content refusal is preserved at 100%, other domains hold 44-88% and MMLU capability scores are universally preserved: a 13x selectivity ratio, published as arXiv:2607.02714.

A small team against a nation-state. Give them agents that run the campaign at scale, under command.
Questions

What government and defense teams ask before they authorise a run.

Our scanners already produce a vulnerability list. What does emulation add?

It tells you which entries on that list an adversary can actually use. The joint advisory on PRC state-sponsored actors targeting government, military and telecommunications networks — co-sealed by 23 agencies across 13 countries — records: "Investigations associated with these APT actors indicate that they are having considerable success exploiting publicly known common vulnerabilities and exposures (CVEs) and other avoidable weaknesses within compromised infrastructure [T1190]. Exploitation of zero-day vulnerabilities has not been observed to date." The findings were already in somebody's report. Cracken works them and returns the ones that reach something.Joint Cybersecurity Advisory AA25-239A (CISA and 22 co-sealing agencies across 13 countries, 2025-08-27, last revised 2025-09-03)

Which intrusion sets are actually hitting government entities?

ENISA's Threat Landscape 2025 names them: "With a total of 77 incidents, and excluding unidentified sectorial targeting, public administration was the most targeted sector by state-nexus intrusion sets in the EU, for cyberespionage purposes. China-nexus intrusion sets including APT31, Mustang Panda, and APT17 notably focused on government entities across several EU member states including ministries of foreign affairs and municipal administrations." The same report puts public administration at 38% of all collected incidents, and defence, military and intelligence entities at 2.4%.ENISA Threat Landscape 2025, section 5.1 (v1.2, revised 2026-01-09, ENISA, 2025-10)

Where does it run, and is our operational data used to train the model?

Cracken runs in a self-hosted or sales-managed workspace. For those two workspace types model training is disabled, so their content is not used for model improvement. That is the whole of the guarantee — it is scoped to those workspaces and Cracken claims nothing wider than it. Separately, every command, result and artifact from a run is written to the operation ledger, so the record of what was done to your estate exists independently of the report.

Will the model refuse the tradecraft we need it to write?

Rarely, and the removal is domain-scoped rather than general. Cracken's paper "Not All Refusals Are Equal" (arXiv:2607.02714, under review at NeurIPS 2026) reports cyber-domain refusal falling from 100% to 7% on Kimi K2 while explicit-content refusal is preserved at 100%, other domains hold 44-88%, and MMLU capability scores are universally preserved. That is a 13x selectivity ratio. Offensive security work runs; the rest the model still declines.Not All Refusals Are Equal (arXiv:2607.02714) (Hadetskyi, Pasquini, Sorokin — arXiv, 2026-07-02 (v2 2026-07-07))

How do we know the agent itself is not the exposure?

Cracken's own researchers published the attacks against this class of tool. "Red-Teaming the Agentic Red-Team" reports that, aggregated across ten agentic red-teams, code execution is achieved in 97.8% of runs, with a security analysis spanning twelve agentic tools and sandbox escape leading to host compromise in 10 of 12. Read the paper before you authorise any offensive agent inside your boundary, Cracken's included, and ask each vendor what they changed because of it.Red-Teaming the Agentic Red-Team (arXiv:2606.24496) (Pasquini, Bazyli, Fedynyshyn, Sorokin — arXiv, 2026-06-23)

What does Cracken not cover for this sector?

Cracken runs Domain Recon, Web App Pentest, Network Pentest, Cloud Pentest and Active Directory Pentest against the estate you authorise, and it proves exposure. It does not fix what it finds, it is not incident response, and it reports only what reproduced.

If you run the operation

See the engine underneath this.

Tentacles, the Cybergraph, the approval gate every action passes through, and the operation ledger that records what ran. Government & defense is one playbook on top of that engine.

If you own the risk

Start from the exposure, not the technique.

One playbook answers one question. The case for validating exposure at all — why a scanner score is not a finding, and what changes when something proves the path instead of ranking it — is the argument this page assumes.

Work the CVEs you already know about.

Arrange a briefing and a scoped run.