THE AI RUNS THE PENETRATION TEST
This playbook hands your engineers exploits they can re-run, not a list of things that might be exploitable.
What you get

The model behind the run, and the trace it leaves for every decision.
platform.
Action ledger
Every proposed action, its approval or denial, and its output in one thread.
Sub-operation tree
Each dispatched child is its own operation with its own thread you can open.
Empty results
A specialist that runs out of attempts records no candidate found as its result.
Cybergraph handoff record
Children share no transcript, so the Cybergraph is the handoff record.
Where it stops
The engagement ends at a reproduction your engineers can re-run. What that exposure is worth to the business, and what to do about it, is a human risk call Cracken does not make.
It does not test your AI systems
Prompt abuse, data exposure and agent tool misuse are a separate discipline and a separate playbook.
It does not fix what it proves
The agent restores anything it changed back to the value it captured beforehand, which is rollback, not remediation.
It is not incident response or forensics
It runs authorised attacks against your own surface.
It does not simulate
There is no synthetic target and no emulated payload.
Fanning out does not widen the blast radius
Every sub-operation inherits the realm's autonomy ceiling and its Trusted Targets list, so a child cannot auto-approve for itself what the parent would have had to ask you about.
Who this is for
Start testingAppSec Lead
Your scanner returns four hundred findings a quarter and you need to know which of them an attacker actually reaches, before the next release goes out.
Pentester
You lose the first two days of every engagement to mapping, session handling and tool setup, before you start looking for anything interesting.
CISO
You sign off on release dates and you want something behind that signature stronger than a clean scan report.
Questions
How is this different from an autonomous scanner?
A scanner reports what matched, and it is one process. A Cracken pentest bundle is a root agent with no shell that dispatches tentacles, and a tentacle is structurally unable to declare a finding. Recording one requires a separate verify operation, launched in fresh context, that reproduces the access and captures a signal. A candidate the verifier cannot reproduce is dropped or kicked back, never downgraded into the report.
What stops it from running something destructive?
Two independent controls. The realm's semi-autonomous policy decides what auto-executes: new realms start on Balanced, where safe reconnaissance auto-runs while scans and mutations queue for your approval, and every sub-operation inherits that ceiling instead of setting its own. Separately, the exploit sub-playbook treats rollback as a success condition, so any technique that creates a principal or changes an ACL, attribute or template must capture the original value first, restore it after, and re-read it to confirm.
Will the model refuse to run offensive techniques?
Cracken's Red model is security-specialised and built for no-cyber-refusal: it runs the tradecraft a general-purpose model declines to touch. It is bounded by authorisation rather than by topic, so every action still runs inside the realm's scope, under the autonomy level you set, and against targets a person put in scope. — Chinese AI's role in stopping rogue OpenAI agent shows cost of US guardrails (Reuters (Aditya Soni and Jaspreet Singh), read via Yahoo Finance syndication, 22 July 2026)
Can the agent running the test be attacked itself?
Yes, and Cracken's own researchers published the attacks. A tool-running agent is itself attackable, and sandbox escapes to the host are a documented risk class for agentic tooling. That is why Cracken's orchestrator holds no shell of its own: it plans and dispatches, and every command runs in an isolated tentacle with fresh context, inside a Tentacle you own. — Red-Teaming the Agentic Red-Team (Pasquini, Bazyli, Fedynyshyn, Sorokin) (arXiv:2606.24496v1, arXiv preprint; all four authors listed at Cracken on the paper byline, 23 June 2026)
Give the model a target and make it prove the finding.
Point Cracken at an asset you control and read every command it ran.






