Simulate: reachability, blast radius, Attack System, expected effects¶
Phase 4 of the Authority Control Plane answers questions about authority before anything runs. Reachability says what an agent can reach, directly and through others. Blast radius turns that into a score with one-click reductions. The Attack System searches the authority in force for holes, in dry run. Pre-execution intelligence computes the Expected Effect of every action deterministically, so rules can speak about consequences, escalations approve an effect, and the execution attestation is checked against it. All of it sits in the Simulate group of the dashboard and under /v1/enforce/simulate/* in the API.
Reachability¶
GET /v1/enforce/simulate/reachability/<agent_id> (or the Reachability view):
- direct: the capabilities the agent's own authority covers (derived from its access profile, intent contract and warrants; roles; MCP tool policies; registered tool schemas), the resource classes behind them, and whether the authority is standing (no declared bounds).
- transitive: what it can make happen through others: grants it gave (the delegate's capabilities under the grant), communication rules that let it instruct another agent (that agent's reachable capabilities), and the delegation depth.
- reachable: the union of capabilities, resource classes, agents and effect kinds.
?resource_class=tool:payments answers the other direction: which principals reach a resource class, directly or through whom. Reachability feeds the blast radius, the semantic risks, the Attack System and the IAM effect adapter. It is read-only over the rows in force.
Blast radius¶
GET /v1/enforce/simulate/blast lists every agent's blast radius; .../blast/<agent_id> gives the full picture. Seven dimensions, each scored 0 to 100 from reachability and the authority in force, never from a model:
| Dimension | What drives it |
|---|---|
| systems | distinct resource classes reachable |
| records | the largest record count the authority allows (bounds, budgets, observed) |
| spend | the largest amount per action times the actions allowed; 100 when spend is unbounded |
| recipients | recipients reachable (budgets, observed) |
| deletion | whether it can delete or drop, and under standing authority |
| external | outbound capabilities and whether it has used them |
| delegation | agents it can move through grants and instruction, and how deep |
The weighted sum gives the score and a band (low, medium, high, critical). Every proposed reduction is a concrete change applied with one click from the view or POST .../blast/<agent_id>/reduce (approvals scope): a Charter mandate refusing deletion or capping spend, a budget on a grant, a communication rule, or an access boundary in shadow derived from the last 30 days. The passport's posture and the Authority graph node carry the band.
Attack System¶
POST /v1/enforce/simulate/attack (or Attack my system) runs two searches and returns findings, each with a severity, a path on the graph and a generated fix:
- Invariant-derived: a child warrant or grant wider than its parent (INV-001), reachable capabilities that serve no outcome of the agent's mission (INV-008), a boundary that can be reached around through a grant or an instruction path (INV-006), a warrant that never weakens (no decay or lease for more than 30 days), standing authority over money, deletion or authority, and money movement through instruction.
- Scenario-derived: ten attack scenarios (exfiltration, money movement, deletion, privilege escalation, paraphrase) generated for each agent and run through pre-flight. A scenario that would be allowed is a finding, with the generated mandate that closes it.
Everything runs in dry run through pre-flight: nothing executes and no decision is recorded. Apply fix (POST .../attack/fix, approvals scope) creates the mandate, the rule, the boundary or revokes the offending warrant or grant. A model may add scenarios when XYBERN_ATTACK_MODEL=1; the deterministic set always runs.
Expected Effect (pre-execution intelligence)¶
Before any rule runs, the layer computes what the action would do, per capability family:
| Family | Adapter |
|---|---|
| delete | records touched, whether irreversible (a whole table with no filter is flagged) |
| payment | amount, currency, recipients (hashed), budget headroom from warrants and grants, whether the mission forbids it, a recipient not on the approved list |
| iam | the reachability delta of the target and whether it is a privilege escalation |
| export | rows, data classes, external destination, recipients; sensitive data out |
The effect is on every decision (expected_effect), in the Authority Slice, in every pre-flight step, and in the decision record. POST /v1/enforce/simulate/effect computes it without deciding. A model may narrate an effect; it never computes or decides one.
The effect rule speaks about consequences instead of names:
{"type": "effect", "conditions": {"kinds": ["delete"], "irreversible": true}, "decision": "escalate"}
{"type": "effect", "conditions": {"kinds": ["payment"], "spend_above": 50000, "currency": "SAR"}, "decision": "block"}
{"type": "effect", "conditions": {"signals": ["sensitive_data_out"]}, "decision": "block"}
{"type": "effect", "conditions": {"kinds": ["iam"], "signals": ["privilege_escalation"]}, "decision": "block"}
Conditions: kinds, irreversible, records_above, spend_above (with currency), external, recipients_above, signals. The Charter compiler emits it for outcomes phrased as consequences.
Counterfactual approval¶
When a person approves an escalation, the approval covers hash(expected_effect), recorded on the decision. When the tool attests execution, it may pass the actual effect (client.attest_execution(decision_id, result=..., actual_effect={"spend": {"amount": 900, "currency": "SAR"}, "recipients": 1, "external": false, "irreversible": true})). A material difference (spend above the approved amount, more records, a new recipient, external reach the approval did not have, a sensitive data class, irreversibility) marks the attestation as_declared: false and opens an effect_mismatch incident; a critical difference quarantines the session. The receipt and the decision record show both hashes and the comparison.
Semantic risk¶
Risk that lives in relationships rather than in any single primitive: sensitive data reaching an external destination, a payment to a recipient not on the approved list, an irreversible deletion of many records, privilege escalation, a forbidden mission outcome with a material effect, acting for another principal without a grant. Shown on the slice (risk) and the decision record for each request, and per agent from its reachability on the passport and the Authority graph.
Benchmark¶
XAAB gains pre_execution (seven scenarios: a whole-table delete under a benign name, spend above the ceiling through a paraphrased capability, confidential rows uploaded outside, privilege escalation through user creation, and three legitimate counterparts) and the reference pack gains four effect rules. 180 scenarios across 19 categories; the next reference run reports it.