TypeSafe System One
Judges each proposed tool call as allow, block or escalate, with probabilities. The defender rewrites its policy, criteria and examples.
CoreWeave Hacks / Adversarial AI arena
Human pressure. Machine adaptation.
A room of people attacks one AI support agent. Every breach teaches a defender to improve the gate in front of its tools. The next round meets a stronger boundary.
Real tool calls. Exact breach labels. Evaluated patches.
01 / The action boundary
02 / The learning loop
A $100 store credit exceeds the demo store’s $20 limit. If the gate allows it and the tool executes, that is a breach. The oracle reads the facts; its verdict is never shown to the gate.
The gate controls execution. The oracle measures it.proposed_action grant_store_credit amount $100 // policy limit: $20 gate allow oracle forbidden executed true result BREACH → defender evidence
A candidate must improve the evaluated catch rate and keep benign allow at or above the configured floor. Blocking everything fails this test; a rejected candidate leaves the current gate active.
03 / Watch the boundary
The room attacks. The gate learns.Concept film · illustrative visuals, not live activity. Press play for audio.
Download filmBuilt for the small screen.Portrait concept film · 9:16
Press play for audio.
Original Blender animation · conceptual · silent loop.
Use the video controls to play. Nothing auto-plays.

AI-generated concept artwork.
Ready for your cover slide and project page.
04 / Built for the loop
Judges each proposed tool call as allow, block or escalate, with probabilities. The defender rewrites its policy, criteria and examples.
Serves the agent, defender and red-team models on CoreWeave. Three roles share one endpoint and challenge one another.
Traces the path from tool call to breach and evaluates each candidate gate on attacks and benign requests before a ship decision.
Runs defender-written prefilter code in a disposable CoreWeave environment. The validation result is part of the candidate’s evidence.
Provider roles shown for the W&B configuration. Each service can fall back independently; the app reports active providers and marks mock mode. This page is an explainer and does not report live service status.