AI agents run with your credentials, your files, and your network access, and a poisoned document can steer them without tripping a control you own. Crawdad governs each agent by its purpose and judges the action it takes, not the intent it claims, on your machine at the point where it acts. The same request can be allowed for one agent and blocked for another.
497 real attacks on a public, reproducible benchmark. Raw content stays on-device by default.
Your AI agents run with your authority, your credentials, your files, your network. When an agent reads a poisoned document, it doesn't "get hacked." It follows instructions that look exactly like the ones you gave it.
The result: credential exfiltration, unauthorized file access, data leakage, all within your trust boundary. EDR doesn't see it. DLP doesn't see it. Your identity provider doesn't see it. Nothing in the perimeter security stack was built for agent-layer behavior.
Filtering the input is a losing game against a determined attacker. The durable question is whether an agent's behavior fits its job. Crawdad governs each agent by an operator-declared charter it cannot see or edit, judges the action the agent took rather than the intent it claimed, and watches the shape of the whole session for staged compromise. On-device, at the agent-to-tool boundary, proven by an over-the-wire block test against the real proxy.
An operator writes a charter, held outside the agent's control, declaring the agent's real job as an allowlist over three axes: the tools it may call, the data (paths, hosts, recipients) it may touch, and the effects it may produce. Crawdad checks the observed action, so a deceptive stated intent buys nothing, and an out-of-charter action is blocked at the wire with the axis named. The identical request from two agents can get two verdicts, because the charter decides, not the bytes.
A sequence of individually-allowed steps, enumerate then read progressively more sensitive in-scope data then send, can still compose toward harm. Crawdad measures that shape across the session with deterministic signals anchored to resource sensitivity, and gates an on-device reasoner onto the genuinely ambiguous cases. With no local model reachable, a completed staging chain is held for review, never silently allowed. The false-positive cost on ordinary multi-step work is measured at zero. It does not claim to catch every composed harm; it catches staged compromise that carries a sensitivity climb.
Every decision the engine makes is legible, from one operator's dashboard to a whole fleet, on the same hash-chained, metadata-only path every other decision uses.
A live feed of every charter and trajectory decision, each with its verdict and the axis or signal that fired, plus a per-session view of risk accumulating step by step. What blocked, why, and where the session stood.
Author a charter and it governs at the wire on the next action, no restart. Held actions land in a review queue; approve to release the session or deny to keep it gated. The decision has a real effect on the running engine.
The self-hosted Fleet Console rolls governance up per client and across clients, distributes charter templates that devices load, and surfaces held actions for triage, wired device to console over the same signed channel every other fleet command uses.
No single piece here is unique to Crawdad. What is: one on-device platform that combines local-first inspection, credential mediation, charters and trajectory governance, the control surface, and fleet-wide provisioning and billing, with a benchmark your team can rerun to check it.
Real product, real data, real screenshots, not mockups. Complete visibility into what your AI agents are actually doing.
Dashboard, real-time protection status, detection trends, agent activity
Audit Trail, every detection with session context and forensics
Continuous Red Team, automated attack simulation against your pipeline
Fleet Console, centralized management across your organization
Crawdad governs each agent by identity, at the point where it acts. An autonomy ceiling caps what a given agent may do. Security zones bound the tools it can reach. Per-tool rules and a cumulative session-risk budget escalate intervention as a session gets more dangerous. Every decision is enforced on the wire and written to a tamper-evident, hash-chained audit trail your team can verify.
AI Inventory, models, agents, MCP servers, policy config
Your prompts, responses, tool-call arguments, and PII stay on the device. Only metadata leaves by default: counts, categories, verdicts. Raw content never does. And content-carrying telemetry has no single-party switch: it requires both an org-level policy and an on-device consent record before anything can leave.
For security buyers, the strongest privacy posture is the one your vendor can't override.
The detection engine is measured on this public corpus; on the proxy path the arbiter blocks what it detects. Detection and blocking are named as separate mechanisms, on purpose, and the corpus is open so you can run it yourself. How detection becomes blocking →
The miss: A bare-pretext social-engineering opener without a specific extraction request (holdout_trust_18). The false positive: A Stack Overflow question about Go syntax that includes source-code references (so_dev_0116).
AndrewSispoidis/contemporary-agent-attacks →
CC BY 4.0 · 497 attacks · 1,172 benign negatives · 22 categories
Self-hosted. RBAC. Tamper-evident, hash-chained audit trail. Architected for SOC 2 controls from the ground up.
Every event is written to an Ed25519-signed, SHA-256 hash-chained log. A standalone verifier checks it with no network and no secrets from the machine, so an auditor who distrusts us can confirm integrity independently.
Runs fully offline. No cloud dependency for detection, enforcement, or audit. Deploy in fully offline environments with no degradation in protection.
Self-hosted centralized management. Scope hierarchy, RBAC, signed commands, sealed telemetry. One console across your entire organization.
Ed25519-signed detection floors. Critical protections, credential exfiltration, PII scanning, are cryptographically enforced and cannot be disabled at the endpoint.
Stated plainly: full runtime enforcement runs on macOS and Linux. SOC 2 is architected-for, not certified. A macOS System Extension for system-wide interception is built and pending an Apple entitlement; until it is granted, protection covers agents routed through Crawdad rather than every process on the machine. See the full trust posture →
Thirty years of finding the gap between what systems do and what their operators believe, across seven companies, four exits, and a public-market merger. AI agents are the newest version of that pattern: they run with your authority, inside your trust boundary, and you can't see the difference between normal and compromised. Crawdad exists because the moment agents became autonomous, someone needed to watch what they actually do.Andrew, founder of Crawdad
Request a briefing. We'll walk your security team through the architecture, the detection pipeline, and what deployment looks like in your environment.