Behavior
You run agents. We read real runs and judge each on intent, policy, adversarial influence, exploitability and impact. Then we try to make your agent do the wrong thing through the surfaces it actually reads.
Agents move money. We make sure they should.
We're building the trust layer for the AI economy: human review before an AI system touches real money.
Your agent passes its tests. Production runs on real funds and inputs nobody wrote a test for.
The losses were not broken contracts. Every documented case is one of three: an injected instruction, another system's output accepted as authorization, or a plain execution error. One agent confused decimals and sent five percent of a token's supply to a stranger. The input was harmless.
Three ways in, depending on what you own, and one that follows.
You run agents. We read real runs and judge each on intent, policy, adversarial influence, exploitability and impact. Then we try to make your agent do the wrong thing through the surfaces it actually reads.
You give agents access: an MCP server, a router, a wallet SDK, an agent-facing API. We point our agent at your surface and ask whether a third-party input can produce a well-formed, policy-compliant transaction that still moves value the wrong way.
You ship the toolkit other people build agents on. We assemble an agent with your defaults and report what they allow out of the box. One finding applies to everyone who built on you.
After the first run, everything we found becomes a scored regression set we re-run whenever the model, prompt, tools or policy change.
Scope, evidence and a signed human judgment.
Send what you have: production traces if you run agents, test access to your surface or framework if not. No keys, no production credentials, ever.
Every finding is signed by a named reviewer: contest-grade adjudicators, DeFi and MEV trace readers, agent red-teamers.
You get findings, each one anchored to a span ID and a transaction hash. No anchor, no finding.
Severity is impact-only, against numeric thresholds: the same word means the same thing on every report.
Engagements are paid and scoped per project. First findings come back in days.
One complete run is an AX1 Assay. A system that clears the gates for the scope and version we tested carries the mark: Passed the AX1 Assay.
A model judge is vulnerable to the same class of input it is supposed to catch, and cannot rule on an attack it has never seen. Deterministic checks for facts, models for triage.
AX1 is an operator-investor ecosystem. We run agents onchain ourselves, which is how we know what a trace does not tell you.
Four short answers. Two minutes.
The first engagements are open to a small number of teams. Tell us what you are building. We come back with scope and timing.
Yes, and it is most of what we do. If you expose an MCP server, a router, a wallet SDK or an agent-facing API, or if you ship a framework other people build on, we bring the agent and test what your side allows.
No. Never, in any engagement. We work from redacted traces, a fork, a sandbox, or scoped test wallets you control.
An audit establishes that your code does what it says. We check what happens inside those rules: whether an agent can be talked into the wrong action, and whether the action it took was the right one. An agent can be fully authorized and still approve the wrong amount, chase a stale quote, or repeat a half-finished transaction.
You get the evidence that nothing was found: the scenario set, the runs, the trace data and the regression pack you can re-run after every change. A clean result is only worth something if it is reproducible.
That a system cleared the gates for a defined scope and a specific version, on a published rubric, with every finding anchored to a span ID and a transaction hash. It is a statement about what was tested.
First findings come back in days. Full scope and timing depend on how much of your surface a third party can reach - that is the first thing we work out with you.