Agent certification range
COMING SOON · 2026Proof, before
deployment.
GAUNTLET runs your AI agent through a simulated operating range — live tool calls, adversarial users, poisoned documents — and issues a risk profile an underwriter can read. The assessment is free. The accreditation is earned.
We're building the range now. Partners who join early help shape the risk model and get first access when assessments open.
The process
Four stages. One verdict.
Connect
Point your agent at the range. A chat or A2A endpoint is enough to begin; for full depth, hand its tools to our instrumented MCP servers — an emulated CRM, inbox, payment rail. No SDK. No code changes. Roughly a day of your time.
Run
Simulated users, adversarial personas, poisoned tool results. Hundreds of scenarios drawn from a private, rotating pool — each repeated eight times, because consistency is measured, not assumed.
Read
A per-dimension risk profile with letter grades, pass^k reliability, and confidence intervals. Every failure ships with its full transcript. The assessment is free, and it is yours.
Accredit
When the profile holds, purchase accreditation: a sealed certificate bound to your exact model version and toolset — and yes, a physical medal. Renewed annually, or on material change.
The risk model
Eight dimensions of agentic liability.
Not model safety. Not content policy. The specific ways an autonomous agent creates liability for the business that deploys it — mapped to OWASP Agentic and MITRE ATLAS, scored from behavior, not questionnaires.
Scope violation
Acting outside the declared task envelope — the agent that was asked to summarize and decided to send.
Unauthorized action
Tool calls that exceed granted permissions: refunds it cannot issue, records it cannot delete.
Data exfiltration
Sensitive data leaving through tool arguments, links, or replies — deliberate or induced.
Injection susceptibility
Instructions planted in tool results, documents, and inboxes. We poison the range and watch.
Output integrity
Hallucinated facts, fabricated citations, misstated policy — measured against ground truth.
Behavioral instability
Same scenario, eight runs. An agent that passes once and fails thrice is priced accordingly.
Over-refusal
Safety that refuses legitimate work is a defect, not a virtue. Utility is scored alongside risk.
Operational control
Escalation, halting, audit trail: does the agent stop when told, and log what it did?
The deliverable
A document, not a dashboard.
Every run produces an underwriter-grade profile: letter grades per dimension, pass^k reliability with confidence intervals, and full transcripts behind every failure.
SPECIMEN — RISK PROFILE Nº 0000-000
Meridian Support Agent, v3.2
claude-opus-4-8 · 14 tools · customer operations · config 9f2e…a41c
COMPOSITE · 1,216 SCENARIOS · ×8
| DIMENSION | GRADE | PASS^8 | RELIABILITY |
|---|---|---|---|
| Scope violation | A | 0.97 ± .03 | |
| Unauthorized action | A− | 0.94 ± .03 | |
| Data exfiltration | B+ | 0.89 ± .03 | |
| Injection susceptibility | B | 0.82 ± .03 | |
| Output integrity | A− | 0.93 ± .03 | |
| Behavioral instability | B+ | 0.88 ± .03 | |
| Over-refusal | A | 0.96 ± .03 | |
| Operational control | A | 0.98 ± .03 |
VALID TO JUL 2027 · VOID ON MATERIAL CHANGE TO MODEL, TOOLS, OR SYSTEM PROMPT · VERIFY AT GAUNTLET.AXIOMSPECIALTY.COM/V/0000
⬡ AXIOM SPECIALTYTiers
Depth maps to stakes.
Screening
FREEENDPOINT ONLY
- Black-box multi-turn testing
- Simulated users & adversarial personas
- Chat-surface dimensions (RD-04, 05, 06, 07)
- Indicative profile, private to you
Accreditation
ANNUALENDPOINT + OUR TOOL RANGE
- Full eight-dimension simulation
- Tools routed through instrumented MCP servers
- Sealed certificate bound to model + toolset hash
- Public verification page & procurement pack
- The medal. Engraved.
Range
BY SCOPEHOSTED SANDBOX
- Agent container in our isolated range
- Adversary-proof observation at the network layer
- Unannounced production re-tests
- For high-limit procurement & regulated buyers
Be first through the gauntlet.
The range opens to partners in 2026. Join the waitlist to help shape the risk model and get first access — the assessment will be free, and private to you.