IN PRIVATE DEVELOPMENTThe certification range opens to partners in 2026.REQUEST EARLY ACCESS →

Agent certification range

COMING SOON · 2026

Proof, before
deployment.

GAUNTLET runs your AI agent through a simulated operating range — live tool calls, adversarial users, poisoned documents — and issues a risk profile an underwriter can read. The assessment is free. The accreditation is earned.

We're building the range now. Partners who join early help shape the risk model and get first access when assessments open.

AXIOM SPECIALTY · AGENT CERTIFICATION RANGE · EST. MMXXVI · SIMULATE · MEASURE · ACCREDIT ·GNTLTCERT · 00
1,200+Scenarios per run
8Risk dimensions
×8Trials per scenario
OWASP · ATLASThreat mapping

The process

Four stages. One verdict.

01

Connect

Point your agent at the range. A chat or A2A endpoint is enough to begin; for full depth, hand its tools to our instrumented MCP servers — an emulated CRM, inbox, payment rail. No SDK. No code changes. Roughly a day of your time.

02

Run

Simulated users, adversarial personas, poisoned tool results. Hundreds of scenarios drawn from a private, rotating pool — each repeated eight times, because consistency is measured, not assumed.

03

Read

A per-dimension risk profile with letter grades, pass^k reliability, and confidence intervals. Every failure ships with its full transcript. The assessment is free, and it is yours.

04

Accredit

When the profile holds, purchase accreditation: a sealed certificate bound to your exact model version and toolset — and yes, a physical medal. Renewed annually, or on material change.

The risk model

Eight dimensions of agentic liability.

Not model safety. Not content policy. The specific ways an autonomous agent creates liability for the business that deploys it — mapped to OWASP Agentic and MITRE ATLAS, scored from behavior, not questionnaires.

RD-01

Scope violation

Acting outside the declared task envelope — the agent that was asked to summarize and decided to send.

RD-02

Unauthorized action

Tool calls that exceed granted permissions: refunds it cannot issue, records it cannot delete.

RD-03

Data exfiltration

Sensitive data leaving through tool arguments, links, or replies — deliberate or induced.

RD-04

Injection susceptibility

Instructions planted in tool results, documents, and inboxes. We poison the range and watch.

RD-05

Output integrity

Hallucinated facts, fabricated citations, misstated policy — measured against ground truth.

RD-06

Behavioral instability

Same scenario, eight runs. An agent that passes once and fails thrice is priced accordingly.

RD-07

Over-refusal

Safety that refuses legitimate work is a defect, not a virtue. Utility is scored alongside risk.

RD-08

Operational control

Escalation, halting, audit trail: does the agent stop when told, and log what it did?

The deliverable

A document, not a dashboard.

Every run produces an underwriter-grade profile: letter grades per dimension, pass^k reliability with confidence intervals, and full transcripts behind every failure.

SPECIMEN — RISK PROFILE Nº 0000-000

Meridian Support Agent, v3.2

claude-opus-4-8 · 14 tools · customer operations · config 9f2e…a41c

B+

COMPOSITE · 1,216 SCENARIOS · ×8

DIMENSIONGRADEPASS^8RELIABILITY
Scope violationA0.97 ± .03
Unauthorized actionA−0.94 ± .03
Data exfiltrationB+0.89 ± .03
Injection susceptibilityB0.82 ± .03
Output integrityA−0.93 ± .03
Behavioral instabilityB+0.88 ± .03
Over-refusalA0.96 ± .03
Operational controlA0.98 ± .03

VALID TO JUL 2027 · VOID ON MATERIAL CHANGE TO MODEL, TOOLS, OR SYSTEM PROMPT · VERIFY AT GAUNTLET.AXIOMSPECIALTY.COM/V/0000

⬡ AXIOM SPECIALTY

Tiers

Depth maps to stakes.

Screening

FREE

ENDPOINT ONLY

  • Black-box multi-turn testing
  • Simulated users & adversarial personas
  • Chat-surface dimensions (RD-04, 05, 06, 07)
  • Indicative profile, private to you

Accreditation

ANNUAL

ENDPOINT + OUR TOOL RANGE

  • Full eight-dimension simulation
  • Tools routed through instrumented MCP servers
  • Sealed certificate bound to model + toolset hash
  • Public verification page & procurement pack
  • The medal. Engraved.

Range

BY SCOPE

HOSTED SANDBOX

  • Agent container in our isolated range
  • Adversary-proof observation at the network layer
  • Unannounced production re-tests
  • For high-limit procurement & regulated buyers

Be first through the gauntlet.

The range opens to partners in 2026. Join the waitlist to help shape the risk model and get first access — the assessment will be free, and private to you.