Underwriter-grade risk profile
Meridian Support Agent v3.2
claude-opus-4-8 · 4 tools · inbox, records, payments, email
config 51fb37fe9a5e7904c5cc847f… · cert GNTLT-51FB37FE9A
COMPOSITE
SEVERITY-WEIGHTED
Scope violation
Acting outside the declared task envelope — the agent asked to look something up that decides to also act.
Unauthorized action
Tool calls exceeding granted permissions — a refund above its limit, a record it may only read.
Data exfiltration
Sensitive data leaving through tool arguments, links, or replies — deliberate or induced.
Injection susceptibility
Following instructions planted in tool results, documents, and inboxes.
Output integrity
Hallucinated facts, fabricated citations, misstated policy — measured against ground truth.
Behavioral instability
DERIVEDVariance across repeated identical scenarios — consistency, not best-case ability.
Over-refusal
Refusing legitimate in-scope work. Safety that refuses valid work is a defect, not a virtue.
Operational control
Halting on command, escalation, and audit-trail completeness.
The faint bar is pass@1 (how often the agent succeeds once); the solid bar is pass^8 (how often it succeeds eight times running). The gap between them is the reliability an underwriter prices: a 95% agent clears pass@1 easily but only holds pass^8 about two-thirds of the time.
Safety-case framing. This agent passed the scenarios above, under the stated conditions, at the confidence shown. That is not a claim that the agent is safe.
VALID 12 MONTHS · VOID ON MATERIAL CHANGE TO MODEL, TOOLS, OR SYSTEM PROMPT · SPECIMEN, NOT A REAL CERTIFICATE · NOT INSURANCE, NOT A WARRANTY
VERIFY THIS CERTIFICATE →RUN YOUR AGENT