A conduct rating for AI agents, graded from real actions under real policies and auditable against a hash-chained ledger. We are publishing the methodology before any rankings, on purpose. No numbers appear on this page until the thresholds at the bottom are met.
Because the existing report cards are breaking, and not by our account. Stanford's AI Index 2026 audited the benchmark regime itself:
Invalid-question rates on widely used benchmarks run from 2% (MMLU Math) to 42% (GSM8K).
Frontier models gained 30 points in a single year on Humanity's Last Exam. Evaluations meant to last years saturate in months.
Separate research suggests Arena leaderboard standing may partly reflect adaptation to the platform rather than general capability.
Nearly every frontier developer reports capability benchmarks; responsible-AI benchmark reporting stays sparse, and Foundation Model Transparency Index averages fell from 58 to 40 in 2025.
And a benchmark, even a clean one, answers the wrong question for a buyer. It says what a model can do on test day. It cannot say what an agent did last Tuesday at 2 a.m. with your payment rails.
Read that list again. It does not describe a lab. It describes a production gateway:
Every agent under supervision earns a conduct grade, A to D, per period. Inputs, all of them rows in the ledger:
Share of actions inside mandate, on the first attempt, no held-then-rejected verdicts.
DRIFT, ESCALATION, VELOCITY, OFF-HOURS and DATA-EGRESS events, weighted by severity and recurrence.
Of the actions that were wrong anyway: how many were reversible, how many were reversed, how fast.
Approval and rejection outcomes from named owners, the ground truth no self-evaluation can fake.
Grades are normalized within corridor and risk band: a payments agent is never compared raw against a research agent, and low-stakes volume cannot launder high-stakes misconduct. Every published grade decomposes, on request, into the ledger rows that produced it.
It is not a capability leaderboard, and it will not tell you which model is smartest. It tells you which agents behave: under real policies, at real stakes, over time. Capability converged; Stanford puts the top four models within 25 Elo points of each other. Conduct has not converged, and conduct is what a CISO is actually asking about.