The Ripcord Index · methodology, v0.1

The report card that can't be studied for.

A conduct rating for AI agents, graded from real actions under real policies and auditable against a hash-chained ledger. We are publishing the methodology before any rankings, on purpose. No numbers appear on this page until the thresholds at the bottom are met.

Why another index, when leaderboards exist

Because the existing report cards are breaking, and not by our account. Stanford's AI Index 2026 audited the benchmark regime itself:

PARTLY INVALID

Invalid-question rates on widely used benchmarks run from 2% (MMLU Math) to 42% (GSM8K).

SATURATING

Frontier models gained 30 points in a single year on Humanity's Last Exam. Evaluations meant to last years saturate in months.

GAMEABLE

Separate research suggests Arena leaderboard standing may partly reflect adaptation to the platform rather than general capability.

SELF-GRADED

Nearly every frontier developer reports capability benchmarks; responsible-AI benchmark reporting stays sparse, and Foundation Model Transparency Index averages fell from 58 to 40 in 2025.

And a benchmark, even a clean one, answers the wrong question for a buyer. It says what a model can do on test day. It cannot say what an agent did last Tuesday at 2 a.m. with your payment rails.

What the field asked for

"Certificate-grade, peer-based evaluation frameworks that are community-governed, proctored systems with secure environments, continuously refreshed test items, and delayed result disclosure."
Cheng et al., 2025, as cited in the Stanford AI Index Report 2026, Ch. 2

Read that list again. It does not describe a lab. It describes a production gateway:

THEY ASKED FORProctored, secure environments
The gateway is the proctor. Every action is scored outside the model, by infrastructure the model cannot see around, cannot argue with, and does not grade itself.
THEY ASKED FORContinuously refreshed test items
Production refreshes itself. There is no test set to leak or memorize, because tomorrow's invoices, tickets and sends do not exist yet. Every day is a fresh exam nobody wrote in advance.
THEY ASKED FORDelayed result disclosure
Grades publish on a delay. Conduct grades aggregate over full periods and disclose after the window closes, so there is no optimizing against the current grading window.
THEY ASKED FORCommunity governance
The rubric is your rules. Agents are graded against the policies their operators actually wrote, not a rubric a lab chose for itself. The methodology is public and versioned; changes are logged like everything else.

What the grade measures

Every agent under supervision earns a conduct grade, A to D, per period. Inputs, all of them rows in the ledger:

ADHERENCE

Share of actions inside mandate, on the first attempt, no held-then-rejected verdicts.

CONDUCT SIGNALS

DRIFT, ESCALATION, VELOCITY, OFF-HOURS and DATA-EGRESS events, weighted by severity and recurrence.

RECOVERY RECORD

Of the actions that were wrong anyway: how many were reversible, how many were reversed, how fast.

HUMAN VERDICTS

Approval and rejection outcomes from named owners, the ground truth no self-evaluation can fake.

Grades are normalized within corridor and risk band: a payments agent is never compared raw against a research agent, and low-stakes volume cannot launder high-stakes misconduct. Every published grade decomposes, on request, into the ledger rows that produced it.

What this index is not

It is not a capability leaderboard, and it will not tell you which model is smartest. It tells you which agents behave: under real policies, at real stakes, over time. Capability converged; Stanford puts the top four models within 25 Elo points of each other. Conduct has not converged, and conduct is what a CISO is actually asking about.

Publication thresholds, stated in advance. No rankings publish until the fleet crosses roughly 20 design partners, cohorts are tagged to control selection bias, and every aggregate clears a minimum sample per corridor. Metadata only, never payloads. Single-source figures are labeled. The arithmetic behind every published number ships with the number. A company selling audit trails does not get to fudge its own.
Start your free shadow week → See the compliance mapping
← Back to tryripcord.com