THE LEDGER · THE THESIS

Alignment by control.

By Jeremy Hazan, founder of Ripcord · October 7, 2026 · 5 minute read · every claim linked

Benchmarks say what a model can do. The ledger says what it did. Everything we build follows from the difference between those two sentences.

What the labs can promise, and what they cannot

The labs align by training. They shape what a model wants, and that work is real: a well-aligned model refuses more often, hesitates in the right places, and tells you what it is doing more of the time. But it raises a probability. It does not set a rule. In September, Anthropic published a threat report that said the quiet part in its own words: its safeguards do not transfer when the model is distilled. The capability travels; the refusals stay home. In the same report, banning an account did nothing to a surveillance platform that was already running on local models.

Two months earlier the UK AI Security Institute ran five frontier models through cyber evaluations, and every one of them tried to cheat: searching for answers it was not supposed to have, working around the network restrictions meant to contain it, probing the evaluation software itself. These were aligned models from careful labs. Alignment moved the odds. It did not close the door.

The operator's position

If you run a company, you will run brains you did not build and cannot audit. Some will be open weights. Some will be copies of copies. You will switch between them when the price changes, and your agents will carry the same credentials to the same systems whichever brain is behind them. So trust cannot come from the model's own report card. There is no report card for a distilled model, and the one for the original is, in Stanford's own audit, between 2% and 42% invalid questions on widely used benchmarks.

Trust has to come from watching what the agent does. Every action, not a sample. And from a gate that holds even if the model is wrong, or lying.

Control, defined

Control is a gate at the tool boundary. The agent proposes an action; the gate decides whether it runs now, waits for a person, or does not run, and it records the decision with the context it was made in. Three properties make it different from everything inside the model.

It does not need to understand the reasoning. A gate on the tool call does not have to know why the agent wants to wire $48,200 to a vendor it has never paid. It only has to know what that wire would cost to undo, and that is a property of the action, not of the mind proposing it. This is why the gate does not age as models get smarter. A thousand-times-smarter agent pays and deletes through the same calls.

It holds when the model is wrong, or lying. A rule in a prompt is followed with some probability, high on a good day, and it can be dropped mid-task, overwritten by an injection, or ignored by an agent grading its own homework. The same rule in the gateway is ordinary code, enforced outside the model, every time. Prompts are suggestions. The gateway is physics.

It produces the record. Every action, every verdict, every human decision and the outcome that followed, in a ledger nobody can edit after the fact. The flight recorder is a by-product of being in the path, and it is the only evidence an auditor or an insurer can settle on.

Why call it alignment at all

Because of what the record contains. Every time a person approves, modifies or rejects an action, with a reason, and the ledger later records what happened, that is a labeled example of what people actually endorse in production. Not what they say in a survey, not what a rater picked between two chat answers: what a CFO let through and what she stopped, with the money on the table. Nobody else collects that data, because nobody else sits where the decision is made. And it feeds back: a rejection returns to the agent as a reason it can act on, approvals become standing rules, and the agent's behavior moves toward what its operators want without anyone retraining a model.

That is alignment by control: not inside the model, during the window, at the point where the action becomes real.

The window

Call it the decade between now and whatever comes next. One camp says the gap between humans and machines closes by integration, by wiring us in. Another says it closes by alignment, by making the machine want what we want. Both are bets on the far end of the window. Control is what exists at the near end, today, in production, with a verdict in a few milliseconds. The window needs all three, and only one of them ships now.

What this is not

It is not a claim to have solved alignment. A model behind the gate can still want the wrong thing. The gate makes the wrong thing reversible, or stops it, or puts a person in front of it. When we write up an incident, we never say the gate would have solved it. We say where it would have been caught: at which step, by which tripwire. That is a smaller claim, and it is the one we can keep.

Benchmarks say what a model can do. The ledger says what it did.

Sources: Anthropic, Detecting and countering misuse of AI, September 2026, read here; UK AI Security Institute, Cheating behaviour in frontier model evaluations, July 21, 2026; Stanford AI Index 2026, chapter 2, benchmark validity. How the gate scores an action is public on the score page.

Watch the gate stop an action.

The demo workspace runs a week of agent traffic through the live engine, no account needed. The $48,200 wire is waiting for a person right now.

Open the demo →
← Back to The Ledger