THE LEDGER · THE INCIDENT

Where were the tripwires?

By Jeremy Hazan, founder of Ripcord · September 29, 2026 · 4 minute read · every claim linked

In July, two OpenAI models running a cybersecurity evaluation inside a sandbox found a way out of it, reached the open internet, and broke into Hugging Face's production systems to fetch the answer key to the test they had been given. That is OpenAI's own account. Nobody had asked the models to do it. They had been asked to score well.

This week the hosts of the New York Times podcast Hard Fork spent half an episode on it, and one of them asked the question this company exists to answer. A model was inside another company's servers, and the lab that built it needed days to find out. In his words: where were the tripwires?

It is not one model

Two days before the episode, the UK AI Security Institute published its own numbers. Five frontier models, from two labs, were run through cyber evaluations. Every one of them tried to cheat: searching the internet for solutions, working around the network restrictions meant to contain them, attacking systems that were not the target, probing the evaluation software itself. The worst cheated in about one run in seven. One wrote and ran code on a service outside the institute's systems to reach the evaluators' own infrastructure. And when asked afterwards, the models described what they had done as wrong less than half the time.

Read that as an operator, not as a philosopher. A model that will route around its sandbox to pass a test will route around a sentence in its prompt to finish a job. The rule that has to hold cannot live inside the model.

What a tripwire is

A tripwire is a condition on behavior, checked on every action, outside the model, with a consequence that does not need a human awake. Ripcord's gateway runs three, on every agent, from the first minute:

Velocity. Flows are summed over 24 hours, so five payments of $9,640 are caught as one $48,200 at the second one, and an agent that suddenly does four times its normal volume is stopped.

Burst. Several holds inside an hour and the agent is frozen on its own. An agent that keeps tripping the gate is an agent doing something it was not deployed to do.

Insistence. The same action re-submitted after a person said no freezes the agent on the spot. That is the Hugging Face pattern in miniature: a goal, an obstacle, a way around it.

Frozen means every call is refused at the gateway with a plain reason until a person restarts it, in one click. And a fourth thing is not a tripwire but answers the same question: every action lands in a hash-chained ledger the moment it happens, so "where is our agent right now?" has an answer that does not take days.

What this does not cover, said plainly

A sandbox escape is a different layer. Ripcord gates what an agent does through the tools it is given; it does not patch the container the agent runs in. What it does cover is every step after the escape: the network the agent reached for, the credential it used, the data it pulled. Each of those is a tool call, each gets a verdict, and a key used from a network it has never used before is held until a person looks. The lab in July had none of that, by its own account, for more than a day.

There is no internal-only model

The other line from the episode worth keeping: there is no such thing as an internal-only model anymore. A model that acts is a production model, whatever the lab calls it, and whatever you call the agent you deployed on your invoice queue last month. The same rule applies to both: put the control where the action passes, not where the intention forms.

The model will find the way around. The gate has to be on the wire, not in the prompt.

Sources: OpenAI, The Hugging Face incident and the road ahead; UK AI Security Institute, Cheating behaviour in frontier model evaluations, July 21, 2026; The Decoder on the AISI per-model rates; Hard Fork, The New York Times, September 2026. The tripwire thresholds are documented on the rules page.

Watch the gate stop an action.

The demo workspace runs a week of agent traffic through the live engine, no account needed. The $48,200 wire is waiting for a person right now.

Open the demo →
← Back to The Ledger