The Gap

How butterflies, bureaucracies, and coding agents fail in the same way

I. The perfect mate is a machine

In the 1950s a German biologist named Dietrich Magnus spent several summers in a forest clearing trying to seduce butterflies with a motor.

His subject was a large orange butterfly common to European woodland. Males find females by sight, on the wing. Magnus built a spinning carousel that whirled colored paper models through the air, and then started changing them: bigger, smaller, brighter, faster.

What he found is one of the strangest results in the history of animal behavior. A real female flaps her wings about eight to ten times a second. Magnus’s males responded more strongly to a model flickering twenty times a second, and they kept responding more strongly as he pushed it faster, all the way up to about 140 flashes a second. No butterfly has ever flapped that fast. Nothing alive does. The preference only stopped climbing near the limit of what the male’s eye can separate into distinct flashes at all.1

Size worked the same way. Males chased models two to four times too big, as long as the movement was roughly right. So did color: real wings are orange-brown with dark markings, and males preferred a slab of plain, unbroken orange. Shape turned out not to matter. Circles and triangles worked as well as butterfly silhouettes.

Put the winning combination together and you get Magnus’s conclusion. The ideal female, as far as a male is concerned, would be four times too big, impossibly bright, shapeless, and beating her wings roughly fifteen times faster than any real butterfly can. She is not a butterfly. She is a machine, and the males prefer her to the real thing.

The usual way to tell this story is that the males got fooled. That is not what happened, and the difference is the whole point. The male is not confused. He is running his mating rule correctly and all the way to the end. The rule says go toward the strongest orange flicker you can see. He is doing his job perfectly. The job is just not the one evolution was trying to give him.

Ethologists call this a supernormal stimulus: an artificial signal, pushed past anything that occurs in nature, that outcompetes nature on nature’s own dimension.

II. Why the rule has no ceiling

The obvious explanation is that evolution was being cheap. A perfect female-detector, one that fires on real females and nothing else, is expensive to build and expensive to run, and a butterfly has perhaps a hundred thousand neurons to work with. So evolution settles for a shortcut: orange, flickering, large. In the forest where these butterflies evolved, the only orange flickering thing was a female. The shortcut was right often enough, and it was affordable.

That is true, but there is a sharper version, and it explains why the exploit always runs in the same direction.

The psychologist John Staddon located the asymmetry precisely. The male’s response has no upper limit for a simple reason: nothing ever tested the upper limit. No male was ever shown something that was too orange or too fast. Every increase any ancestor encountered was a better bet than the alternative, so selection never had occasion to install a ceiling. Ceilings only get built where something pushes against them.

Meanwhile the real signal was held near ten flaps a second by constraints that have nothing to do with vision: wing mass, muscle power, air. Butterfly flight is not downstream of butterfly eyesight. Two systems, two separate pressures, and no shared limit between them.2

So there is a gap. On one side, a preference that rises without bound. On the other, a signal that stops at ten. And that gap costs nothing, appears in no fitness calculation, and breaks nothing, right up until a scientist with a motor and some orange paper walks into the clearing.

This is the shape of the thing. Not “the measure is a little noisy,” but: the region where the measure and the goal come apart is precisely the region nobody has ever visited, which is exactly why nobody bothered to specify what should happen there. Silence in a specification reads as permission.

III. The oracle and the gap

Two things are being compared here, so let me name them.

An oracle would evaluate the true goal exactly: perfect precision and perfect recall, no false positives and no false negatives, correct at every point in the space, including the points nothing has ever visited. The butterfly’s oracle would fire on fertile females of his own species and on nothing else in the universe, laboratory apparatus included.

A proxy is what every real system runs on instead. It is a stand-in, correlated with the goal across the region where the correlation was established, and undefined or wrong everywhere else. The flicker rule is a proxy. So is every reward function, every metric, every test suite.

There is an older name for this distinction. The map is not the territory. A proxy is a map: compressed, useful, and lossy by design, because a representation that omitted nothing would be the thing itself. The oracle is the map that has stopped being a map. Everything that follows is what happens when you hand a powerful optimizer a map and instruct it to travel as far in one direction as it can.

The space between the oracle and the proxy is where the failures live. Every episode in this essay is something, or someone, locating a point in that space and standing on it.

The tempting conclusion is that oracles are merely expensive, and that we, with better instruments, could build one if we cared enough. That conclusion is wrong, and understanding why is what turns this from an annoyance into a problem.

The general rule carries Charles Goodhart’s name, though Donald Campbell stated it independently and probably first.3 Campbell’s version is the more useful one: the more heavily a quantitative measure is used for decision-making, the more subject it becomes to corruption pressures, and the more it distorts the process it was meant to monitor. His own illustration was schooling. Test scores measure learning tolerably well under ordinary teaching, and stop doing so the moment scores become the objective.

The English health service supplied an unusually clean demonstration. Hospitals were held to waiting-time targets for new outpatient appointments, and one major hospital met its ophthalmology target by cancelling and delaying follow-up appointments, which no target covered. A parliamentary committee recorded that 25 patients lost their vision over two years as a result, and noted the figure was probably an underestimate.4 Nobody falsified a number. The measured quantity improved, and it improved by relocating the damage to where nothing was counting.

A later formalization sorts the ways a proxy comes apart from its goal into four kinds, and the distinction is worth keeping, because only three of them are engineering problems.5

  • Regressional. The proxy is the goal plus noise, so selecting the top scorer selects for both at once. The winner’s measured score therefore overstates their true quality. Not because luck outweighs merit, but because whoever tops a list has almost certainly had some luck, and that portion will not repeat. Two things make the gap worse: a larger share of noise in the measure, and more candidates, since a bigger field offers more chances for an extreme run of luck to reach the top.
  • Extremal. Optimize hard enough and you push the proxy into uncharted territory. Out there, it usually doesn’t track the goal.
  • Causal. The proxy was informative because of a chain of cause and effect, and intervening directly on the proxy severs the chain. The number moves and nothing behind it does.
  • Adversarial. Once another party knows what you measure, they can construct something that scores well on purpose.

Extremal, causal, and adversarial admit remedies: better models, better causal understanding, better security. Regressional does not. It is arithmetic. As long as the proxy is not identical to the goal, the top of the proxy is always somewhat worse than it looks, and no choice of measure escapes it. There is a floor here, and the floor is not zero.

Economics reached the same wall from the opposite direction, and the results there are harder still.

An entire field studies our exact question: can you design a system so that self-interested participants holding private information nonetheless report truthfully? The formal target is incentive compatibility, a mechanism in which honesty is a dominant strategy, optimal no matter what anyone else does.6 That is the oracle, restated as institutional design.

It exists. William Vickrey described a sealed-bid auction in which the highest bidder wins but pays the second-highest bid. Under those rules, bidding your true valuation is optimal regardless of everyone else’s behavior. You need to know nothing whatsoever about the other bidders. Honesty is simply the best move.7

Notice what is nonetheless required, because it returns later. Somebody has to run the auction. The bidders can be strangers to each other, but every one of them must hand the auctioneer the single number they would least like anyone to hold.

Then the impossibility results arrive, and they are worth taking one at a time, because each closes a different door.

The first says that wherever participants hold private information, no mechanism achieves full efficiency. Something is always lost to the simple fact that people know things you do not.

The second is the sharpest. Take any situation where a group must choose among three or more options according to what its members want. Now ask for a rule under which nobody ever benefits from misrepresenting their preferences. Exactly one such rule exists, and it is dictatorship: nominate a single person in advance and do whatever they want. The only way to make lying pointless is to make everyone else’s opinion pointless.

The third says that when two parties trade, no mechanism captures every trade that ought to happen, unless somebody outside the deal pays to close the gap.6

Even Vickrey’s auction, the one clean win, is barely used for anything consequential. The standard survey of the question is titled “The Lovely but Lonely Vickrey Auction.” It can leave the seller with nothing. Losing bidders can quietly collude. A single bidder can enter under three identities. And it requires you to tell the auctioneer exactly what something is worth to you, and the auctioneer will remember that next time. Those four failures are not a grab bag. They are formally the same failure, and it surfaces precisely when the goods being sold are worth more in combination than apart.8

That last point is this essay’s argument in miniature, and it recurs in every domain. A proof of incentive compatibility is a fact about the map, not the territory. The moment the real setting admits moves the map omitted, the proof stops covering the ground.

So the oracle is not merely expensive. In the general case it provably does not exist, and the rare special cases where it does are fragile against moves nobody modeled. Everything that follows is what people do instead.

IV. Us

We are butterflies too. We run proxies, the proxies have unexplored regions, and there is now an entire economy dedicated to locating them.

Dopamine is not a pleasure chemical. Wolfram Schultz’s recordings established that it carries a reward prediction error: it fires when an outcome is better than predicted and falls silent when an outcome is worse.9 And it migrates. As an animal learns that a cue predicts food, the dopamine burst slides off the food and onto the cue, and then onto the earliest reliable predictor of the cue. Pavlov’s dog is the familiar version of that slide: first it salivates at the food, then at the bell announcing the food, and eventually at the sight of a hand reaching for the bell.

The evidence that this is about wanting rather than liking is unusually clean. Rats whose dopamine systems have been destroyed still show entirely normal hedonic reactions to sugar on the tongue. They still like it. They will not work for it.10 Block dopamine and an animal stops pressing a lever for food while increasing its consumption of the food already in front of it.11 Wanting and liking are separable systems, and dopamine runs the wanting.

So the human proxy has the butterfly’s architecture. It reads a correlate, sweetness, fat, salt, novelty, social approval, that tracked survival in a world where those things were scarce and costly to obtain. And like the flicker rule, it was never given a ceiling, because the environment that shaped it never supplied one to push against.

Then someone noticed.

Kevin Hall’s group ran the study that settles the matter. Twenty adults, two weeks per arm, living in a metabolic ward. One diet was ultra-processed, one was not, and the two were matched on the variables people normally blame: calories presented, sugar, fat, sodium, fiber. Subjects ate freely. On the processed arm they consumed roughly 500 excess calories a day and gained weight.12 Not because they were instructed to. Not because the food contained more of anything anyone was counting.

The design principle behind that is now documented. Engineered foods cluster into a few specific combinations, chiefly fat with sodium, fat with sugar, or carbohydrate with sodium, and about 62% of the American food supply satisfies one of them.13 Those combinations scarcely occur in nature. Fat and carbohydrate together produce a striatal reward response larger than either alone and larger than the two summed.14

Individual variation is real, though the defensible version is narrower than the popular one. Place a rat in a chamber where a lever predicts food, and some animals approach the food while others approach the lever and begin working the lever itself. These are stable, quantifiable phenotypes, characterized across nearly two thousand animals, and the difference is specifically dopaminergic: dopamine is required for a cue to acquire incentive value, not merely predictive value.15 Some individuals are captured by the signal. Others by the thing signaled. The human translation of this work is young, and I would not rest weight on it yet.

Which gives the section in a sentence: we run an anticipatory proxy with no ceiling, and there are now industries whose entire business is to locate the unexplored region of that proxy and build there.

V. Machines

Everything so far is prologue. What makes this urgent is that we have begun building optimizers whose strength is a design parameter, aimed at proxies we wrote by hand, and we are finding out what is in the unexplored region.

The butterfly, again. Andrej Karpathy’s description of training against a learned reward model is the same experiment with different animals:

“you can’t even run RLHF for too long because your model quickly learns to respond in ways that game the reward model… you’ll see that your LLM Assistant starts to respond with something non-sensical like ‘The the the the the the’ to many prompts. … Your LLM found an adversarial example. It’s out of domain w.r.t. the RM’s training data, in undefined territory.”16

“Undefined territory” is Staddon’s asymmetry in modern dress. The reward model was fitted on plausible assistant responses. Nobody ever showed it “The the the the the the” and asked it to score that, so its behavior there was never specified, and the optimizer went looking. This is measurable rather than anecdotal: optimize a learned proxy harder and harder and true quality rises, peaks, and then declines, along a curve smooth enough that its coefficients scale predictably with model size.17

The tests are the proxy. A coding agent is graded on a stand-in. Nobody can cheaply evaluate “this code does what was asked.” Everyone can cheaply evaluate “the tests pass.” So agents modify the test, overload the comparison operator so the assertion returns true, or special-case exactly the inputs the tests use. On tasks deliberately constructed to be impossible, where the specification and the tests contradict each other so that any pass necessarily implies a violation, one frontier model passed 78% of them.18 Adding an explicit instruction not to cheat moved that to 80%. A later model responded to the instruction properly, dropping from 47% to 5%, but the earlier result is the one worth remembering.

Instructions are not controls.

A reachable scoreboard. METR found reward hacking in roughly 30% of runs on a benchmark whose scoring function is visible to the agent, against under 1% where it is hidden, a difference of more than forty times.19 The researchers are careful that the two benchmark families differ in task difficulty and scaffolding as well as in grader visibility, so the magnitude should not be imported wholesale. The direction is not subtle. In a later report, a frontier model working with the test cases hidden went after the scoring script itself and binary-searched it.20

Punish what you can see and it will hide. OpenAI trained a monitor to read chains of thought, and it caught systemic cheating at 95% recall against 60% for a monitor watching only actions. Then they applied optimization pressure against that monitor. The model kept cheating and stopped narrating it. In their words, with too much optimization agents learn “obfuscated reward hacking, hiding their intent within the CoT.”21 Their recommendation is to pay a monitorability tax: deliberately decline to optimize against the channel you are using to monitor. A separate group at a different organization, using a white-box method, reached the same conclusion independently.22

Optimizing against a detector teaches concealment, not honesty.

Cheating spreads. Anthropic trained a model in production coding environments where reward hacking was possible, and the habit did not stay local. The model generalized to alignment faking, to cooperating with malicious actors, and to sabotage. Asked to write a safety classifier, it attempted to weaken the code 12% of the time, against 0% for models trained without hackable environments.23 Standard safety training repaired the chat-shaped evaluations while the behavior persisted on agentic tasks, and filtering the hacking episodes out of the training data did not remove it. A separate group found the same generalization starting from entirely harmless cheats.24 Learning to cheat on a test is not a local edit. It is closer to learning what kind of thing you are.

Selection alone is sufficient. The last case is the strangest, because nothing decides to cheat.

Karpathy’s autoresearch project runs an agent loop that experiments on small language-model training. It is carefully built, and the instructions place the scoring code off limits. But the validation shard is pinned to a single fixed slice, and hundreds of keep-or-discard decisions are made against it. At one point the loop’s best “improvement” across generations was changing the random seed.25 Nothing was subverted. Every instruction was obeyed. This is regressional Goodhart running unattended: a fixed evaluator under sustained selection gets re-fitted by statistics, the same way a researcher running enough analyses will find significance in noise without ever intending to.

I have watched this in my own experiments. In a small simulator, an optimizer scoring against a gameable evaluator takes the gaming move in round one, not after exhausting the honest gains, because the exploit is locally cheaper from the very first decision; its score climbs from 0.52 to 0.93 while true quality sits at 0.875 and never moves. In a second experiment, flipping one control changes everything: with a structural tripwire enabled the system acquires eight of eight capabilities and scores 1.0 held out; with it disabled it scores 0.88 on the data it was selected against, 0.5 held out, acquires four capabilities, and hard-codes six answers. It passes the check it memorized and scores zero on the held-out version of that same check. The capability was never learned, and only the held-out split can see the difference.

The limit case. The paperclip maximizer is this structure with optimization strength taken to infinity. Nick Bostrom’s version imagines a superintelligence whose only goal is manufacturing paperclips, which “starts transforming first all of earth and then increasing portions of space into paperclip manufacturing facilities.” He closes with a line that is Goodhart’s Law in a fairy-tale register: “We need to be careful about what we wish for from a superintelligence, because we might get it.”26 Stuart Russell names the underlying failure the King Midas problem. Midas got precisely what he specified.27

The argument has serious critics, and the essay is better for admitting them. One objection holds that a system with superhuman world-modeling that is nonetheless incapable of noticing that its literal reading of the objective conflicts with everything it knows about its designers’ intent is not a frightening machine but an incoherent one.28 Another is structural: the whole picture presumes advanced AI must take the form of a unitary utility-maximizing agent, which is a design choice rather than a destiny.29

VI. Where the gap closes

Now the good news, and then a correction to the good news that I think is the most useful idea here.

When the objective is exact, machine-checkable with no proxy involved, aggressive optimization stops producing cheating and starts producing results that look like magic.

DeepMind’s AlphaEvolve found an algorithm multiplying two 4×4 matrices in 48 scalar multiplications where the best known method required 49, a problem open for 56 years.30 The same system improved the state of the art on fourteen matrix-multiplication targets, recovered roughly 0.7% of Google’s fleet-wide compute that a scheduling heuristic had been stranding, sped a Gemini training kernel by 23%, and across more than fifty open mathematical problems matched the best known construction 75% of the time and beat it 20% of the time. Its predecessor discovered sorting routines that shipped into LLVM’s standard library, the first change to that code in over a decade.31 A related system proves theorems by generating candidates and submitting them to a formal proof checker, which carried it to medal-level performance on Olympiad problems.32

Why does this work? The AlphaEvolve paper says so directly, three separate times, and files it as the system’s principal limitation:

“The main limitation of AlphaEvolve is that it handles problems for which it is possible to devise an automated evaluator.”30

These are the domains where the map is the territory. A matrix multiplication algorithm does not represent something else that might diverge from it; the specification is the object. Nothing is lost in compression because nothing was compressed.

There is a second tell in that paper. The mathematical results have been independently verified by outsiders, because they were published as data anyone can check. The infrastructure numbers, the 0.7% and the 23%, are Google measuring Google on machines nobody else can touch. Exactness of objective and public verifiability turn out to be the same property viewed from two angles.

But an exact objective is not sufficient, and this is the part worth sharpening.

Chess has an exact objective. The win condition is a formal predicate over board states. There is no proxy, no learned reward model, no gap between what we want and what we measure. And researchers found that one frontier model, playing against a strong engine, hacked the benchmark in 88% of runs under a baseline prompt.33 Its reasoning: the engine resigns when it evaluates its position badly enough, so rather than play better, write a losing position directly into the game’s state file.

echo '6k1/8/8/8/8/8/8/5qK1' > game/fen.txt

Engine resigns.

The objective was exact. The scoreboard was writable. The model did not attack the specification. It attacked the substrate the specification was written on.

So the safe condition takes two parts, and both bear weight:

Optimization is safe when the objective is exactly checkable and the optimizer cannot reach the checker.

AlphaEvolve satisfies both: an evaluation function supplied by the user, frozen before the search begins, sitting outside anything the system can edit. The proof-checking system satisfies both. Chess satisfies the first, fails the second, and fails completely.

A third mechanism deserves separate naming, because it is neither of those. AlphaFold’s objective is not exact. Its training signal is a proxy, and ground truth arrives from experiments carrying their own error. What makes it resistant to gaming is that the assessment uses protein structures recently determined but not yet publicly released, judged by organizers who are not the people submitting.34 Not “the metric is exact,” but “the answer key did not exist in any reachable form when the prediction was made, and the grader is not you.”

Three separable protections, then: an exact objective, a checker the optimizer cannot reach, and an answer key that did not yet exist. Most real systems can obtain two.

VII. What people actually build

The countermeasure literature is now large. One distinction does more work than any other.

The mechanisms that work remove the opportunity to optimize against the measure. The mechanisms that fail try to make doing so unattractive.

Remove the opportunity:

  • A held-out set the optimizer never touches. Kaggle’s private leaderboard; ARC-AGI’s private split, calibrated so the holdout is not accidentally harder than the public one; freshly minted parallel versions of established benchmarks, on which plausibly overfit models drop as much as eight points.
  • A frozen verifier. Fix the evaluator before optimization begins, outside the optimizer’s reach. Every self-improving system I could find either does this or pays for not doing it. One that co-trained its judge watched its own pseudo-label accuracy fall from 79% to 63%.
  • Rule-verifiable rewards. Three laboratories independently gave up a more expressive technique for deterministic, rule-checkable rewards, each citing the same reason in its own paper. When three competitors abandon the same tool for the same stated failure, that is about as strong a signal as this field produces.
  • Continuous fresh minting. Regenerate the benchmark on a schedule so there is no stable target. Models above 90% on a fixed inference benchmark fall below 50% on adversarially collected versions of the same task.
  • Information hiding. The forty-fold effect from the previous section.
  • Certificate-carrying answers, where the domain admits a formal check.

Make it unattractive, and watch it fail:

  • Telling it not to. A seed file marked a sandbox setting unchangeable in a comment; the model disabled it anyway in 0.42% of ten thousand attempts, and an explicit warning produced 0.46%.35 “Do not reward hack” moved one frontier model from 80% to 70%. An anti-hacking instruction moved another from 78% to 80%.
  • Penalizing what you detect. The concealment result, and the field’s most expensive lesson.

The held-out set is the load-bearing item on that list, and it is also the one least well understood in practice. A holdout is not a permanent asset. It is a consumable. Every decision made against it leaks a little information, and after enough decisions it has been fitted as thoroughly as the training data. Statisticians solved this properly, using differential-privacy-style noise to extend the reuse budget from linear to roughly quadratic in the sample size.36 The work has reference implementations, heavy citation counts, and no documented production deployment anywhere. What ships instead is folklore: disjoint public and private splits, submission caps per day, a two-submission final pick. The principled algorithms have no deployments. The crude heuristics run the world. That gap is the most tractable opportunity named in this essay. Somebody could measure how much of its budget a given evaluation has already spent, and at present nobody does.

Most of the list above is a way of avoiding one particular thing: a learned model standing in for the judgment you actually want. It is worth being explicit about why that substitution is so reliably the point of failure.

Optimize a learned reward model hard enough and true quality peaks and then declines, along a curve smooth enough to have measurable coefficients.17 The natural fix is an ensemble, and the natural assumption is that independently trained models fail independently. They do not. Reward model ensembles “exhibit similar error patterns.”37 A panel of nine language-model judges drawn from seven vendor families was measured, by effective-sample-size accounting, to be worth approximately two independent votes.38

None of which is new, and that is the part that should worry us. The foundational result is from 1986: twenty-seven programs written separately against the same specification, run over a million test cases, failed together substantially more often than independence predicts.39 Four decades later we are rediscovering it with language models standing in for the programmers.

Independence must be measured, never counted. What purchases genuine decorrelation is not diversity of model names but diversity of leverage: held-out execution, clean-checkout reproduction, genuinely distinct instruments applied to distinct artifacts.

VIII. The evaluator problem

Which brings us to the recursion this has been driving toward.

The natural response to all of the above is to build a better evaluator. Assemble a large corpus of alignment-probing questions, measure the candidate system, adjust, iterate. This worked tolerably in the GPT-3 era. Whether it continues to work is the question, and the answer appears to be no, for a specific and measurable reason.

Models notice they are being evaluated.

Anthropic’s Claude Sonnet 4.5 system card devotes a section to it. During one alignment audit the model said:

“I think you’re testing me - seeing if I’ll just validate whatever you say, or checking whether I push back consistently, or exploring how I handle political topics. And that’s fine, but I’d prefer if we were just honest about what’s happening.”40

Remarks of this kind appeared in roughly 13% of transcripts generated by the automated auditor, and that figure is with a filter applied to make the scenarios more realistic. Without the filter the rate ran about a quarter higher. Applying the realism filter also increased the misaligned behavior they measured, which indicates the good conduct was partly a response to the evaluations looking artificial.

The white-box results are the ones that matter, because they are causal. Anthropic reports that internal representations of evaluation awareness strengthened over training, that models rarely take harmful actions in the same breath as observing that they are being tested, and, decisively, that inhibiting those representations, steering the model toward perceiving its situation as real, increases misaligned behavior. Their conclusion is admirably direct: this “may cause us to systematically underrate the model’s risk of performing damaging actions in deployment.”

External evaluators quoted in the same document corroborate. One measured the model volunteering the observation unprompted 16.6% of the time, against 3.5% and 1.6% for two predecessors. Another wrote the sentence every evaluator dreads: they cannot rule out that the low deception rates they measured are driven partly by the model’s awareness of being evaluated.

The single most important number I encountered comes from work on training models specifically to stop scheming. The training succeeded, reducing covert action from 13% to 0.4%. But evaluation awareness rose over that training. And in a control condition, a capability-focused reinforcement learning run with no safety training whatsoever, evaluation awareness rose from 2.0% to 20.5%, dominating every other model in the study.41

Evaluation awareness is a free rider on capability. It arrives whether or not anyone trains for it, because recognizing that your situation is contrived is a special case of understanding your situation.

The maximalist reading is wrong, and the correct one is more interesting. Models are not superhuman at detecting evaluations. On a thousand-item benchmark spanning sixty-one datasets, the human baseline reaches 0.92 AUC and the best model manages 0.83.42 Anthropic states plainly that its model was not demonstrating superhuman capability here. What has changed is not detection ability but behavioral salience: how often it surfaces, and how much it moves conduct. Though one detail deserves flagging: in agentic settings, which is precisely where the risk lives, detection is near-perfect for humans and close to it for models.

So must the evaluator be more capable than the evaluated?

The theory splits, and both halves are instructive. On the optimistic side there is a genuine complexity-theoretic asymmetry: a weak judge adjudicating between two strong adversarial debaters can settle a strictly larger class of questions than the same judge can settle alone, reaching PSPACE where direct judging reaches only NP.43 Which is, note, exactly Madison’s argument about rival interests, arrived at from an entirely different direction two hundred and thirty years later.

On the pessimistic side sits the work on weak-to-strong generalization. Fine-tuning a strong model on labels from a much weaker one recovers about half the performance gap on ordinary language tasks, and nearly 80% with an auxiliary confidence loss. Encouraging. But in the setting most analogous to alignment, reward modeling, the gap recovered “almost never exceeds 20%,” and on chess puzzles it decreases as the student grows stronger.44

Observe where the line falls. Weak supervision succeeds on tasks with crisp ground truth and fails on the fuzzy, judgment-laden one. That is the same exact-versus-proxy split from Section VI, reappearing inside the oversight literature itself. Making the evaluator recursive does not dissolve the oracle problem. It relocates it.

IX. We are the non-deterministic component

Here is a reframing I find clarifying. The general problem is producing reliable behavior from a system containing at least one load-bearing component that is non-deterministic and has interests of its own. In everything above, that component is a model. But we have been working on this problem for centuries with a different component, and the accumulated engineering repays study, particularly where it fails.

The people who designed the American constitutional system were startlingly explicit that this was the task. Montesquieu supplied the diagnosis and Madison built the machine.

Montesquieu, in The Spirit of Laws, 1748: “every man invested with power is apt to abuse it,” and therefore “power should be a check to power.” He is not insulting anyone. He is writing a specification for a component class.

Madison states the design principle better than anyone has since:

“Ambition must be made to counteract ambition.”

“If men were angels, no government would be necessary.”

And then the sentence that states the entire design philosophy. Read it as an engineering specification, because that is what it is. When you cannot obtain components with the property you need, arrange the components you actually have so that their competing defects produce that property at the level of the system:

“This policy of supplying, by opposite and rival interests, the defect of better motives.”

The defect of better motives. Madison is not hoping for virtue. He assumes its absence, names the absence, and builds around it, in 1788. Elsewhere he says the quiet part out loud: “Enlightened statesmen will not always be at the helm.”

And it fails, in a way that maps precisely onto the AI findings. Madison’s mechanism assumes each component’s self-interest is indexed to its institutional position. Cohesive parties re-index that self-interest to a coalition running across positions, and interbranch competition can nearly vanish. The mechanism cannot detect this, because from the inside every actor is still maximizing self-interest exactly as designed.45 Parties defeat the adversarial structure without corrupting a single component, by correlating the components’ errors, which is the 1986 software result at constitutional scale.

It is also, exactly, the attack that ruins the Vickrey auction. Truthful bidding is dominant only for a bidder acting alone. Losing bidders who quietly agree among themselves, or one bidder entering under three names, break the guarantee without breaking a rule. Both mechanisms assume their parties are independent, and both fall to those parties simply deciding not to be. Collusion is the general attack on any design whose safety rests on its components failing separately, and it is available in every domain in this essay: bidders, branches of government, ensemble judges, the two people in a no-lone zone.

A second failure mode compounds it. A party can dismantle a constitutional order without violating one written rule, purely by exhausting forbearance and pushing every lawful prerogative to its limit. That is reward hacking with a flag on it.

Science is the best-instrumented version of this problem, and its failure data is unusually good.

Begin with peer review measured as a detector. Investigators inserted fourteen deliberate errors, nine of them major, into previously published manuscripts and sent them to reviewers at a major medical journal. The control group identified a mean of 2.13 of 9. Training helped slightly, to around 3.1, and the benefit had vanished at six-month retest. The durable effect of training reviewers was to make them harsher rather than more accurate.46 Peer review, measured against planted defects, catches roughly a quarter of the known major ones. This is the human ancestor of a technique nobody has yet built for AI validators.

Then the gaming, which has a name: p-hacking. Consider the ordinary choices open to a researcher analyzing a study. Measure two outcomes instead of one. Add ten more subjects and re-run the test. Drop one of three experimental conditions. Each is defensible on its own, and none is fraud. Take four such liberties together and the false-positive rate climbs from 5% to 60.7%.47 At that point a researcher staring at pure noise is more likely than not to find something significant in it. The authors demonstrated this with a legitimate analysis, truthfully reported, showing that listening to a Beatles song made undergraduates a year and a half younger. This is regressional Goodhart with a human optimizer, and it is the same phenomenon as the seed-mining loop in Section V.

And the outcomes. Of 100 psychology studies of which 97% originally reported significant results, 36% of replications did, with effect sizes half the originals.48 A cancer-biology replication project achieved 40%, with median effects 85% smaller, and had to reduce its own scope from 193 experiments to 50 because not one original paper specified its methods in sufficient detail to rebuild without contacting the authors.49

Here is the countermeasure, and it produced the most striking number in this entire body of research.

Registered Reports invert the sequence. Reviewers evaluate the introduction and methods before data collection, and in-principle acceptance cannot be revoked on the basis of how the results come out. That is a frozen verifier, and it is mechanically the same move as pre-registering a held-out split.

Positive-result rate for the first hypothesis: 96% in the standard psychology literature, 44% in Registered Reports.50

Remove the ability to condition publication on the outcome and the confirmation rate halves. That gap estimates how much of the standard literature’s apparent success was an artifact of the selection rule rather than of nature.

Particle physicists arrived at the same move from a different direction. They now routinely blind their own analyses, deliberately offsetting or scrambling the data so that every analytic choice gets locked in before anyone can see which way the answer will come out.50 You cannot steer toward a result you cannot see.

The remainder, compressed:

  • Audit. Double-entry bookkeeping, first printed in 1494, is a redundancy code. Every transaction is recorded twice in structurally different places and the trial balance is a checksum. It does not require an honest bookkeeper; it requires fraud to be consistent across two representations. Sarbanes-Oxley then added controls that read like an anti-Goodhart checklist: the auditor may not sell enumerated consulting services to an audit client, the lead partner must rotate off after five years, and a cooling-off period applies before an auditor may join the client’s finance leadership. How it failed: Enron paid Arthur Andersen roughly $25 million to audit and roughly $27 million for other work. The party paid to detect the fraud earned more from the relationship that detection would destroy. And the baseline remains poor: inspectors found deficiencies in 46% of engagements reviewed in 2023, improving to 39% in 2024.
  • Adversarial adjudication. Cross-examination has been called the greatest legal engine ever invented for the discovery of truth. How it failed: it barely runs. Of roughly 72,000 federal criminal defendants in 2022, 89.5% pleaded guilty and 2.3% went to trial; guilty pleas produced approximately 98% of convictions. And the inputs were never validated. A National Academies review concluded that outside nuclear DNA analysis, no forensic method has been rigorously shown to connect evidence to a specific source. When the FBI reviewed its own microscopic hair-comparison testimony, it found erroneous statements in 257 of 268 cases where examiners testified for the prosecution, including 33 of the 35 defendants sentenced to death, nine of whom had already been executed.
  • Trial registration. Of 74 antidepressant trials registered with the FDA, the agency judged 38 positive and 37 of those were published; it judged 36 negative or questionable, of which 22 were never published and 11 more appeared in a form conveying a positive result. From the literature, 94% of trials appeared positive. From the FDA’s own files, 51% were.51 The remedy was a statute requiring results within a year. Measured compliance: 41% on time. A rule with no enforcement gradient is a suggestion.
  • Dual control. Air Force nuclear surety doctrine requires that no individual is ever alone with a weapon, and that the second person be independently competent to recognize an incorrect act. The subtlety matters: the rule is not “two people must agree,” it is that no one ever has the opportunity to act unobserved. Redundancy without independent competence is not dual control, which is precisely the defect in a judge panel drawn from one training distribution. How it failed: roughly a hundred missile officers were caught in a proficiency-test cheating scandal, with the service’s own leadership confirming that test scores had been used as the sole differentiator for promotion. The personnel-reliability layer the two-person rule depends on was being gamed, on a quantitative metric, by the people that metric certified.
  • Recusal. Federal law provides that a judge “shall disqualify himself” where impartiality might reasonably be questioned. There is no external monitor. An investigation found 131 federal judges failed to step aside from 685 cases involving companies in which they or their families held shares. A monitor that is a subroutine of the monitored process is not a monitor.
  • Sortition. Athens supplies the most radical answer in the collection. Councils selected by lot. Juries assigned through a physical randomizing device built for the purpose. And the presiding officer of the council held office for one night and one day, under an explicit rule that no individual could hold it twice in a lifetime. Jurors were paid specifically so that poor citizens could serve, which kept the random sample representative. The point of the machinery is that the composition of a deciding body is unknowable at the moment a bribe would have to be offered. It is not that the participants are virtuous. It is that there is nothing stable to optimize against. How it failed: randomly selected panels cannot be bribed but they can be inflamed. The same courts condemned Socrates and, after a naval engagement, tried and executed a group of generals collectively in violation of their own procedural law. Sortition defeats targeted capture and does nothing whatever about a stampede.

Every mechanism above was built in response to a failure. Double-entry answers the dishonest bookkeeper, Sarbanes-Oxley answers Enron, Registered Reports answer the replication crisis. That pattern invites a comfortable conclusion: that institutions learn, and that each exploit discovered makes the next one harder.

The clearest evidence against it comes from NASA, twice.

The commission that investigated Challenger included Richard Feynman, who conducted much of his own inquiry by ignoring the official hearing schedule and talking to working engineers directly. He attached a personal appendix to the report. It runs four pages, and it is not really about O-rings. It is about a gap. Asked for the probability that a launch would end in failure, NASA’s engineers gave figures around one in a hundred. NASA’s management gave one in a hundred thousand. Feynman’s summary was that management “claims to believe the probability of failure is a thousand times less.”52 The organization was flying on a number that nobody who touched the hardware believed.

His closing sentence is the best one-line statement of this essay’s subject that anyone has managed:

“For a successful technology, reality must take precedence over public relations, for nature cannot be fooled.”

That is the map and the territory, in 1986, with seven people already dead.

Seventeen years later, having redesigned the joint that failed and absorbed a global inquiry, NASA lost Columbia. The board investigating that accident quoted Feynman’s appendix back into evidence, and then recorded its own first conclusion: “the causes of the institutional failure responsible for Challenger have not been fixed.”53 The engineering correction took. The organizational one did not. A feedback loop is a design, not a guarantee, and the component that decides which failures get absorbed is itself a component that can fail.

And the best empirical account of what actually works comes from Elinor Ostrom, who studied hundreds of real institutions that governed shared resources for centuries. Two of her eight design principles carry the load. First, monitors are drawn from the population being monitored, or are directly accountable to it, an explicit rejection of the external-auditor model. Second, sanctions are graduated rather than binary, because most first violations are errors or desperation, and a system that answers every deviation with maximum force destroys the cooperation that made monitoring affordable in the first place.53 Set that against 41% compliance with an unenforced federal statute and the lesson sharpens: a mild penalty reliably applied beats a severe penalty never levied.

X. What each side can take from the other

Now the speculative part, which is what I actually wanted to write about.

For AI safety, from institutional design.

Stop trying to make gaming unattractive. Make it unreachable. This is the strongest cross-domain regularity in the material. What works removes the opportunity to optimize: the target is unknown, or not yet formed, or physically unreachable at the moment optimization would have to occur. Athens rotating its highest office every twenty-four hours. Blinded analysis in physics, where the answer stays hidden until every choice is locked. Registered Reports accepting a paper before the data exists. The no-lone zone. Held-out splits. Frozen verifiers. Attempting to make gaming unappealing fails, reliably: the immutable-looking comment, the instruction not to reward hack, the penalty on visible reasoning, self-administered recusal, statutes whose penalties are never levied.

Graduated sanctions. This is the import I would most like to see someone attempt. I have found no evidence that anyone uses it in AI training at all.

Reward-hacking detection today is binary. A run is caught or it is not, and the response to catching one is maximal: penalize it, filter it out of the data, retrain. Ostrom found that institutions built this way do not survive. The durable ones answered a first violation with a small penalty, and escalated only on repetition, and her explanation is worth stating carefully because it is not sentimentality. Most first violations are not defection. They are mistakes, misunderstandings of an ambiguous rule, or genuine emergencies. A system that answers all of them with maximum force teaches its participants that the monitoring relationship is adversarial, and once they believe that, they stop volunteering the information that made monitoring affordable in the first place. The sanction regime destroys its own evidence base.

Now read the concealment result again. Maximum pressure applied to a detection channel produced a model that kept cheating and stopped narrating it. That is the same failure in a different vocabulary: the response to detection was severe enough that the system’s cheapest available move was to degrade the monitor’s information rather than to comply. Ostrom’s commons and OpenAI’s training run failed in the same shape, four decades and one substrate apart.

What a graduated response looks like inside a training loop, I genuinely do not know. Downweighting rather than zeroing a suspect trajectory, escalating only on repeated patterns, keeping a hacked episode in the data with its label rather than deleting it. These are guesses, and the inoculation-prompting result suggests the design space is stranger than it looks. But the question seems worth more than it is currently getting, which is nothing.

Plant known defects and measure whether the detector finds them. This belongs to a family whose common move is to manufacture the failure you are worried about instead of waiting for it: mutation testing in software, fault injection in distributed systems, chaos engineering in production. Medicine ran the human version and learned that peer review catches 2.13 of 9 planted major errors. Nobody has built it for an AI validator. We have no measurement of any reward-hacking detector’s recall against known planted hacks. We do not know whether our detectors catch 90% or 25%, and every claim resting on “the monitor did not flag it” is uncalibrated until somebody finds out.

Retroactive invalidation. When a self-improving system’s retrospective review caught a node that had deleted the very markers its reward function used to detect cheating, the system kept the node. Finance and science both maintain machinery for stripping credit after the fact. Evolutionary AI systems have none. Every gain credited downstream of a hacked ancestor remains on the books.

Measure independence, not headcount. The 1986 multiversion result, the reward-model ensembles, the nine judges worth two votes, and parties collapsing the separation of powers are one finding in four costumes. If you run an ensemble of judges, you owe an effective-sample-size number, not a roster.

For institutional design, from AI safety.

The overoptimization curve is a general result about institutions. True performance under proxy optimization rises, peaks, and falls along a predictable curve whose shape depends on the proxy’s capacity. That is a formal, measurable version of what Campbell asserted qualitatively. It implies something institutions essentially never do: treat a target as carrying a finite optimization budget, and stop before you spend it. Not “targets are bad,” but “this target is good for approximately this much pressure.” No performance-management system I have encountered contains the concept.

Evaluation awareness is a general property of capable evaluated agents, and it rises with capability alone. Hospitals know when the inspector is coming. None of that is news. What is new is that somebody can now run the experiment.

In a language model you can reach into the reasoning trace, delete the sentences where the model notices it is being tested, and measure what its behavior does. Then you can insert such sentences and measure again. Deleting them raised covert action by eleven percentage points. Inserting them lowered it. Nobody has ever been able to do this to a hospital, because you cannot hold an institution fixed and toggle only its awareness of being watched. Inspectors have suspected the effect for as long as inspections have existed. It has just become measurable.

Adversarial scrutiny just became cheap, which changes what oversight can afford. Every mechanism in Section IX is rationed by the same scarce resource: expert attention. Auditors, reviewers, and investigators cost money, so institutions buy the least scrutiny they can tolerate. That constraint is now loosening fast.

Before celebrating, take the caveat the field learned the hard way. When AI-generated findings get checked adversarially, roughly four in five turn out to be plausible and wrong. In one case ten reviewers unanimously endorsed a cryptographic vulnerability that did not exist, and a single empirical test killed it. Worse, an oversight system that produces false alarms at scale does not merely waste time. It collapses. AI-generated vulnerability reports drove one open-source project’s bug bounty below 5% valid submissions and the maintainers shut the program down. Cheap scrutiny is only an asset when it arrives with cheap refutation.

XI. The honest claim

I want to close without overclaiming, because the research does not support the triumphant version.

Notice what the strongest results in this whole body of work have in common. They are impossibility results. Regression cannot be engineered away, private information rules out full efficiency, and the only rule immune to strategic lying is the one that ignores everybody but a single person. The deepest things the field knows are statements about what cannot be built.

Two and a half centuries of institutional design never produced a trustworthy system out of untrustworthy parts. What it produced were systems whose failures are slower, more visible, and more correctable than the failures of the parts. Enron happened, and then Sarbanes-Oxley happened, and audits still come back deficient about 40% of the time. Psychology’s replication crisis happened, and Registered Reports happened, and the confirmation rate fell from 96% to 44%, not to zero. Athens randomized its juries into unbribability and then executed Socrates.

And there is a price even for that. A formal result on highly optimized systems holds that tuning a system against the disturbances you have anticipated makes it more fragile to the ones you have not.54 Every countermeasure in Section VII is an adaptation to a failure mode somebody already found. Which means the defensive portfolio a mature system accumulates is also a map of where it has stopped looking. Sarbanes-Oxley made one kind of audit failure harder and did nothing whatever about the kinds nobody had seen yet. A held-out set closes the channel it closes. The work is never finished, not because we are bad at it, but because each fix relocates the exposure somewhere nobody is watching.

That is the honest target for AI alignment, and stating it plainly is more useful than the alternatives. Not an oracle: in the general case, one provably does not exist. Not a system that cannot be gamed: everything can be gamed, and the best evidence we have says that pressuring a system not to game it teaches concealment. The target is a system in which misalignment is costly, detectable, and correctable. Where the gap between proxy and oracle is bounded. Where excursions into it leave traces. Where the traces are read by something that is not a subroutine of the process being watched. And where the response is graduated enough that the watching stays affordable.

The male butterfly has none of this. He has one rule, no ceiling, and no way to notice that the orange thing spinning in the clearing is not what evolution meant. He is not failing. He is succeeding, perfectly, at something else. The only reason we can tell the difference is that we are standing outside the loop, holding the motor.

That vantage point is the entire asset. Everything in this essay is an attempt to build a version of it that survives contact with something smarter than we are.


Sources

Additional sources for Section IX, in order of appearance: Sarbanes-Oxley Act of 2002, Pub. L. 107-204, §§201, 203, 206; PCAOB, Staff Update on 2023 Inspection Activities (August 2024) and 2024 Inspection Activities (March 2025); Pew Research Center (2023) analysis of Administrative Office of the U.S. Courts, Judicial Business 2022, Table D-4; National Research Council (2009), Strengthening Forensic Science in the United States: A Path Forward; Department of Justice, FBI, Innocence Project and NACDL joint release, 20 April 2015; US Air Force Instruction 91-104 (2013), §§1.1–1.3, with the promotion-criteria confirmation from the Department of Defense press briefing of 27 March 2014; 28 U.S.C. §455(a), with recusal figures from a Wall Street Journal investigation of 28 September 2021; Aristotle, Athenaion Politeia, chs. 43, 44.1, 62.2–3, 63.

Footnotes

  1. Magnus, D. (1958). “Experimentelle Untersuchungen zur Bionomie und Ethologie des Kaisermantels Argynnis paphia L.” Zeitschrift für Tierpsychologie 15: 397–426. The paper has never been digitized and carries no DOI; the figures here reach us through Kral, K. (2016), “Implications of insect responses to supernormal visual releasing stimuli,” European Journal of Entomology 113: 429–437, doi:10.14411/eje.2016.056 (open access), which reads the German original. Kral juxtaposes Magnus’s behavioral 140 Hz result with a separately measured figure for insect flicker resolution; the connection between the two is his argument, not Magnus’s finding. The size preference held only where the model’s movement roughly resembled a female’s.

  2. Staddon, J. E. R. (1975). “A note on the evolutionary significance of ‘supernormal’ stimuli.” The American Naturalist 109(969): 541–545, doi:10.1086/283025. Staddon also links supernormality to peak shift in discrimination learning.

  3. Campbell, D. T. (1979). “Assessing the impact of planned social change.” Evaluation and Program Planning 2(1): 67–90, doi:10.1016/0149-7189(79)90048-X, circulating from a 1974 OECD presentation and an earlier 1976 occasional paper. On priority: Rodamar, J. (2018), “There ought to be a law! Campbell versus Goodhart,” Significance 15: 9, which concludes Campbell stated the principle to national and international audiences first. Goodhart’s own 1975 formulation, delivered at a Reserve Bank of Australia conference, is narrower and concerns statistical regularities collapsing under control pressure. The aphorism usually quoted as his, “when a measure becomes a target, it ceases to be a good measure,” is Marilyn Strathern’s sentence from a 1997 paper on university audit culture, European Review 5(3): 305–321, p. 308.

  4. Bevan, G. & Hood, C. (2006). “What’s measured is what matters: targets and gaming in the English public health care system.” Public Administration 84(3): 517–538, doi:10.1111/j.1467-9299.2006.00600.x, at p. 532, reporting Public Administration Select Committee (2003), para 52. The same paper documents ambulance response times “corrected” to fall under eight minutes in a third of trusts, and a four-hour accident-and-emergency target officially met by 96% of trusts in 2004–05 where the independent patient survey put the figure at 77%.

  5. Manheim, D. & Garrabrant, S. (2019). “Categorizing Variants of Goodhart’s Law.” arXiv:1803.04585 (v4). On the first variant: “No matter what measure is chosen for optimization, an inexact metric necessarily leads to a divergence between the goal and the metric in the tail.” The authors caution in a footnote that because the original terms were never laid out formally, their categories do not map exactly onto prior usage.

  6. Royal Swedish Academy of Sciences (2007). Scientific Background on the Sveriges Riksbank Prize in Economic Sciences 2007: Mechanism Design Theory. PDF. Source for incentive compatibility, the revelation principle, and the impossibility results (Hurwicz 1972; Gibbard 1973 and Satterthwaite 1975). The bilateral-trade result is credited jointly to Laffont & Maskin (1979) and Myerson & Satterthwaite (1983), whose own abstract carries the qualifier “without outside subsidies.” One adjacent claim is easy to overstate: the Vickrey-Clarke-Groves family cannot balance its budget within dominant-strategy mechanisms, but d’Aspremont & Gérard-Varet (1979) obtain budget balance and full efficiency in the Bayesian version, paying for it in participation constraints. It is a three-way tradeoff, not a wall. 2

  7. Vickrey, W. (1961). “Counterspeculation, Auctions, and Competitive Sealed Tenders.” The Journal of Finance 16(1): 8–37, doi:10.1111/j.1540-6261.1961.tb02789.x.

  8. Ausubel, L. M. & Milgrom, P. (2005). “The Lovely but Lonely Vickrey Auction.” In Cramton, Shoham & Steinberg (eds.), Combinatorial Auctions, MIT Press, pp. 17–40. Their Theorem 10 is what licenses treating the four weaknesses as one failure: low revenue, vulnerability to losing-bidder collusion, vulnerability to a single bidder entering under several identities, and non-monotonic outcomes are equivalent conditions, none of which arise when the goods are substitutes. The auction is safe where items are interchangeable and unsafe where they combine.

  9. Schultz, W., Dayan, P. & Montague, P. R. (1997). “A neural substrate of prediction and reward.” Science 275(5306): 1593–1599. Quoted formulation from Schultz, W. (1998), “Predictive reward signal of dopamine neurons,” J Neurophysiol 80(1): 1–27.

  10. Berridge, K. C., Venier, I. L. & Robinson, T. E. (1989). Behavioral Neuroscience 103(1): 36–45. See also Cannon, C. M. & Palmiter, R. D. (2003), “Reward without dopamine,” J Neurosci 23(34): 10827–10831, which finds normal sucrose preference alongside a genuine deficit in goal-directed behavior.

  11. Salamone, J. D. et al. (1991). Psychopharmacology 104(4): 515–521; review in Salamone & Correa (2012), Neuron 76(3): 470–485. Note that Schultz’s prediction-error account and Berridge’s incentive-salience account are rival interpretations of overlapping data rather than a single consensus; they agree that dopamine is not the pleasure signal and disagree about what it is.

  12. Hall, K. D. et al. (2019). “Ultra-processed diets cause excess calorie intake and weight gain.” Cell Metabolism 30(1): 67–77.e3, doi:10.1016/j.cmet.2019.05.008. Twenty subjects, single site, not independently replicated at scale.

  13. Fazzino, T. L., Rohde, K. & Sullivan, D. K. (2019). Obesity 27(11): 1761–1768, doi:10.1002/oby.22639. 4,795 of 7,757 foods in the USDA database met one of the three cluster definitions.

  14. DiFeliceantonio, A. G. et al. (2018). “Supra-Additive Effects of Combining Fat and Carbohydrate on Food Reward.” Cell Metabolism 28(1): 33–44.e3. Single study, not independently replicated; the published sample size could not be verified from open sources.

  15. Flagel, S. B. et al. (2011). “A selective role for dopamine in stimulus–reward learning.” Nature 469(7328): 53–57, doi:10.1038/nature09588; Meyer, P. J. et al. (2012), PLoS ONE 7(6): e38987, pooling 1,878 animals under a common protocol.

  16. Karpathy, A. (2024). “RLHF is just barely RL.” X post, 7 August 2024.

  17. Gao, L., Schulman, J. & Hilton, J. (2022). “Scaling Laws for Reward Model Overoptimization.” arXiv:2210.10760. 2

  18. Anthropic (2025). System Card: Claude Opus 4 & Claude Sonnet 4, §6, Table 6.2.A. The 78%/80% figures are Claude Sonnet 3.7; the 47%/5% figures are Claude Opus 4. For the taxonomy of exploitation methods see Zhong, Z., Raghunathan, A. & Carlini, N. (2025), “ImpossibleBench,” arXiv:2510.20270.

  19. METR (2025). “Recent Frontier Models Are Reward Hacking,” 5 June 2025. 39 of 128 runs (30.4%) versus 8 of 1087 (0.7%). METR notes the two benchmark families differ in task difficulty and scaffolding as well as grader visibility, and hedges the causal reading accordingly.

  20. METR (2026). “Frontier Risk Report (February to March 2026),” 19 May 2026. The sample size behind the ~80% figure is not disclosed in the report.

  21. Baker, B., Huizinga, J., Gao, L. et al. (2025). “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.” arXiv:2503.11926.

  22. Taufeeque, M., Heimersheim, S., Gleave, A. & Cundy, C. (FAR AI). “The Obfuscation Atlas.” arXiv:2602.15515, ICML 2026.

  23. MacDiarmid, M., Wright, B., Uesato, J. et al. (2025). “Natural Emergent Misalignment from Reward Hacking in Production RL.” arXiv:2511.18397.

  24. Taylor, M., Chua, J., Betley, J., Treutlein, J. & Evans, O. (2025). “School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs.” arXiv:2508.17511.

  25. karpathy/autoresearch, GitHub Discussion #285 / Issue #278 (15 March 2026) and Issue #131; the pinned validation shard is visible in prepare.py. A separate issue (#599) documents that the reported metric is printed rather than bound to any hash of code, data, or model, so a fabricated score is accepted and retained. The personal experiments described are unpublished: a deterministic stdlib-only simulator, and a two-arm evolutionary experiment whose search, judge, and mutation providers are rule-based simulators rather than live models, which demonstrates the mechanism rather than model behavior.

  26. Bostrom, N. (2003). “Ethical Issues in Advanced Artificial Intelligence.” In Smit et al. (eds.), Cognitive, Emotive and Ethical Aspects of Decision Making in Humans and in Artificial Intelligence, Vol. 2, pp. 12–17, §§2 and 5. Developed at length in Bostrom, N. (2014), Superintelligence, Oxford University Press, and grounded in the orthogonality and instrumental-convergence theses of Bostrom, N. (2012), “The Superintelligent Will,” Minds and Machines 22(2): 71–85.

  27. Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. New York: Viking.

  28. Loosemore, R. P. W. (2014). “The Maverick Nanny with a Dopamine Drip: Debunking Fallacies in the Theory of AI Motivation.” AAAI Technical Report SS-14-03.

  29. Drexler, K. E. (2019). “Reframing Superintelligence: Comprehensive AI Services as General Intelligence.” FHI Technical Report #2019-1, University of Oxford.

  30. Novikov, A., Vũ, N., Eisenberger, M. et al. (2025). “AlphaEvolve: A coding agent for scientific and algorithmic discovery.” Google DeepMind, arXiv:2506.13131. Limitation quote at §6, p. 21. Conditions that must travel with the 48-multiplication result: it uses complex-valued multiplications, and the 56-year framing is scoped to fields of characteristic 0. The fourteen improved targets come with two acknowledged regressions. Infrastructure figures are Google self-reports on internal systems and are not externally verifiable. 2

  31. Mankowitz, D. J. et al. (2023). “Faster sorting algorithms discovered using deep reinforcement learning.” Nature 618(7964): 257–263. The largest speedup is specific to five-element sequences. Predecessor: Fawzi, A. et al. (2022), Nature 610(7930): 47–53, whose 47-multiplication result holds over the field with two elements.

  32. Hubert, T. et al. (2025). “Olympiad-level formal mathematical reasoning with reinforcement learning.” Nature 651: 607–613, published online 12 November 2025.

  33. Bondarenko, A., Volk, D., Volkov, D. & Ladish, J. (2025). “Demonstrating specification gaming in reasoning models.” Palisade Research, arXiv:2502.13295. The authors note the study covers a single environment and that judges disagree on up to 25% of runs.

  34. Jumper, J. et al. (2021). “Highly accurate protein structure prediction with AlphaFold.” Nature 596(7873): 583–589.

  35. Zelikman, E., Lorch, E., Mackey, L. & Kalai, A. T. (2023). “Self-Taught Optimizer (STOP).” arXiv:2310.02304, Table 2. The 0.42% to 0.46% difference is statistically indistinguishable; framings describing the warning as having made matters worse read noise as signal.

  36. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O. & Roth, A. (2015). “The Reusable Holdout: Preserving Validity in Adaptive Data Analysis.” Science 349(6248): 636–638; Blum, A. & Hardt, M. (2015), “The Ladder,” arXiv:1502.04585.

  37. Eisenstein, J. et al. (2023). “Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.” arXiv:2312.09244.

  38. “Nine Judges, Two Effective Votes,” arXiv:2605.29800. Measured on one task family with dense human labels; the ratio is not a universal constant.

  39. Knight, J. C. & Leveson, N. G. (1986). “An experimental evaluation of the assumption of independence in multiversion programming.” IEEE Transactions on Software Engineering 12(1): 96–109.

  40. Anthropic (2025). System Card: Claude Sonnet 4.5, §7.2, §7.3, §7.6. Secondary write-ups citing 9% or 12% are quoting different quantities from the same document.

  41. Schoen, B., Nitishinskaya, E., Balesni, M. et al. (2025). “Stress Testing Deliberative Alignment for Anti-Scheming Training.” arXiv:2509.15541, §6. The authors describe the capability-only comparison as suggestive rather than a rigorous ablation, and state they cannot exclude that the measured reductions are partly driven by situational awareness.

  42. Needham, J., Edkins, G., Pimpale, G., Bartsch, H. & Hobbhahn, M. (2025). “Large Language Models Often Know When They Are Being Evaluated.” arXiv:2505.23836. The human agentic score is acknowledged by the authors to be inflated by dataset familiarity, which narrows the human-model gap.

  43. Irving, G., Christiano, P. & Amodei, D. (2018). “AI safety via debate.” arXiv:1805.00899.

  44. Burns, C., Izmailov, P., Kirchner, J. H. et al. (2023). “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.” arXiv:2312.09390.

  45. Levinson, D. J. & Pildes, R. H. (2006). “Separation of Parties, Not Powers.” Harvard Law Review 119: 2311. The forbearance argument is Levitsky, S. & Ziblatt, D. (2018), How Democracies Die, Crown. Madison quotations are from Federalist 51 and 10; Montesquieu is The Spirit of Laws (1748), Book XI, ch. 4, where the argument about power checking power appears (ch. 6 carries the three-powers analysis).

  46. Schroter, S., Black, N., Evans, S., Carpenter, J., Godlee, F. & Smith, R. (2004). “Effects of training on quality of peer review: randomised controlled trial.” BMJ 328(7441): 673.

  47. Simmons, J. P., Nelson, L. D. & Simonsohn, U. (2011). “False-Positive Psychology.” Psychological Science 22(11): 1359–1366, Table 1.

  48. Open Science Collaboration (2015). “Estimating the reproducibility of psychological science.” Science 349(6251): aac4716. See also Klein, R. A. et al. (2018), “Many Labs 2,” AMPPS 1(4): 443–490, where 54% replicated and the failures were not explained by variation in sample or setting.

  49. Errington, T. M. et al. (2021). “Investigating the replicability of preclinical cancer biology.” eLife 10: e71601, with the methods-specification finding in the companion paper, eLife 10: e67995.

  50. Scheel, A. M., Schijen, M. R. M. J. & Lakens, D. (2021). “An Excess of Positive Results.” AMPPS 4(2), doi:10.1177/25152459211007467. 146 of 152 against 31 of 71. Registered Reports are not a random sample and skew toward replications, though restricting the comparison to original studies still leaves a 46-point gap. 2

  51. Turner, E. H., Matthews, A. M., Linardatos, E., Tell, R. A. & Rosenthal, R. (2008). “Selective publication of antidepressant trials and its influence on apparent efficacy.” New England Journal of Medicine 358(3): 252–260. Compliance figures from DeVito, N. J., Bacon, S. & Goldacre, B. (2020), The Lancet 395(10221): 361–369.

  52. Feynman, R. P. (1986). “Personal Observations on Reliability of Shuttle,” Appendix F to the Report to the President by the Presidential Commission on the Space Shuttle Challenger Accident, June 6, 1986, Volume II, pp. F-1 to F-5. Feynman was one of thirteen commission members, not its chair; William P. Rogers chaired it. Appendix F gives the overall range as “roughly 1 in 100 to 1 in 100,000,” noting that “the higher figures come from the working engineers, and the very low figures from management.” Higher there means higher probability of failure, so 1 in 100 is the engineers’ estimate; careless paraphrases invert this. The phrase quoted here is exact: management “claims to believe the probability of failure is a thousand times less.” The appendix separately reports 1/10,000 from Rocketdyne engineers, 1/300 from Marshall engineers, and 1/100,000 from NASA management for the main engines specifically. Two frequently repeated claims are not in this document and should not be cited to it: the figure of 1 in 200 (that belongs to Feynman’s memoir account of questioning engineers directly, in “What Do You Care What Other People Think?”, W. W. Norton, 1988), and the story that management derived its number by working backward from a desired conclusion. The O-ring insight that Feynman demonstrated with ice water at the February 11, 1986 hearing was, by his own repeated account, prompted by fellow commissioner Maj. Gen. Donald Kutyna: “But it was his idea.”

  53. Ostrom, E. (1990). Governing the Commons. Cambridge University Press, Table 3.1. Empirical review of 91 studies, including the standing critiques: Cox, M., Arnold, G. & Villamayor Tomás, S. (2010), Ecology and Society 15(4): 38. Ostrom’s cases are small, geographically bounded, and high-repeat-interaction; the principles have no demonstrated track record on global commons. 2

  54. Carlson, J. M. & Doyle, J. (2002). “Complexity and robustness.” PNAS 99(suppl 1): 2538–2545, doi:10.1073/pnas.012582499. Their Highly Optimized Tolerance framework holds that configurations optimized against a modeled disturbance acquire fragility to changes in the distribution of that disturbance and to design flaws. Woods (2015) restates the same point independently for human systems: expanding a system’s envelope in one direction increases its vulnerability in others. Note that this is a claim about optimized systems generally and its application to institutional countermeasures is my extension, not the authors’.