The Proof Was Never the Point

Twenty-five Fields medallists just described a measurement failure that is sitting on your dashboard

Line drawing of a chalkboard proof beside a token meter, the two connected by a cord that has been cut
the-proof-was-never-the-point Twenty-five Fields Medal winners say the AI benchmark race is damaging mathematics. The real finding is narrower and applies to every metric you own: a proxy works only while producing it stays expensive. AI benchmarks, Goodhart law, proxy metrics, Erdos problems, measurement, engineering metrics, Fields Medal declaration, PageRank

Paul Erdős paid for his problems out of his own pocket. Ten dollars for a small one, thousands of dollars for a hard one, settled in cash when somebody solved it. Those same problems are now priced in tokens. The unit changed, and the unit was the whole measurement.

TL;DR

Write the correspondence clause for every metric you track: "this number tracks X because producing it requires Y." Price Y today, not when you adopted the metric. A spread that bunches up is a reason to look, not a verdict: ask whether the metric can now be raised by a route that does not touch the outcome. Take a decoupled metric out of the compensation plan before you take it off the dashboard.

On September 11, twenty-five holders of the Fields Medal signed a joint declaration. The signatures run from Pierre Deligne, who won in 1978, to Yu Deng, who won this year. Terence Tao, Peter Scholze, Maryna Viazovska, Maxim Kontsevich, Cédric Villani. You do not assemble that list for a press cycle.

Their sentence is short: "the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community."

The easy reading is guild protectionism. Famous people, threatened by a machine, dressing up their status anxiety as concern for the discipline. I understand why that reading is tempting, because I have used it myself on other professions, and because a letter signed by twenty-five laureates does look like an institution defending its gate.

The easy reading costs you something specific. These are the people with the oldest and most precisely stated benchmark in intellectual work, which makes them the first to notice it might have stopped telling them what it used to. Not that anyone cheated. That the thing it measured and the thing it now measures have come apart. Every proxy metric you own is exposed the same way, without the centuries of baseline that let them see it.

What the bounty was actually buying

Erdős posed thousands of problems and attached money to many of them, a playful incentive from a man who lived out of a suitcase and gave away most of what he earned. The declaration describes what those problems became:

"Famous problems have often served as landmarks and lighthouses against which one can measure an improved understanding of this landscape. Solving one of these problems has been a certain sign of new insights and interesting methods, which would then be studied by a community of mathematicians, through a long and arduous process of talks, discussions, simplifications."

Read that as a measurement spec, because that is what it is. The claim is not that solving a famous problem is valuable in itself. It is that a solution was a certain sign of something else: new insight, transferable method. The proof was the receipt. The method was the goods.

That correspondence held for an unglamorous reason. Producing the receipt and acquiring the goods took the same scarce input, a trained person's years. You could not have one without the other, so counting one told you about the other. Nobody enforced it. It was arithmetic about where the cost sat.

The list itself is recent. In early 2023 a mathematician named Thomas Bloom started collecting Erdős problems into a website, mostly so he could look them up from anywhere, expecting nobody to use it. By August 2025 he had catalogued close to a thousand. Over 2024 and the first eight months of 2025, 111 problems moved from open to solved on his list, though some of those had been solved years earlier and the change only recorded a rediscovery. Then the list became a scoreboard.

The unit changed

On May 20, 2026, OpenAI announced that an internal model had produced a counterexample to the unit distance problem, which Erdős conjectured in 1946. It was the first historically significant proof to come from an AI model. It was not a trick: the result brought in ideas from a distant branch of mathematics that nobody had successfully applied to that problem, and within days related techniques were being used elsewhere. Human mathematicians substantially improved on it within weeks. On August 1 OpenAI announced ten more advances from an unreleased model called Astra, including three further Erdős problems.

Here is the sentence from Quanta's account that I have not been able to put down. As the problems became an informal benchmark, their solutions "are being discussed in terms of their 'per-problem cost' — the price of the tokens needed to solve them."

Erdős paid dollars, out of his own pocket, for insight.
We pay tokens, on a corporate card, for the artifact.
Same problems. Different denominator.

A benchmark denominated in tokens is still measuring a real thing. It is measuring inference spend. Whether it is still measuring what the bounty was for is the question, and it is not one the price answers.

The metric can fail without anyone cheating

The reflex here is to say Goodhart's law and move on. The phrasing everyone quotes is Marilyn Strathern's, from a 1997 essay on audit culture in British universities: "When a measure becomes a target, it ceases to be a good measure."

Almost everyone stops at that line. Her next sentence is the useful one: "The more a 2.1 examination performance becomes an expectation, the poorer it becomes as a discriminator of individual performances."

Discriminator. Strathern is not primarily describing cheating. She is describing a measure losing its power to separate one thing from another. The name Goodhart's law sits on that same page, where she reports Hoskin using it for exactly this, so what follows is a distinction inside the idea rather than an escape from it.

The distinction still decides what you do next. The story most people carry is the adversarial one: somebody distorts their behaviour to hit the number, the number detaches from reality, and the fix is better incentives or a harder-to-fake measure. Nothing like that happened here. Nobody gamed the Erdős problems. The proofs are real. The unit distance counterexample was checked. The 1196 result carries Tao's name as a co-author. There is no fraud to find, and if you go looking for one you will waste the quarter.

What collapsed was the spread. A measure discriminates only across a population for whom producing it is expensive in the same way the underlying quality is expensive. Change what it costs to produce, and the measure keeps returning valid numbers that no longer separate anybody from anybody. It does not go wrong. It goes quiet, while continuing to report, and unlike gaming it leaves no fingerprints to find.

Brin and Page wrote their correspondence down, and it still broke

I was writing crawlers and servers from scratch by 1993, so I had a front-row seat to the last time this happened at scale, and it is worth being precise about how it went.

The 1998 paper that introduced PageRank is unusual in that its authors stated their assumption out loud. PageRank, they wrote, is "an objective measure of its citation importance that corresponds well with people's subjective idea of importance. Because of this correspondence, PageRank is an excellent way to prioritize the results of web keyword searches."

Because of this correspondence. The method was never a claim that links are good. It was a claim that links corresponded to judgement, and in 1998 placing one cost you a few minutes and your own page's standing next to it.

A link was a vote while placing one cost the linker something.
When placing one cost nothing, it was still a link.
It was no longer a vote.

They were already defending against cheapness in the design. PageRank does not count backlinks equally; a page divides its score between everything it points at, so pointing at everything is worth less per link. Watch it iterate on a four-page graph and you can find one where the page with three backlinks loses to the page with two. The defence was real, and link farms outgrew it anyway. My reading is that the scarce thing being spent was the linker's own standing, and a farm has none to spend. That is interpretation, not a variable in the algorithm. The correspondence clause is the part actually in the paper.

Then emitting a link became nearly free, the correspondence decayed, and the response was not to count links harder. It was to rework the signal and surround it with others that were still expensive to fake. Google's own documentation says how PageRank works "has evolved a lot since then, and it continues to be part of our core ranking systems." That is a better ending than abandonment: when a correspondence breaks you rebuild it and add independent evidence. What you do not get to do is keep reading the old number the old way.

Brin and Page are unusual only in having printed their correspondence clause. Every metric you run has one, almost none of them written down, which means nobody on your team can tell you whether it still holds.

The half that did not get cheaper

Here is the part that does not resolve itself, and it is not about mathematics at all.

Producing proofs got cheap. Reading them did not. Bloom, who built the list, put it plainly: "We're seeing a lot more of these 100- to 200-page papers that people are posting. 'I solved this theorem; I got AI to generate the proof and check the proof and write the paper.' But no human has read it, and no human is going to read it."

On Christmas morning a contributor posted what he believed was the first fully autonomous LLM resolution of an Erdős problem. Hours later somebody pointed out that Erdős had resolved it himself in a paper published in 1977. His reply is the most honest sentence in the story: "As someone who has fallen for this twice now, it's quite gut-wrenching."

That is not a hobbyist problem, and the corporate version is more interesting than the embarrassment it got reported as. In October 2025, OpenAI vice-president Kevin Weil posted that GPT-5 had "found solutions to 10 (!) previously unsolved Erdős problems and made progress on 11 others." Bloom called it "a dramatic misrepresentation" and explained why, and his explanation is the whole essay in one sentence: a problem marked open on his site only means "I personally am unaware of a paper which solves it."

So the model had not failed at anything. In Bloom's words it "found references, which solved these problems, that I personally was unaware of." An OpenAI researcher conceded that "only solutions in the literature were found" and then said the thing I keep coming back to: "I know how hard it is to search the literature."

He is right, and that is the point. The label on that website meant "not known to one careful human" because that was the only thing it could affordably mean, and everyone read it as "not known to anybody." The artifact was real and correct. The inference drawn from it was the part that had quietly stopped being safe.

The economics: when two outputs share one scarce input, their costs move together and either can stand in for the other. Introduce a substitute for one input and the two prices come apart. That is the whole mechanism, and it is worth being exact about which half got a substitute here. Checking that a proof holds together is not the bottleneck; Barreto had a tool called Aristotle certify one, and formal checking keeps getting cheaper. What has no substitute yet is deciding whether a result matters, noticing it was already known, and carrying it into the canon so the next person can use it. That is the expensive half, and it is the half a token price does not buy.

Audit your own correspondences

I spent 1992 and 1993 being paid to find bugs, first on games at Sierra and then on PageMaker at Aldus. The metric was bugs filed. It worked, and it worked for a boring reason: filing one required reproducing it by hand, and reproducing it by hand required understanding the product. The count tracked comprehension because the only route to the count ran through comprehension. Take that route away and the same number means nothing, which is the story of every automated-submission quality program I have watched since.

So here is what to do instead. The question is not whether this is a good metric, which nobody can answer. It is narrower.

For each number your team is measured on:
  1. Write the correspondence clause. One sentence, this exact shape: "This number tracks X because producing it requires Y." If nobody can write the sentence, you are not tracking anything, and that is the finding.
  2. Price Y today. Not when the metric was adopted. Today, including what a model does for free.
  3. Then go and check. A cheap Y is a reason to look, not a verdict.

Step three is the one people skip. It starts as a query:

-- SQLite. A screen, not a verdict: has the spread bunched up?
-- Bunching is a reason to go and look at the per-service numbers.
-- It is not on its own evidence that anything has decoupled.
SELECT
  strftime('%Y-%m', measured_on)                  AS month,
  COUNT(DISTINCT service)                         AS services,
  ROUND(AVG(coverage_pct), 1)                     AS mean_pct,
  ROUND(MAX(coverage_pct) - MIN(coverage_pct), 1) AS spread_pct,
  ROUND(AVG(escaped_defects), 2)                  AS escaped
FROM service_quality
GROUP BY month
ORDER BY month;

That is SQLite. Only strftime needs translating for Postgres, which takes an output alias in GROUP BY perfectly happily. One row per service per month is assumed; if your table stores several samples per service, aggregate to one row first or the spread will measure your sampling rather than your services.

Now the important part, because I got this wrong in the draft and a reader caught it. A rising mean with a collapsing spread and a flat outcome is not a verdict. Here is a table that trips all three conditions while the metric is working perfectly:

JanuaryA 40% / B 60% / C 80%
FebruaryA 94% / B 96% / C 98%
escaped defects, both monthsA 6 / B 4 / C 2
SPREAD40 points, then 4

Every alarm I originally wrote goes off. The metric still ranks the three services in exactly the right order, with a correlation of minus one against escaped defects in both months. Nothing has decoupled. The services improved, unevenly, which is what improvement looks like.

Bunching has at least three innocent explanations: real convergence, a ceiling, or noise now larger than the gaps that remain. What the query cannot tell you is the thing you care about, which is whether the number still tracks the outcome per service rather than on average.

What to do when the screen fires:
  1. Stop aggregating. Line the metric up against the outcome one unit at a time, over a horizon long enough that the outcome could have moved.
  2. Ask whether the ordering still holds. If the high-coverage services still ship fewer escaped defects, the metric is doing its job on a narrower range.
  3. Ask whether the metric can now be raised by a route that does not touch the outcome. That is the real question, and automation is the most common new route.
  4. Only then touch anything. A screen that fires is a reason to open a review, never a reason to change somebody's compensation.

This is the same machinery underneath the test coverage lie and observability theater, and it is why legibility used to be free is a sentence about cost rather than about design. It is also most of what makes startup metrics theater so durable: the numbers in a board deck are rarely fabricated. They are usually just no longer discriminating.

Where the medallists may be wrong

There are exceptions to the gloom, and the declaration is more careful than its coverage. It says outright that AI "offers the potential of enhancing and accelerating genuine mathematical study and understanding," and that whether the change helps or harms "will in large part be determined by the decisions of the humans in control of this new technology." That is a claim about governance, not a prophecy about machines.

The counter-evidence is strong. OpenAI's own announcement says the unit distance proof was checked by external mathematicians who have "also written a companion paper explaining the argument and providing further background and context for the significance of the result." Checked, explained, contextualised, handed on. That is precisely the long human process the declaration fears losing, and on the most famous example anyone has, it ran. One active contributor describes his own practice as reading an idea from a model and then digesting, simplifying and generalising it, which is the same chain with a machine in it.

So I want to be careful about what I am claiming, because an earlier draft of this was not. I wrote that the collapse of the benchmark was settled. It is not, and the declaration does not say it is. Here is the claim I will defend instead: an output can stay good while what it tells you about whoever produced it gets worse, and the price of producing it is a hint that this may have happened rather than evidence that it has. Whether it has is an empirical question about your own numbers, and it is answerable.

Audit One Metric Before Monday

One number, four questions. The whole point is that you can finish this before the week starts.

  • Pick the single number your team is judged on, and write its correspondence clause in one sentence: "this number tracks X because producing it requires Y." If nobody can write the sentence, that is the finding.
  • Price Y as it costs today, including whatever a model now does for free. Compare that with what it cost when the metric was adopted.
  • Query the spread, not the average. Group the metric by month and track max minus min alongside the outcome it stands for.
  • If the range is collapsing while the outcome stays flat, pull the metric out of the compensation plan first. Unwinding comp takes a quarter; changing a dashboard takes an afternoon.
  • Keep the metric for finding gaps. A decoupled proxy is still a fine flashlight and a terrible scoreboard.

The Bottom Line

A proxy metric is a bet that producing the artifact stays expensive in the same way the real goal is expensive. That bet is never written down, and it is the entire basis of the measurement.

Machines that make artifacts cheap are not gaming your metrics. They are settling that bet quietly, in your favour on cost, and possibly against you on information. Possibly. The falling price is the reason to go and look; it is not the finding.

So pick the one number your team is judged on. Write its correspondence clause in a single sentence. Price what that sentence assumes is expensive. If the answer is "not much, anymore," go and check whether the number still tracks the outcome one unit at a time. Keep it if it does. If it does not, change it, and change the compensation plan last and deliberately rather than first and in a hurry.

"A proxy metric is a bet that producing the artifact stays expensive in the same way the real goal is expensive. That bet is never written down, and it is the entire basis of the measurement."

Sources

  • Declaration — Math and AI — Joint declaration signed by twenty-five Fields Medal recipients, from Pierre Deligne (1978) to Yu Deng (2026), on the misalignment between AI companies' benchmark race and the goals of mathematics.
  • Why the Legendary Erdős Problems Are Falling to AI — Konstantin Kakaes on how Thomas Bloom's erdosproblems.com became an informal AI benchmark: the May 20 2026 unit distance counterexample, the Astra results of August 1, the 111 problems reclassified over 2024-2025, and solutions now discussed in terms of per-problem token cost.
  • 'Improving ratings': audit in the British University system — Marilyn Strathern, European Review 5(3), 1997, pp. 305-321. Source of the canonical phrasing of Goodhart's law and, on the same page, the sharper claim that the measure degrades as a discriminator of individual performances.
  • The Anatomy of a Large-Scale Hypertextual Web Search Engine — Brin and Page, 1998. States the PageRank correspondence assumption explicitly: the measure works because citation importance corresponds to people's subjective idea of importance.
  • OpenAI claims it solved an 80-year-old math problem — for real this time — Rebecca Bellan on the unit distance disproof, and on the earlier episode in which an OpenAI vice-president claimed GPT-5 had solved ten open Erdos problems that turned out to be already solved in the literature.

Disagree? Have a War Story?

I read every reply. If you've seen this pattern play out differently, or have a counter-example that breaks my argument, I want to hear it.

Send a Reply →