When Solving Hard Problems No Longer Proves Expertise
25 Fields Medalists and 31% of the consulting market are pointing to the same breakdown.
BusinessThe Declaration Uses the Term “Proxy Metric”
On September 11, 2026, 25 Fields Medal laureates put their names to a joint declaration. Its argument: the way AI companies treat the solving of hard math problems as a performance benchmark is badly out of step with what mathematics as a discipline actually exists to do. The signatories span generations — from Pierre Deligne, the 1978 laureate, to Wi Deng, the 2026 laureate — and the list also includes Korean mathematician Professor June Huh.
The heart of the declaration isn’t criticism — it’s a single-line definition. Solving problems, it says, is merely a tool and a proxy metric1 for gauging whether the real goal — conceptual understanding — has actually been reached. Mistake the tool for the goal, and the tool ends up undermining the goal itself.
Two days later, on September 13, the number of signatories had grown to 4,432. The signing page only adds a name after it’s been verified through an ORCID login or a confirmed institutional email address. Even a document announcing that proxy metrics have collapsed, in other words, still had to bolt on its own mechanism for confirming that whoever signed it was a real researcher.
I couldn’t read this declaration as an event confined to the world of mathematics. I’d already been watching the same breakdown play out in a different market.
Hard Problems Used to Be Mathematics’ Measuring Stick
Famous unsolved problems in mathematics have always done two jobs at once. One job faces inward. Solving a hard problem was a signal that new methods and concepts had been forged along the way, and mathematicians turned those ideas into shared tools through presentation, debate, and formalization.
The other job faces outward. I can’t explain what a Galois representation is, but I know Fermat’s Last Theorem got solved. Hard problems were a gauge that let even people who don’t understand math recognize genuine ability when they saw it.
What the declaration objects to isn’t the fact that AI can solve hard problems. It’s that when solutions get rushed into the open, there’s no proper writeup, no citation of prior work, no process of extracting and refining the underlying ideas. The signatories wrote that this is exactly where attribution problems and plagiarism creep in. When a result appears but nobody is left who can understand it and carry it forward, the gauge survives while whatever it was measuring disappears.
Put simply, what mathematicians are calling dangerous is this: if the habit of skipping every intermediate step and just grabbing the final answer keeps repeating, the essence and underlying logic of mathematics itself start to blur.
Most of the mechanisms used to sort out real experts were exactly this kind of proxy metric — and AI has driven the cost of faking that metric down to nearly zero.
Take a math problem whose answer is -1, 0, or 1: the odds of guessing correctly are 33.3%. A student who genuinely has the skill will work through the formula and show their reasoning, while someone else will just guess the short answer. Once AI-style shortcut solutions become the norm, the question becomes: who’s still going to bother studying math?
Mathematics is simply where this is most visible. In markets where things carry a price tag, the cracks had already been showing for a long time.
The advisory market was already running on self-reporting
There’s a market called expert networks2. It’s where investment firms and consulting firms buy hour-long calls with practitioners from specific industries. GuidePoint says it has a pool of over 2 million experts, and GLG claims over 1 million.
On December 4, 2025, a research firm that studies this industry published survey results from 1,368 experts who had participated in advisory calls. About 31 percent of respondents said they had taken part in a call on a topic they weren’t sufficiently qualified to speak on. That’s a little over 420 out of 1,368 people. The report identified three causes: expertise being self-declared, the 24-to-48-hour turnaround required to fill requests, and matching based on keywords.
Here’s something worth pausing on. AI doesn’t even appear on the list of causes behind under-qualified calls. Verification wasn’t broken by AI — it had been resting on self-reporting from the start. The same survey found that while clients pay $1,200 per call, the expert’s cut comes to less than 20 percent of that billed amount — and from the buyer’s side, that gap was effectively the price paid for verification. Since this is a self-reported survey of industry participants, the exact figures should be read broadly, but there’s no reason to doubt the direction of the finding.
The compliance procedures expert networks run before and after calls are rigorous. But they exist to prevent the exchange of undisclosed material information. They aren’t designed to measure whether this person actually did that job at that time, or whether what they’re saying now reflects their own judgment.
Both the buying side and the selling side are turning into the model
Start with demand, and the shift is already fast. GuidePoint announced on March 26, 2026 that it had passed 100,000 expert interview transcripts, and on May 5 it plugged that library into Claude via MCP3. By August 26, it had connected over 120,000 transcripts to Gemini Enterprise. The library grows by 5,000 transcripts a month, so the pile only gets bigger over time. AlphaSense and Tegus advertise AI-conducted expert calls at 70 percent cheaper than the traditional format, and GLG offers AI-led calls too.
An answer obtained by asking a person becomes a transcript; the transcript becomes input to a model; the next question goes to the model instead of a person. A single hour of someone’s time, once purchased, turns into an asset that keeps getting resold.
It isn’t hard to guess which way prices will move. There’s no reason to pay for a call to answer a question the existing transcripts can already answer, so the price of common questions starts falling first. Conversely, a premium attaches to questions the library doesn’t cover — judgment calls nobody has put on record yet. To be clear, this isn’t something the companies have stated; it’s my own projection based on how the structure works.
The supply side is moving the same direction. A large share of advisory work happens in writing, and writing is exactly the format a model handles best. Switching to video doesn’t change much either. So the buyer can’t easily tell whether they bought a person or a model, and the seller has no good way to prove they’re human. Both sides lose out — and the transaction happens anyway.
So What Can Actually Be Proven
At this point, only one answer remains. Not the correct answer, but what someone actually went through. But verifying experience is oddly slippery. Split it into three layers and you can see exactly where the gap sits.
| Layer | What can be proven | Who confirms it |
|---|---|---|
| Affiliation and tenure | Where you were, for how long | Company, school, institution |
| Record of judgment | What you decided when, and what turned out wrong | Right now, only the person themselves |
| Attribution of outcomes | Whether the result was actually due to your judgment | No verification system exists |
The first row is verifiable but says nothing about skill. The third row speaks to skill but can’t be verified. So the real test comes down to the second row. For researchers, ORCID IDs and citation records fill this gap — but practitioners have nothing equivalent.
The same logic operates well beyond advisory calls, too. Resumes, proposals, portfolios, interview answers — all of these are proxy indicators, and being well-crafted no longer screens out anything. The era when a well-written document served as proof of ability was short-lived. Markets like Korea’s, where evaluating people through written materials and one-off presentations is deeply entrenched, are hit by this shift more directly than most.
If you’re the one buying advice, changing the question is the cheapest response. Models are faster, and often more accurate, at market forecasts or structural explanations. But ask instead about a decision the person once opposed, a case that failed and the standards they changed afterward, or where in this very answer they’re most likely to be wrong — and the character of the response shifts entirely. A choice made under one’s own name at a specific point in time, and the cost that came with it, simply doesn’t turn up in a search. If you’re the one selling advice, keeping a timestamped record of exactly those choices and their costs is what becomes your asset.
Some Read This Declaration as a Red Flag Act
There’s an opposing reading too. In 1865, Britain passed a law requiring a person carrying a red flag to walk ahead of automobiles — it’s often cited as the classic case of horse-and-carriage interests slowing down a new technology. On that reading, the Fields medalists’ declaration is ultimately just a protest from people about to lose their standing. Beyond that, there’s also the argument that solving math problems will, going forward, remain an intellectual pastime — like chess.
Look at what the declaration actually demands, though, and the gap between the two readings becomes clear. The Red Flag Act regulated the act of using the technology itself. The declaration, by contrast, opens by acknowledging that AI has real potential to expand genuine mathematical understanding, and then takes issue with the practice of rushing solutions into public view while skipping organization, citation, and attribution. What it’s targeting isn’t the technology — it’s the publication process.
Timing matters here too. In 2002, a math teacher named Paul Lockhart wrote an essay called “A Mathematician’s Lament.” He argued that school mathematics had stripped away the process of discovery and kept only the results. Mathematics, in his view, isn’t the fact that a triangle’s area equals half of base times height — it’s the single line of insight that gets you there — yet schools simply make students memorize the formula and move on. His conclusion was that there’s no mathematics in math class. The same complaint existed 24 years ago, and back then the target wasn’t AI — it was the curriculum that the mathematics community itself had built. That’s why it’s hard to read this declaration as a defense hastily thrown together the moment a new technology showed up.
But bringing Lockhart into the picture doesn’t let the mathematics community off the hook, either. The procedural mathematics he indicted — stamping out true and false according to a fixed format — happens to be exactly what current models do best. And it was the education system and academia that turned that format into a standard and graded it through exams. What got automated this time isn’t mathematical thinking — it’s the format we chose to grade in place of mathematical thinking. **The advisory market works the same way. It’s been buying and selling written reports built to a fixed template, so of course the template is the first thing to get automated. That’s actually why, when I set up my own consulting firm, the first thing I did was decide not to sell reports. AI can write a report in no time anyway. Better to spend that time building a working PoC/MVP or running a workshop.
The intellectual-pastime argument, if anything, actually reinforces this piece’s thesis. Chess kept a human league even after computers beat the best players. But what that league sells is spectacle, not a certificate of skill. People pay money to watch matches between players weaker than the computer. Saying that problem-solving becomes a pastime means that, at that very point, its function as proof of skill falls away — and that is precisely the process by which a proxy metric disappears. The only difference is that in the advisory market, there’s no one willing to pay an admission fee to watch.
Oswarld’s Lens
I serve as an expert advisor for outfits like GLG and Silver Lake. Recently I’ve been hearing a similar complaint from the people who run those programs. Ever since LLMs became widespread, there’s almost no way left to tell experts apart. Written consultations get answered by AI, and even video sessions increasingly feature frontier models that can respond in real time — that’s just the reality now.
So the only way left to verify an expert is experience — but I keep getting stuck on how vague it is to actually certify that experience. You can confirm which company I worked at and for how long, but what judgment calls I made there, and whether those calls were right, has no evidence behind it besides my own word. Other than filling in the second row of the table above for myself, I still haven’t found a tool I can actually use right now.
One thing seems clear, though. What Fields Medal winners are trying to protect and what advisory-program managers are worried about are the same kind of problem. Both are losing the ability to recognize the person who produced a piece of work — not its quality, but its author. The difference is that mathematics already has verification machinery in place — papers, citations, institutional affiliation — while the market for working professionals has nothing but self-reporting. That’s why I think this market will collapse first, and more quietly.
Closing
Proxy metrics were always stand-ins we used instead of measuring real ability. Solving problems, writing a good answer on paper, explaining something fluently — all of it worked that way. Now that models have taken over those stand-ins, what’s left is what someone actually went through and the choices they made in the moment. The problem is that we don’t yet have a system to prove any of that.
If you’re in a position where you have to judge someone’s expertise, for now it’s safer to ask about the path an answer took to get made, not just the quality of the answer itself. I’ll keep watching to see whether the number of people who freeze up when asked that question is actually growing.
💬 The last time you had to judge someone’s expertise, what did you look at to decide? If you’ve got a criterion that failed you, I’d love to hear about that too.
📨 If you have a colleague who works with advisors or outside experts, send them this piece. Even just swapping in this list of questions is enough to put it to use right away.
Keep the perspective, not the noise.
We choose one consequential shift and trace what sits beneath it, every other day.
Confirm once to finish subscribing.
Already a subscriber? Sign in to join the conversation
References & Further Reading
Primary sources
- Math and AI, “A Severe Misalignment of AI in Mathematics”, September 11, 2026. ··· The full text of the declaration, the list of 25 initial signatories, and the signing process are all on one page
- Terence Tao, “A Severe Misalignment of AI in Mathematics”, What’s new, September 11, 2026. ··· Comes with a short account from one of the signatories himself on how the declaration came together
- Woozle Research, “The State of the Expert Economy 2025”, December 4, 2025. ··· Breaks down the denominator behind the 31 percent unqualified-response figure and analyzes its causes
- Guidepoint, “Guidepoint Launches MCP on Claude”, May 5, 2026. ··· The company explains directly how expert-call transcripts become model input
Background
- GLG, “Expert Calls”. ··· Shows what compliance procedures actually check — and what they don’t
- Paul Lockhart, “A Mathematician’s Lament”, 2002. ··· The full 25-page original. Just the first two pages, which open with a nightmare about music education, reveal the roots of this declaration
- Guidepoint, “Guidepoint Accelerates Enterprise AI Transformation with Gemini Enterprise”, August 26, 2026. ··· Lets you compare how transcript volume grew over just five months
📝 Glossary
Footnotes
-
Proxy: An observable indicator used as a stand-in for something hard to measure directly. Solving a hard problem stands in for understanding; a certificate stands in for actual skill. ↩
-
Expert network: A service that brokers paid calls between investment firms or consultancies and practitioners in a given industry. GLG, Guidepoint, and AlphaSights are leading examples. ↩
-
MCP (Model Context Protocol): A connection standard for attaching external data or tools to an AI model. It’s used when a company wants a model to directly query its own proprietary materials. ↩

Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?