🔍 Read the full analysis: Why AI Output Is Cheap But Confidence Still Costs on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI reported producing 722 mathematical manuscripts this week, while an earlier result from the same programme drew careful scrutiny from five leading mathematicians. The contrast reflects a wider pattern in software and legal workflows: AI can generate work quickly, but human review, accountability and the ability to judge whether a result is fit for purpose remain limited.
OpenAI published 722 mathematical manuscripts this week, according to source material describing the programme, while a counterexample to an older Erdős conjecture from the same programme received careful verification by five leading mathematicians. The contrast points to a growing operational problem across AI-assisted work: generating results is becoming faster and cheaper than checking whether they are correct, relevant and safe to use.
The source says OpenAI’s model was given about 4,000 problems and produced manuscripts grouped into 372 families, with an average result taking about three hours of compute. Some manuscripts were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification “could have issues.” Formal checking can validate a proof against its stated claim, but it does not by itself determine whether the claim is important or answers the right question.
In software, the source cites several studies and industry analyses with different methods. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analyzing 8.1 million pull requests across 4,800 organizations, found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests had no human review before being merged or closed.
The figures should not be treated as directly comparable: they come from separate analyses, and the source notes that several providers sell code-review tools. Still, the reported pattern is consistent across the examples: more work reaches reviewers, while review capacity and confidence do not rise at the same pace.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets AI’s Limits
For organizations adopting AI, the practical constraint may be how much output qualified people can verify, rather than how much a model can produce. If review becomes a bottleneck, teams may delay useful work, accept changes without adequate scrutiny or rely too heavily on the system that generated the material to decide what deserves attention.
The source also identifies a workforce risk. Junior staff traditionally build judgment by drafting code, contracts or research and receiving feedback. If AI takes over much of that first-hand work, employers may weaken the training path to senior expertise even as they need more people capable of evaluating AI output. The source presents this as a risk, not a demonstrated outcome across all workplaces.
That could increase the value of experienced reviewers, auditors, specialist lawyers and safety assessors. Their role is not simply to check whether output looks polished; it is to judge whether it meets the real need and to take responsibility for decisions. The idea of a “referee premium” is the source author’s economic interpretation, rather than a measured wage trend.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Review Gap
The examples span mathematics, software and contracting. In mathematics, formal proof tools can test whether a derivation follows from defined premises, while human researchers assess the choice and significance of the proposition. In software, automated tests check specified behavior, but their coverage depends on what developers thought to test. In contracting, a model can draft or evaluate clauses, but a professional still has to check jurisdiction, approvals and the client’s obligations.
The source describes OpenAI’s partnership with contract-software company Ironclad and says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, an improvement over a previous model. That result is attributed to the source material; it does not provide the underlying evaluation, comparison score or full criteria. The figure therefore signals remaining gaps in that particular assessment, not a general failure rate for contract work.
The common distinction is between checking an answer and judging its purpose. A formal system can verify a proof of a stated theorem, and tests can check behavior they encode. Neither automatically confirms that the theorem, requirements or tests reflect the problem people actually needed to solve.
formal proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Evidence Cannot Yet Show
The cited figures do not establish one universal rate of AI error or review failure. The studies cover different organizations, time periods and definitions of review, and some of the industry sources sell review products. The source does not provide links or full methods for every statistic, so the numbers should be read as reported findings rather than a single, independently comparable dataset.
It is also unclear whether the reported review delays and no-review rates will persist as tools and workplace processes change. The source offers a warning that AI may erode the experience junior workers need to become expert reviewers, but it does not quantify that effect or establish how widely it is occurring. Nor does it provide the full methodology behind the Ironclad evaluation or the mathematical programme’s selection of manuscripts for publication.
As an affiliate, we earn on qualifying purchases.
Building the Next Review Pipeline
The immediate test for organizations is whether they can expand review capacity alongside AI use: setting clear standards for human approval, tracking which work receives scrutiny and preserving responsibility for consequential decisions. The cited studies do not establish which approach works best, but they show why output volume alone is an incomplete measure of productivity.
Employers will also need to decide how junior staff gain experience if AI handles more routine drafting and coding. Training that includes supervised creation, revision and review could help preserve the route to expert judgment; the source frames that as a response to a risk, not an established industry policy. For mathematics and other research fields, the outstanding question is how formal verification and specialist peer review can scale without treating automated checks as substitutes for human adjudication.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the central development in this report?
The report describes a widening gap between AI’s ability to generate work and people’s capacity to verify and judge it. Its lead example is OpenAI’s reported publication of 722 mathematical manuscripts this week.
Does formal verification prove that AI-generated work is useful?
No. Formal verification can establish that a proof follows from its stated premises. It does not by itself show that the claim is useful, relevant or the right one to address. In software, tests similarly check only the behavior they cover.
What did the cited software analyses find?
Faros AI reported higher pull-request merging alongside longer review time during high-AI-adoption periods. LinearB reported longer waits before review and lower acceptance rates for AI-generated changes in its analysis. These are separate datasets, not one combined result.
Are the figures proof that AI makes software less reliable?
No. The figures describe review timing, acceptance and review coverage in particular analyses; they do not establish a general reliability rate for AI-written software. The source also cautions that some providers cited sell code-review tools.
Why might AI use affect junior workers?
The report argues that people often develop judgment by doing the drafting and coding that AI can now assist with or take over. If entry-level workers get fewer opportunities to practice and receive feedback, organizations could have a harder time developing experienced reviewers. The scale of that effect remains unknown.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
