Why AI Output Is Cheap But Confidence Still Costs
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Output Is Cheap But Confidence Still Costs on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI reported producing 722 mathematical manuscripts this week, while an earlier result from the same programme drew careful scrutiny from five leading mathematicians. The contrast reflects a wider pattern in software and legal workflows: AI can generate work quickly, but human review, accountability and the ability to judge whether a result is fit for purpose remain limited.

OpenAI published 722 mathematical manuscripts this week, according to source material describing the programme, while a counterexample to an older Erdős conjecture from the same programme received careful verification by five leading mathematicians. The contrast points to a growing operational problem across AI-assisted work: generating results is becoming faster and cheaper than checking whether they are correct, relevant and safe to use.

The source says OpenAI’s model was given about 4,000 problems and produced manuscripts grouped into 372 families, with an average result taking about three hours of compute. Some manuscripts were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification “could have issues.” Formal checking can validate a proof against its stated claim, but it does not by itself determine whether the claim is important or answers the right question.

In software, the source cites several studies and industry analyses with different methods. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analyzing 8.1 million pull requests across 4,800 organizations, found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests had no human review before being merged or closed.

The figures should not be treated as directly comparable: they come from separate analyses, and the source notes that several providers sell code-review tools. Still, the reported pattern is consistent across the examples: more work reaches reviewers, while review capacity and confidence do not rise at the same pace.

At a glance
reportWhen: Reported this week; cited software rese…
The developmentA source report compares OpenAI’s latest mathematical output with verification demands and cites software and contract-work data to argue that AI generation is outpacing human review.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets AI’s Limits

For organizations adopting AI, the practical constraint may be how much output qualified people can verify, rather than how much a model can produce. If review becomes a bottleneck, teams may delay useful work, accept changes without adequate scrutiny or rely too heavily on the system that generated the material to decide what deserves attention.

The source also identifies a workforce risk. Junior staff traditionally build judgment by drafting code, contracts or research and receiving feedback. If AI takes over much of that first-hand work, employers may weaken the training path to senior expertise even as they need more people capable of evaluating AI output. The source presents this as a risk, not a demonstrated outcome across all workplaces.

That could increase the value of experienced reviewers, auditors, specialist lawyers and safety assessors. Their role is not simply to check whether output looks polished; it is to judge whether it meets the real need and to take responsibility for decisions. The idea of a “referee premium” is the source author’s economic interpretation, rather than a measured wage trend.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Review Gap

The examples span mathematics, software and contracting. In mathematics, formal proof tools can test whether a derivation follows from defined premises, while human researchers assess the choice and significance of the proposition. In software, automated tests check specified behavior, but their coverage depends on what developers thought to test. In contracting, a model can draft or evaluate clauses, but a professional still has to check jurisdiction, approvals and the client’s obligations.

The source describes OpenAI’s partnership with contract-software company Ironclad and says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, an improvement over a previous model. That result is attributed to the source material; it does not provide the underlying evaluation, comparison score or full criteria. The figure therefore signals remaining gaps in that particular assessment, not a general failure rate for contract work.

The common distinction is between checking an answer and judging its purpose. A formal system can verify a proof of a stated theorem, and tests can check behavior they encode. Neither automatically confirms that the theorem, requirements or tests reflect the problem people actually needed to solve.

Amazon

formal proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evidence Cannot Yet Show

The cited figures do not establish one universal rate of AI error or review failure. The studies cover different organizations, time periods and definitions of review, and some of the industry sources sell review products. The source does not provide links or full methods for every statistic, so the numbers should be read as reported findings rather than a single, independently comparable dataset.

It is also unclear whether the reported review delays and no-review rates will persist as tools and workplace processes change. The source offers a warning that AI may erode the experience junior workers need to become expert reviewers, but it does not quantify that effect or establish how widely it is occurring. Nor does it provide the full methodology behind the Ironclad evaluation or the mathematical programme’s selection of manuscripts for publication.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Building the Next Review Pipeline

The immediate test for organizations is whether they can expand review capacity alongside AI use: setting clear standards for human approval, tracking which work receives scrutiny and preserving responsibility for consequential decisions. The cited studies do not establish which approach works best, but they show why output volume alone is an incomplete measure of productivity.

Employers will also need to decide how junior staff gain experience if AI handles more routine drafting and coding. Training that includes supervised creation, revision and review could help preserve the route to expert judgment; the source frames that as a response to a risk, not an established industry policy. For mathematics and other research fields, the outstanding question is how formal verification and specialist peer review can scale without treating automated checks as substitutes for human adjudication.

Amazon

AI mathematical proof assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the central development in this report?

The report describes a widening gap between AI’s ability to generate work and people’s capacity to verify and judge it. Its lead example is OpenAI’s reported publication of 722 mathematical manuscripts this week.

Does formal verification prove that AI-generated work is useful?

No. Formal verification can establish that a proof follows from its stated premises. It does not by itself show that the claim is useful, relevant or the right one to address. In software, tests similarly check only the behavior they cover.

What did the cited software analyses find?

Faros AI reported higher pull-request merging alongside longer review time during high-AI-adoption periods. LinearB reported longer waits before review and lower acceptance rates for AI-generated changes in its analysis. These are separate datasets, not one combined result.

Are the figures proof that AI makes software less reliable?

No. The figures describe review timing, acceptance and review coverage in particular analyses; they do not establish a general reliability rate for AI-written software. The source also cautions that some providers cited sell code-review tools.

Why might AI use affect junior workers?

The report argues that people often develop judgment by doing the drafting and coding that AI can now assist with or take over. If entry-level workers get fewer opportunities to practice and receive feedback, organizations could have a harder time developing experienced reviewers. The scale of that effect remains unknown.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of Webcams: 10 Best AI 4K Devices In 2026

Discover the 10 best AI-enhanced 4K webcams of 2026, including features, performance, and value for streamers, professionals, and content creators.

MartyPC Is A Cross-platform Emulator Of Early PCs Written In Rust

MartyPC, a new emulator for early PCs built in Rust, now supports multiple platforms, promising improved performance and accessibility for vintage computing enthusiasts.

Unlocking The Potential Of ChatGPT Ads In Europe’s AI Industry

OpenAI announces expansion of ChatGPT Ads into Europe, but specific countries, launch dates, and details remain unconfirmed, raising questions about impact and scope.

Future-Proof Your Business With Insights Into AI Futures

OpenAI announced AI Futures, a new research initiative to examine how advanced AI could influence governance, power distribution, and individual rights.