AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Makes It Cheap To Create—and Costly To Verify on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says its model produced 722 mathematical manuscripts this week after being posed about 4,000 problems, while a related result drew careful scrutiny from five leading mathematicians. Studies of software development and a new legal-workflow evaluation point to the same tension: AI can increase output, but people and institutions still have to decide what is reliable and usable.

OpenAI published 722 mathematical manuscripts this week produced by a model posed about 4,000 problems, according to material supplied by the company. The release highlights a growing challenge across research, software and professional services: AI can generate work faster than people can check it, leaving review capacity as a constraint on what organizations can safely use.

The manuscripts cover 372 problem families, and the source says the average result took about three hours of compute to produce. Some results were formally checked using Lean, a proof assistant. OpenAI cautioned that some results without formal verification “could have issues.” The supplied account contrasts that volume with the response to an earlier result from the same programme: a counterexample to an old Erdős conjecture that received careful verification by five leading mathematicians. These examples have different scopes, and the comparison does not establish how much total human review the new manuscripts received.

Software studies cited in the source describe a similar strain. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analyzing 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A separate peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before they were merged or closed. Faros also reported a 31.3% rise in merges with zero review during high-adoption periods.

The figures do not all measure the same thing, and several cited software-data providers sell code-review tools. The source says to read their results with care. In professional services, an OpenAI partnership with contract-software company Ironclad tested GPT-6 Astra on 11 contracting tasks. Astra met 55% of evaluation criteria on average, according to the supplied account; that leaves substantial criteria unmet, though the source does not specify how the tasks or criteria were weighted.

At a glance
analysisWhen: This week; software and legal-workflow…
The developmentOpenAI’s publication of 722 AI-produced mathematical manuscripts, alongside reported increases in software review burdens and gaps in legal-workflow evaluations, highlights a widening gap between AI output and human verification.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

More generated work does not automatically mean more useful work. Each proof, code change or contract draft still needs a decision about whether it answers the right question, fits its real-world setting and can be relied on. When review cannot keep up, organizations face choices that can carry real costs: leave work waiting, accept it with limited scrutiny or reject it based on broad suspicion rather than its merits.

The effects could reach beyond immediate productivity. If senior specialists become the people every AI-assisted workflow must wait for, their time may become more valuable and harder to scale. That is an interpretation of the reported pattern, not a measured estimate of wages or staffing demand. The source calls this a “referee premium”: increased value for people who can assess work and take responsibility for approving it.

There is also a training concern. Junior professionals often develop judgment by doing the work that more experienced colleagues later review: writing code, proving results or drafting contracts. If AI takes over much of that practice, organizations may have to find deliberate ways for newer workers to build those skills. The cited figures do not show whether that is already happening or quantify its long-term effect.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Bottleneck

In mathematics, formal proof checking can establish that a proof follows the stated rules and proves the stated theorem. It does not, on its own, show that the theorem addresses an important question or that the result is meaningful. The supplied source summarizes the challenge with the phrase “verification abundance, adjudication scarcity.” It attributes that framing to a recent paper, but does not provide the paper’s title, authors or publication details.

Software tests can check whether code meets specified conditions, but they cannot establish that the conditions capture what users or systems actually need. Similarly, a contract review has to take account of its intended use and applicable rules. Across these fields, technical checks can help, but they do not transfer professional or legal responsibility from the people and institutions using the work.

The evidence cited here comes from different kinds of sources: a company account of its mathematics programme, analyses from software-industry providers, a peer-reviewed study, and a product evaluation associated with an OpenAI partnership. Their findings should not be treated as one unified measurement of AI’s effect. Together, however, they raise a shared operational question: who has the time and expertise to assess the additional output?

““verification abundance, adjudication scarcity””

— A recent paper, as characterized in the supplied source

Amazon

formal verification software for AI research

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The supplied material does not state how many of the 722 manuscripts were independently reviewed, how many were formally verified, or how the five mathematicians’ work on the earlier counterexample compares in effort with review of the larger batch. It also does not provide enough methodological detail to compare the software studies directly or determine whether their findings apply across industries and team sizes.

The 55% figure for GPT-6 Astra is an average across 11 tasks, but the source does not name the evaluation criteria, explain their relative importance or say whether a human reviewer corrected the missed criteria. It is also unclear how review costs change as AI systems improve, whether review tools can reduce the burden, and how much oversight is required for different levels of risk. The evidence points to a verification bottleneck, but does not establish its eventual scale.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review Alongside Output

Organizations adopting AI can monitor not only how much work gets produced, but also how long it waits for review, what share receives human scrutiny and how often reviewers find consequential errors. For AI-generated research, useful next details would include independent assessments, the number of results checked with formal tools and the standards used to select manuscripts for publication.

For professional workflows, the next question is whether evaluations expand beyond task completion rates to test errors, review time and the consequences of missed requirements. The supplied material gives no timetable for further results from OpenAI, Ironclad or the cited studies. Until those details emerge, the central issue remains practical: how quickly can qualified reviewers assess the work AI systems produce?

Amazon

AI manuscript verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI publish this week?

According to the supplied source, OpenAI published 722 mathematical manuscripts across 372 problem families after posing about 4,000 problems to its model.

Were all the mathematical results formally verified?

No. The source says some results were checked in Lean, while OpenAI cautioned that some unformalized results “could have issues.” It does not give a count for either group.

What do the software review figures show?

The cited analyses report higher pull-request volume alongside longer waits or limited review in some settings. Their methods and measures differ, and the source notes that several providers sell code-review tools, so the figures need to be interpreted with care.

The supplied account says GPT-6 Astra met 55% of evaluation criteria on average across 11 contracting tasks. It does not identify the criteria or explain their weighting, so the figure alone cannot show how usable the outputs were.

Why can’t AI simply verify its own work?

Automated checks can test a proof against formal rules or code against specified tests. They may not establish that the original question, requirements or tests were appropriate. People also remain responsible for many professional and legal decisions.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Speech-to-Text Devices Support Accessible Learning

An overview of how speech-to-text devices enhance accessible learning and why they are essential for inclusive classrooms and better student engagement.

What AI Means for University Admissions and Advising

An exploration of how AI transforms university admissions and advising, offering personalized, equitable insights that could reshape your educational journey—discover the full impact.

AI-Powered Note-Taking Apps: A Back to school Guide

Discover how AI-powered note-taking apps boost efficiency, accuracy, and integration. Learn the latest features, trends, and practical tips to upgrade your notes.

AI in Coding Bootcamps for Teens

Greatly transforming teen coding bootcamps, AI introduces innovative tools and skills that will shape your future in technology—discover how inside.