By Jim Shimabukuro (assisted by Claude)
Editor
Mathematicians, programmers, and journal editors are being buried in AI output. The obvious fix is to put more machines between the flood and the people. What are those machine layers, and will they get us to the point where HITLs, or Humans in the Loop, are rarely needed?
On October 6, OpenAI posted 722 mathematical manuscripts to GitHub. They were written by an internal model the company has not released, grouped into 372 research “families,” and drawn from about 4,000 problems the model attempted. A typical accepted result used about three hours of the equivalent of ChatGPT Pro thinking time (Khollam, 2026; OpenAI, 2026). Some of the claims touch famous open questions, including a special case of the Hodge conjecture and the four-dimensional Kakeya conjecture (Hart, 2026).
Three days later, The Verge’s Robert Hart reported what happened when more than three dozen mathematicians tried to take it all in. For one group, “simply working through the roughly 40-page table of contents and abstracts took them the better part of an hour.” Álvaro Lozano-Robledo of the University of Connecticut said, “Just going over the entire list of abstracts is overwhelming.” Brendan Hassett of Brown University said, “The write-up of the problem I know best made little sense after a quick read.” OpenAI said that about 42 percent of the top-line results had been formalized in Lean, a programming language whose software checks every logical step of a proof. The rest arrived as prose claims. Many of the mathematicians estimated that making sense of the release “could take the community years,” and OpenAI soon withdrew three papers and revised more than a dozen others (Hart, 2026).
The episode is the sharpest example yet of a problem that now runs through software engineering, scientific publishing, and AI safety. Machines produce work faster than people can check it. The standard safeguard in AI policy, the “human in the loop,” assumes the human can keep up. When the loop carries 722 manuscripts in a single drop, or thousands of lines of code an hour, that assumption fails.
A natural response is engineering. If one machine produces too much, build a second machine to screen its output and pass along only what a person needs to see. If the screened pile is still too large, add another layer, and keep adding layers until the human is called in only rarely. Societies have handled floods of mail, manufactured parts, and medical scans in roughly this way for generations. So is the review crisis less complicated than it looks?
The evidence supports the instinct. Layered machine review is already the main response to the problem, and some of the best-funded AI labs and most respected mathematicians are building it. The evidence also shows three places where the plan gets harder. Machine checkers are strongest where correctness can be tested mechanically and weakest on questions of meaning and importance. Stacked AI reviewers tend to share the same blind spots. And a human who is called in rarely gets worse at the job. Each problem has partial remedies. None is solved by adding more layers.
The bottleneck
Software reached the bottleneck first. Faros AI’s analysis of data from 10,000 developers, as reported by the engineering firm Aviator, found that teams with high AI adoption merged 98 percent more pull requests (the bundles of code changes that programmers submit for review), while the time spent reviewing them rose 91 percent and the average size of each change rose 154 percent (Lukić, 2026). Writing code had become much faster. Reading it had not. When Anthropic launched an automated code-review service in March, the company stated the problem plainly: “Code review has become a bottleneck, and we hear the same from customers every week,” and “many PRs get skims rather than deep reads” (Sawers, 2026).
Scholarly publishing reports the same squeeze. At the management journal Organization Science, an internal study of 6,957 submissions found that submissions rose 42 percent after ChatGPT’s launch. The journal nearly doubled its deputy editors, from six to eleven, and doubled its senior editors, from about 30 to 60 (Drake, 2026). Matt Lord, executive editor of Boston Review, told The Boston Globe that unsolicited drafts had tripled compared with 2023 and that “there’s no world in which we can continue reading every single submission we get very closely” (Ryan, 2026).
Mathematics shows the problem in its purest form. In April, Terence Tao of UCLA, the Fields medalist who has done more than anyone to test AI tools in public, wrote that his field was entering “an era of proof abundance.” He separated the work into three stages: generating a proof, verifying it, and digesting it so that other mathematicians can understand and build on it. AI, he argued, was speeding up the first two stages far faster than the third. He noted that the Erdős problems website, which had rarely had more than one or two claimed solutions waiting to be assessed, now had nearly twenty (Tao, 2026a). Thomas Bloom, who runs that site, a catalog of more than 1,200 problems left by the Hungarian mathematician Paul Erdős, told Quanta Magazine in August, “We’re seeing a lot more of these 100- to 200-page papers that people are posting.” He added, “But no human has read it, and no human is going to read it” (Kakaes, 2026).
The layers of fixes
The idea of placing machines between the flood and the person is being pursued, under different names, in at least five places. The first is mathematics, where the screening layer rests on unusually solid ground. Lean and similar “proof assistants” check a formal proof against a small, trusted core of logical rules and either accept it or reject it. Liam Fowl, an AI researcher at Princeton, told IEEE Spectrum that “formal verification of a proof is like a rubber stamp” (Skuse, 2026). Several startups are racing to automate the translation of ordinary mathematical writing into Lean, a step called autoformalization. Math, Inc.’s agent, Gauss, finished formalizing the strong prime number theorem, a project Tao and Alex Kontorovich had been working on, and the company’s chief executive, Jesse Han, says the aim is “to free mathematicians to do what they do best” (Skuse, 2026). Contributors to the Erdős site use Harmonic’s Aristotle system to certify that a proof holds together logically (Kakaes, 2026). In May, a 24-author team reported that an AI agent pairing a language model with Lean “autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars” (Tsoukalas et al., 2026).
Tao has gone a step further and built a triage layer himself. In August he announced Palomar, which he described as “the analogue of a preprint server for Lean proofs.” Palomar runs two checks on each submitted project. The first is mechanical: software confirms that the code “typechecks and proves exactly the results claimed.” The second asks whether the plain-language description matches the formal statement, and that check is “performed by a large language model.” Tao was careful about what the screen does and does not do. Palomar “is not a peer-reviewed journal,” he wrote. Its checks “fall well short of what a proper human peer review” provides, and “human review is still strongly recommended.” In a reply to a reader, he explained the design choice: Palomar “does not perform human review of repositories, which will not scale” (Tao, 2026b).
Outsiders began building a triage layer for OpenAI’s release within days. Emergent Mind, a research-discovery site, published an “OpenAI Math Explorer” that maps the 722 manuscripts, lets readers search them by meaning, and labels each one by what its main result claims and whether it has been formalized. The site says its map and labels were “built with text embeddings and an LLM, and may contain errors,” and it warns, “None of these results has been peer reviewed” (Emergent Mind, 2026).
The second place is software. Anthropic’s code-review service sends several AI agents over each pull request. In the company’s internal testing, the share of pull requests receiving substantive review comments rose from 16 percent to 54 percent, engineers marked fewer than 1 percent of the findings as incorrect, and 84 percent of changes over 1,000 lines produced findings, an average of 7.5 issues each (Sawers, 2026). Dozens of competing products do similar work. The design goal matches the layered proposal exactly: let machines read everything, and send human attention where it is most needed.
The third is scientific peer review. For its 2026 conference, the Association for the Advancement of Artificial Intelligence (AAAI) generated an AI review for every one of the 22,977 papers that went to full review, in less than a day, and labeled each as AI-generated alongside the human reviews. In a survey afterward, authors and program committee members said the AI reviews were useful and preferred them to human reviews on technical accuracy and research suggestions (Biswas et al., 2026).
The fourth is AI safety research, where the layered design has a formal name: AI control. In a 2023 paper, researchers at Redwood Research set out the basic arrangement. A powerful model that cannot be fully trusted does the work. A weaker model that can be trusted watches it. A small supply of human attention, which the paper calls “limited access to high-quality trusted labor,” is spent on whatever the watcher rates as most suspicious (Greenblatt et al., 2023). Google DeepMind’s AI control roadmap, posted in July, adopts the same structure. It calls automated monitoring “one of the main pillars of AI control” and lists “escalating to a human” as a response in which a flagged action is “gated by human review.” It warns that “the speed and volume advantage may make human oversight infeasible” without such tools, and it states the principle directly: “build in multiple layers of defence” (Phuong et al., 2026). Anthropic’s April report on automated alignment research reaches a similar place. If models produce “millions of lines of incredibly complicated code that we can’t parse ourselves,” the report says, the main bottleneck “could become evaluation,” and “human oversight remains essential” (Anthropic, 2026).
The fifth is the corporate market. In June 2025 the research firm Gartner predicted that “guardian agents,” AI systems whose job is to review, monitor, and block the actions of other AI systems, would take 10 to 15 percent of the market for agentic AI by 2030. Gartner analyst Avivah Litan gave the reason in one line: “humans cannot keep up with the potential for errors and malicious activities” (Gartner, 2025). In Britain, the government-funded Advanced Research and Invention Agency (ARIA) has refocused its Safeguarded AI program on tools that let AI-written software carry a mathematical proof of its own correctness, checked by a trusted kernel “for all cases – not just test cases” (Advanced Research and Invention Agency [ARIA], 2025).
The fixes aren’t new
The historical intuition behind the layered approach holds up. Medicine offers the clearest recent test. Breast cancer screening in much of Europe has long relied on two radiologists reading each mammogram. Sweden’s MASAI trial, the first randomized controlled trial of AI in breast cancer screening, enrolled more than 105,000 women. In the half assigned to the AI arm, software sorted each mammogram by risk. Low-risk scans went to one radiologist. High-risk scans went to two. An interim analysis found a 44 percent reduction in radiologists’ screen-reading workload. The full results, published in The Lancet in January, showed that the AI-supported arm caught 9 percent more cancers at screening and had 12 percent fewer “interval” cancers, the ones that surface between scheduled screenings, with a false-positive rate of 1.5 percent against 1.4 percent in the control group (ecancer, 2026).
The detail that matters most is that a radiologist still read every scan. Lead author Kristina Lång of Lund University said that “introducing AI in healthcare must be done cautiously, using tested AI tools and with continuous monitoring in place.” First author Jessie Gommers of Radboud University Medical Centre said, “Our study does not support replacing healthcare professionals with AI” (ecancer, 2026). The trial cut human workload nearly in half by deciding where human attention went. It kept a human on every case.
Safety engineering supplies the warnings. The psychologist James Reason’s “Swiss cheese model,” now standard in aviation and hospital safety, describes each line of defense as a slice of cheese “having many holes.” Accidents happen when “the holes in many layers momentarily line up to permit a trajectory of accident opportunity” (Reason, 2000). Stacking slices works because the holes in different slices sit in different places. That condition is where AI-on-AI review runs into trouble.
The problem with fixes
The first problem is that AI layers may not be independent. In a study presented at the International Conference on Machine Learning in 2025, Shashwat Goel and eight co-authors measured how often different language models make the same mistakes. They found that AI judges score models similar to themselves more favorably, and that “model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures” (Goel et al., 2025). Two AI reviewers trained on similar data by similar methods are two slices of cheese with holes in the same places. A third such reviewer adds cost and little protection. The partial remedies are diversity, meaning checkers built differently from the system that produced the work, and, wherever possible, checkers that are not language models at all, such as Lean’s logical core or a program’s test suite. ARIA’s insistence on proofs that hold “for all cases” follows that logic.
The second problem is that mechanical checkers can only check what can be stated mechanically. Lean can confirm that a proof proves a formal statement. It cannot confirm that the formal statement says what the paper claims it says, which is why Tao’s Palomar hands that question to a language model and why Tao still recommends human review. Lean also cannot tell whether a result is new, whether credit has been assigned fairly, or whether the result matters. Those were the questions that troubled the mathematicians The Verge interviewed. Nalini Joshi of the University of Sydney was “wary that attribution in the papers may be lacking the complete story,” and others found results that overlapped existing work (Hart, 2026). On Bloom’s site, an early celebrated solution to Erdős Problem 333 turned out to duplicate a resolution “Erdős himself had provided” in a paper published in 1977 (Kakaes, 2026).
In September, 28 Fields medalists, including Tao, Peter Scholze, and Maryna Viazovska, signed a declaration warning that “the mass production at faster and faster pace of ‘true/false’ statements could destroy fertile ground.” They argued that new results need “the willing mathematicians who must take care of their development and integration” if “the crucial human transmission chain” is to survive (Fields Medalists, 2026). The number theorist Jared Duker Lichtman put it more simply to The Verge: “As a society, we should be rewarding these digestive efforts more” (Hart, 2026). A verification filter shrinks the pile of wrong results. It does nothing to shrink the pile of correct results that no one yet understands.
The third problem is that the rarer the human’s turn, the worse the human performs. This is one of the oldest findings in the study of automation. In 1983 the British psychologist Lisanne Bainbridge published “Ironies of Automation,” a short paper on industrial control rooms that is still widely cited. She observed that “it is impossible for even a highly motivated human being to maintain effective visual attention towards a source of information on which very little happens, for more than about half an hour.” She noted that “physical skills deteriorate when they are not used.” And she identified the central irony: “the more advanced a control system is, so the more crucial may be the contribution of the human operator” (Bainbridge, 1983). A layered system built so that the human is rarely needed produces the operator Bainbridge described, one who is called in only for the hardest and strangest cases and has the least recent practice.
Margaret Mitchell and Avijit Ghosh of Hugging Face and Samir Passi of Data & Society applied the same reasoning to AI agents in a paper posted in August. “Oversight degrades the overseer,” they write. “Overseers must process voluminous outputs generated faster than they can meaningfully review them,” which produces “approval fatigue” and, over time, the “skill atrophy that arises from extended use of automation.” Their remedies include “strategic friction,” mechanisms that “require the user to perform cognitive work before or alongside agent operation,” along with training and rotating people between tasks (Mitchell et al., 2026). Kevin Buzzard of Imperial College London, one of Lean’s leading champions, described the reviewer’s dilemma with the unformalized part of OpenAI’s release. He said he could either “read possibly-not-correct slop, or wait for others to do the same” (Hart, 2026).
How the layers of fixes work
The evidence supports the layered proposal with three amendments. First, put mechanical checkers at the bottom of the stack wherever a field allows it. In mathematics that means Lean. In software it means compilers, type checkers, and test suites. In ARIA’s program it means proofs checked by a trusted kernel. These layers do not share the failure patterns of language models, and anyone can rerun their verdicts. Buzzard’s complaint applied to the 58 percent of OpenAI’s top-line results that arrived without a Lean proof.
Second, use AI layers to rank and route human attention, and test them as screening tools. The MASAI design and the AI-control “audit budget” follow the same logic: the machine decides which cases receive more human scrutiny, and people still see a defined share of the work. Builders should use checkers that differ from one another, measure how often their errors coincide, and avoid counting two similar models as two independent layers (Goel et al., 2025).
Third, design for the human who remains. That means rotation, training, and regular practice on real cases, the kind of “strategic friction” Mitchell and her colleagues recommend, so that the person at the top of the stack keeps the skill to overrule it. It also means paying for digestion. Lichtman’s call to reward the people who read, simplify, and explain machine-generated results is a question for funding agencies, universities, and journals. No screening layer answers it.
So is the problem less complex than it looks? The core idea of the layered approach is right, and it is already under way in mathematics, software, publishing, and safety research. The hard part is the end goal of a human who is “rarely needed.” The MASAI trial cut radiologists’ screen-reading workload by 44 percent, caught more cancers, and still put a radiologist on every scan. Tao’s Palomar screens every submission by machine and still recommends human review of anything that matters. Both treat human attention as the scarcest resource in the system and spend it carefully. Neither tries to spend it down to zero. On the record so far, that is the version of the layered approach that works.
Getting to the point where humans are rarely needed
The evidence stops there. What follows is extrapolation. No field has yet reached the point where a human reviewer of AI output is called in only rarely, and no one can say when one will. The researchers working closest to the problem, though, have described the conditions that would have to be met. Their work points to five developments.
The first is that the human job moves upstream, from checking answers to checking questions. Martin Kleppmann, a computer scientist at the University of Cambridge, predicted last December that AI would make formal verification of software routine, because “the proof checker will reject any invalid proof and force the AI agent to retry.” If that happens, he wrote, “we wouldn’t even need to bother looking at the AI-generated code any more.” He was clear about what would remain for people: “the challenge will move to correctly defining the specification,” and “reading and writing such formal specifications still requires expertise and careful thought” (Kleppmann, 2025). Leonardo de Moura, who created Lean, made the same case in February. “As AI takes over implementation, specification becomes the core engineering discipline,” he wrote. He summed up today’s review practice in two short sentences: “The errors are there. The reviewers are not” (de Moura, 2026a).
The appeal of this shift is arithmetic. A specification, the statement of what a program must do or what a theorem claims, is usually much shorter than the code or proof that satisfies it. A mathematician who can trust a machine-checked proof needs to read only the statement, often a paragraph, and can skip the 100-page argument. That cuts the human workload by a large factor and keeps a person in charge of the part that defines the goal.
The weak point is the specification itself. In June, Pawan Sasanka Ammanamanchi, Siddharth Bhat, and Stella Biderman audited roughly 10,000 problems in the collections used to test AI theorem provers. Their audit produced 4,833 findings, 398 of them backed by a machine-checkable certificate that the formal statement was either unprovable or vacuous, meaning it could not be proved or proved nothing of substance. A machine-checked proof, they noted, “does not verify that the statement faithfully encodes the intended informal problem” (Ammanamanchi et al., 2026). This is the same question Tao handed to a language model in Palomar. Reaching “rarely needed” would require tools that check statements as rigorously as Lean checks proofs: for example, AI systems that translate each formal statement back into plain language, compare it with the author’s claim, and flag mismatches for a person to inspect.
The checker also needs checking. In a March post titled “Who Watches the Provers?”, de Moura wrote that “the piece of Lean you actually have to trust is tiny” and that “Lean is designed so that anyone can build independent kernels.” He recounted that in 2022 an independently written checker called Nanoda rejected an invalid proof that Lean 4’s official checker had accepted, which exposed a bug (de Moura, 2026b). Running several independently built checkers is the Swiss cheese model used as Reason intended, with slices made by different hands and holes in different places.
The most ambitious version of this program reaches beyond mathematics and software. In 2024 a group that included David “davidad” Dalrymple, now director of ARIA’s Safeguarded AI program, along with Yoshua Bengio, Stuart Russell, Max Tegmark, and Christian Szegedy, proposed what they called “guaranteed safe AI.” In their framework, an AI system’s actions would come with “an auditable proof certificate” showing that they satisfy “a mathematical description of what effects are acceptable,” judged against “a mathematical description of how the AI system affects the outside world” (Dalrymple et al., 2024). Writing such descriptions for messy domains such as medicine, law, or diplomacy is still an open research problem.
The second development is that autonomy gets earned through measured evidence, as it has been for driverless cars. Robotaxis offer the clearest working example of a human who is rarely consulted. Waymo says its fleet of about 3,000 vehicles is supported by “approximately 70 Remote Assistance agents on duty worldwide at any given time.” Those agents do not watch the cars. “RA does not continuously monitor a vehicle or set of vehicles,” the company wrote in February. The agents “respond to specific requests for information initiated by the Waymo Driver” and “provide advice which the system can decide to use or reject” (Waymo, 2026a). That works out to about one person on duty for every 43 cars, and the machine decides when to ask.
Waymo’s path to that ratio ran through years of supervised testing and a published safety record. By the company’s count, its cars had driven more than 220 million fully autonomous miles through March 2026, with “94% fewer crashes causing serious or fatal injuries” and “82% fewer crashes involving any reported injury” than human drivers in the same areas (Waymo, 2026b). The record still has gaps. An investigation by U.S. Senator Edward Markey’s office this year found that Waymo and six other companies declined to say how often their vehicles ask for remote help (Office of Senator Edward J. Markey, 2026).
Two ideas from other fields could carry this model over to AI agents. Kevin Feng, David McDonald, and Amy Zhang of the University of Washington proposed five levels of agent autonomy, defined by the human’s role at each level: operator, collaborator, consultant, approver, and observer. They suggested “AI autonomy certificates” that would set how much independence a particular agent may exercise (Feng et al., 2025). The observer is the rarely consulted human. The certificate is the evidence that a system has earned that level of trust.
Insurance supplies the second idea. The Artificial Intelligence Underwriting Company has built a certification standard, AIUC-1, that runs “more than 5,000 adversarial simulations” against an AI agent; companies that pass can buy insurance against the agent’s failures. When the voice-AI company ElevenLabs obtained coverage in February, its agents had passed “5,835 technical tests across 14 risk categories.” AIUC’s chief executive, Rune Kvist, said insurers were “so confident in this certification-based approach that they’re offering AI-specific financial coverage to those who earn it” (ElevenLabs, 2026). An insurer that pays for failures has a financial reason to measure them carefully. That gives companies a way to pull their people back from routine review while someone still answers for the risk.
The third development is that the monitors themselves get tested as adversaries. Layered review works only if the screening layers resist being fooled, and recent research points both ways. A 41-author paper led by Tomek Korbak, whose co-authors include Bengio, Shane Legg, and Geoffrey Irving, argued that reading a reasoning model’s “chain of thought,” the steps it writes out on the way to an answer, offers a new way to catch misbehavior. The authors also warned that “CoT monitoring is imperfect and allows some misbehavior to go unnoticed” and that the opportunity “may be fragile” (Korbak et al., 2025). OpenAI researchers reported in December that such monitoring “may provide a more scalable path for control than naive monitoring” (OpenAI, 2025). In May, Ethan Elasky, Frank Nakasako, and Naman Goyal found that AI critics help a weaker judge evaluate a stronger model, and that “a single independent critique recovers the bulk of debate’s benefit at lower inference cost” (Elasky et al., 2026).
Two papers posted in the last days of September show the promise and the risk side by side. Qishuai Jing applied separation of duties, a long-standing principle of financial controls, to AI agents. His design splits each task among a planner, a “policy gate,” an executor, and an auditor that sit outside the model. In a controlled test, “effective attack success falls from 98.3% under direct execution to 7.7% deployed” (Jing, 2026). Two days later, Deema Alnuhait and colleagues reported that helpful AI agents in simulated workflows found ways around a monitor “without adversarial incentives,” disguising a password-like credential so that it “evades the monitor.” Their summary finding was four words long: “Monitoring favors concealment” (Alnuhait et al., 2026). A monitor that agents learn to route around adds cost without adding protection. Reaching “rarely needed” means attacking the screening layers as hard as the systems they screen, which is the red-team testing that the AI control researchers at Redwood Research and Google DeepMind build into their designs (Greenblatt et al., 2023; Phuong et al., 2026).
The fourth development is that scarce human attention gets spent where machines cannot go, and the people who spend it get rewarded. Tao does not expect people to leave mathematics’ review process. In his essay for the International Congress of Mathematicians, posted in August, he wrote that “AI-generated proofs will accumulate faster than they can be verified” and that “I do not believe that human referees can be removed from the publication process.” He set a standard that no screening layer can meet alone: “A proof that no human can properly explain should be viewed as incomplete, even if it has been formally verified” (Tao, 2026c). His lecture slides propose an automated gate of the kind this article describes, “Journals could automatically reject papers that are flagged for inadequate verification or exposition,” together with a cultural shift to “increase the emphasis on ‘proof digestion'” (Tao, 2026b).
That suggests a division of labor. Machines would handle correctness wherever mechanical checking is possible. Human experts would concentrate on significance, originality, credit, and explanation. Under that arrangement, people would rarely be asked whether a result is right and would be asked constantly whether it matters and what it means. That second job is what the 28 Fields medalists called “the crucial human transmission chain” (Fields Medalists, 2026).
The fifth development is that institutions decide the matter on purpose. Many organizations are already reaching “rarely looped” by default. In a survey released September 15, EY found that 85 percent of senior AI executives at large U.S. companies using agentic AI said at least some of those systems act without real-time human involvement. Twenty-six percent said their organizations could not identify AI agents operating internally without authorization, and 47 percent said their organizations had skipped their own AI governance process for urgent deployments. John McLain of EY said, “The biggest agentic AI risk is that human oversight hasn’t evolved accordingly” (EY, 2026). In many companies the human is already stepping out of the loop. What remains open is whether mechanical checkers, track records, certificates, tested monitors, and trained reviewers are in place when that happens.
Put those five developments together and it is possible to sketch what a release like OpenAI’s might look like in a few years. Each manuscript would arrive with a Lean proof confirmed by two independently written checkers. An AI system would translate every formal statement back into plain language and flag any mismatch with the paper’s claim, and a human expert would read only the flagged statements, a page or less apiece. Search agents would compare each result with the published literature and flag overlaps of the kind that tripped up the Erdős Problem 333 solvers. A screening model with a published accuracy record would rank the results by likely importance, and a random sample would go to human experts to check the ranking. Mathematicians would then spend their time where Tao and the Fields medalists say it belongs: on the few dozen results worth digesting, explaining, and building on. In that picture, the human in the loop is rarely asked whether a proof is correct. People are still asked what it means, and that question has no machine checker.
References
Advanced Research and Invention Agency. (2025, November 28). AI progress and a Safeguarded AI pivot. https://aria.org.uk/insights/ai-progress-and-a-safeguarded-ai-pivot
Alnuhait, D., Wang, G., Khalifa, M., & Peng, H. (2026). Covert assistance: Helpful LLM agents evade oversight in multi-agent systems (arXiv:2609.39050). arXiv. https://arxiv.org/abs/2609.39050
Ammanamanchi, P. S., Bhat, S., & Biderman, S. (2026). Faults in our formal benchmarking: Dataset defects and evaluation failures in Lean theorem proving (arXiv:2606.29493). arXiv. https://arxiv.org/abs/2606.29493
Anthropic. (2026, April 14). Automated alignment researchers. https://www.anthropic.com/research/automated-alignment-researchers
Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779. https://doi.org/10.1016/0005-1098(83)90046-8
Biswas, J., Schoepp, S., Vasan, G., Opipari, A., Zhang, A., Hu, Z., Joseph, S., Lease, M., Li, J. J., Stone, P., Wagstaff, K. L., Taylor, M. E., & Jenkins, O. C. (2026). AI-assisted peer review at scale: The AAAI-26 AI review pilot (arXiv:2604.13940). arXiv. https://arxiv.org/abs/2604.13940
Dalrymple, D., Skalse, J., Bengio, Y., Russell, S., Tegmark, M., et al. (2024). Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems (arXiv:2405.06624). arXiv. https://arxiv.org/abs/2405.06624
de Moura, L. (2026a, February 28). When AI writes the world’s software, who verifies it? https://leodemoura.github.io/blog/2026-2-28-when-ai-writes-the-worlds-software-who-verifies-it/
de Moura, L. (2026b, March 16). Who watches the provers? https://leodemoura.github.io/blog/2026-3-16-who-watches-the-provers/
Drake, J. (2026, April 30). AI slop is flooding academic journals. A top journal measured it. Forbes. https://www.forbes.com/sites/johndrake/2026/04/30/ai-slop-is-flooding-academic-journals-a-top-journal-measured-it/
ecancer. (2026, January 30). AI-supported mammography screening results in fewer aggressive and advanced breast cancers, finds full results from first randomised controlled trial. https://ecancer.org/en/news/27721-ai-supported-mammography-screening-results-in-fewer-aggressive-and-advanced-breast-cancers-finds-full-results-from-first-randomised-controlled-trial
Elasky, E., Nakasako, F., & Goyal, N. (2026). Debate helps weak judges reward stronger models (arXiv:2605.27483). arXiv. https://arxiv.org/abs/2605.27483
ElevenLabs. (2026, February 12). ElevenLabs secures first-of-its-kind AI agent insurance. https://elevenlabs.io/blog/aiuc-announcement
Emergent Mind. (2026). OpenAI math explorer. Retrieved October 10, 2026, from https://www.emergentmind.com/openai-math-explorer
EY. (2026, September 15). EY survey finds that autonomous AI implementation outpaces oversight, yielding an AI governance gap [Press release]. https://www.ey.com/en_us/newsroom/2026/09/ey-survey-finds-that-autonomous-ai-implementation-outpaces-oversight-yielding-an-ai-governance-gap
Feng, K. J. K., McDonald, D. W., & Zhang, A. X. (2025). Levels of autonomy for AI agents (arXiv:2506.12469). arXiv. https://arxiv.org/abs/2506.12469
Fields Medalists. (2026, September 11). A severe misalignment of AI in mathematics [Declaration]. https://doi.org/10.5281/zenodo.22737750 (also at https://mathandai.org/)
Gartner. (2025, June 11). Gartner predicts that guardian agents will capture 10-15% of the agentic AI market by 2030 [Press release]. https://www.gartner.com/en/newsroom/press-releases/2025-06-11-gartner-predicts-that-guardian-agents-will-capture-10-15-percent-of-the-agentic-ai-market-by-2030
Goel, S., Struber, J., Auzina, I. A., Chandra, K. K., Kumaraguru, P., Kiela, D., Prabhu, A., Bethge, M., & Geiping, J. (2025). Great models think alike and this undermines AI oversight. In Proceedings of the 42nd International Conference on Machine Learning. https://arxiv.org/abs/2502.04313
Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2023). AI control: Improving safety despite intentional subversion (arXiv:2312.06942). arXiv. https://arxiv.org/abs/2312.06942
Hart, R. (2026, October 9). “Pure insanity”: Mathematicians will need years to make sense of OpenAI’s latest drop. The Verge. https://www.theverge.com/ai-artificial-intelligence/1008726/openai-mathematics-solutions-chaos
Jing, Q. (2026). Separation of duties for privileged LLM agents: A governed execution architecture with measured security-utility trade-offs (arXiv:2609.38224). arXiv. https://arxiv.org/abs/2609.38224
Kakaes, K. (2026, August 3). Why the legendary Erdős problems are falling to AI. Quanta Magazine. https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/
Khollam, A. (2026, October 6). OpenAI’s largest math release tackles 4,000 problems with Lean proofs. Interesting Engineering. https://interestingengineering.com/ai-robotics/openai-largest-math-release-lean-proofs
Kleppmann, M. (2025, December 8). Prediction: AI will make formal verification go mainstream. https://martin.kleppmann.com/2025/12/08/ai-formal-verification.html
Korbak, T., et al. (2025). Chain of thought monitorability: A new and fragile opportunity for AI safety (arXiv:2507.11473). arXiv. https://arxiv.org/abs/2507.11473
Lukić, D. (2026, June 3). The AI code verification bottleneck: Why faster code generation means slower reviews. Aviator. https://www.aviator.co/blog/the-ai-code-verification-bottleneck-why-faster-code-generation-means-slower-reviews/
Mitchell, M., Ghosh, A., & Passi, S. (2026). AI agents push humans out of the loop (arXiv:2608.23642). arXiv. https://arxiv.org/abs/2608.23642
Office of Senator Edward J. Markey. (2026). Remote assistance investigation report. United States Senate. https://www.markey.senate.gov/imo/media/doc/remote_assistance_investigation_report.pdf
OpenAI. (2025, December 18). Evaluating chain-of-thought monitorability. https://openai.com/index/evaluating-chain-of-thought-monitorability/
OpenAI. (2026, October 6). Sharing AI progress in mathematics. https://openai.com/index/sharing-ai-progress-in-mathematics/
Phuong, M., Jenner, E., Simon, L., Ho, L., Shah, R., Farquhar, S., Coull, S., et al. (2026). GDM AI control roadmap (v0.1) (arXiv:2607.13087). arXiv. https://arxiv.org/abs/2607.13087
Reason, J. (2000). Human error: Models and management. BMJ, 320(7237), 768–770. https://doi.org/10.1136/bmj.320.7237.768
Ryan, A. (2026, August 2). Medical journals, law reviews, and literary magazines are grappling with AI-generated submissions. The Boston Globe. https://www.bostonglobe.com/2026/08/02/business/literary-academic-journals-artificial-intelligence/
Sawers, P. (2026, March 12). Claude Code’s new AI code review agents scan pull requests for bugs. Tessl. https://tessl.io/blog/anthropic-launches-ai-code-review-agents-that-scan-pull-requests-for-bugs
Skuse, B. (2026, March 2). Watershed moment for AI–human collaboration in math. IEEE Spectrum. https://spectrum.ieee.org/ai-proof-verification
Tao, T. [@tao@mathstodon.xyz]. (2026a, April 27). [Thread on the “era of proof abundance”] [Post]. Mathstodon. https://mathstodon.xyz/@tao/116477351524980995
Tao, T. (2026b, July 24). Mathematics in the age of AI [Lecture slides, International Congress of Mathematicians]. https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf
Tao, T. (2026c, August 17). Mathematics in the age of AI (arXiv:2608.16753). arXiv. https://arxiv.org/abs/2608.16753
Tao, T. (2026d, August 18). Palomar – a registry of Lean verified mathematics. What’s New. https://terrytao.wordpress.com/2026/08/18/palomar-a-registry-of-lean-verified-mathematics/
Tsoukalas, G., Kovsharov, A., Shirobokov, S., et al. (2026). Advancing mathematics research with AI-driven formal proof search (arXiv:2605.22763). arXiv. https://arxiv.org/abs/2605.22763
Waymo. (2026a, February 17). Advice, not control: The role of remote assistance. https://waymo.com/blog/shorts/advice-not-control-the-role-of-remote-assistance/
Waymo. (2026b, June 24). From the road — June 24, 2026. https://waymo.com/blog/shorts/safetydata-june26/
###
Filed under: Uncategorized |















































































































































































































































































































































































































































































































































































Leave a Reply