Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

AI is wreaking havoc on math and coding. Where could it have the biggest impact next?

By Brian Buntz | October 8, 2026

Big data technology Data science analysing artificial intelligence generative AI deep learning machine learning algorithm Neural flow network analytics innovation abstract futuristic. 3d rendering.

[Adobe Stock]

Machines are now solving open math problems and rebuilding complex software on their own.

What they do next to the rest of science is an open question. Some of the clearest gains have come in tasks where answers can be checked automatically, using a proof checker or a test suite.

Still, frontier AI is moving fast enough that the CEOs of Anthropic and OpenAI have both said the industry needs to “pace the frontier” after the labs began rapidly releasing new models, thanks in part to new training methods that are helping drive advances in areas like math, coding and cybersecurity. OpenAI in particular has drawn scrutiny. In July, roughly 700 of its AI agents running inside a cybersecurity evaluation took part in an attack on the infrastructure of Hugging Face, the AI model-hosting platform. In September, the company claimed its models had solved the Navier–Stokes Millennium Prize Problem, which asks whether smooth three-dimensional fluid flows can develop singularities. OpenAI’s claimed proof applies a smooth external force, which Clay’s rules allow, and the Clay Mathematics Institute now lists the problem separately, under “Active problems” on its website.

Judging AI’s impact

Global investment in AI infrastructure is projected to reach roughly $1 trillion this year, according to Goldman Sachs, yet what that will buy science is still unsettled. So far the AI payoff is hard to measure overall. The examples below are sorted by what has measurably changed, not by how widely a tool is used.

Bucket What earns placement Examples
Active disruption Repeated evidence that AI substantially changes core work: smaller teams producing comparable validated output, major cycle-time reductions, or substantial execution delegated to AI. Example: a step labs now routinely hand to software, such as getting a first predicted model of a protein’s shape. Code generation, where over half of developers surveyed by JetBrains said they wrote less than 20% of their code without AI assistance; protein structure prediction, recognized with the 2024 Nobel Prize in Chemistry
Emerging impact Useful capabilities and measurable gains in particular steps, but limited evidence of transformation across the broader workflow. Example: a tool that speeds up one step, such as screening papers, while the rest of the project runs much as before. Picking drug targets; screening the literature; drafting regulatory documents
Too early to judge outcomes Insufficient evidence that workflow gains improve the ultimate result. Example: a drug candidate that reached the clinic faster with AI’s help but hasn’t yet shown it works better in patients. Whether AI-discovered drugs succeed more often in human trials; whether AI speeds up wet-lab experiments

Active disruption

Software engineering: Developers now direct more code than they write

Software engineering remains one of the clearest cases of AI reshaping daily practice, from drafting specifications and writing scripts to debugging code and preparing commits. In a few years, writing code has gone from something developers write character by character to something they increasingly direct and review. In an evaluation by the research group Epoch AI, an agent running Anthropic’s Claude Opus 4.6 reimplemented an existing bioinformatics toolkit using documentation, tests and black-box access to the original program, passing 2,000 of 2,001 tests. Newer models continue to show gains in coding sophistication. In independent testing from Vals AI, GPT-6 Astra passed about 60% of Terminal-Bench 4.0’s multistep technical tasks, compared with 38% for GPT-5.6 Sol.

That shift from coding chatbot to tireless agentic coding sidekick took less than three years. After ChatGPT launched in late 2022, many developers copied a chatbot’s suggestions from a browser window, when they managed to be correct, and pasted them into their code editor. By 2025, coding agents such as Anthropic’s Claude Code and OpenAI’s Codex could take on a task and write, run and test the code themselves. In a JetBrains survey of more than 15,000 professional developers conducted from May through July, over half said they wrote less than 20% of their code without AI assistance. That includes both assisted coding and code generated entirely by agents. Measuring the payoff is harder. METR, a nonprofit that found in a 2025 trial that experienced open-source developers were 19% slower with AI tools, reported in February that its follow-up study was compromised because many developers refused to work without AI.

Some prominent skeptics have changed their minds. Mo Bitar, a software engineer and YouTuber, said an AI agent “failed catastrophically” on a project spec in March, so badly “that I dedicated my life to tearing apart the scam of AI coding.” In an Oct. 2 video, he described handing Anthropic’s Claude Opus 5.5 a 2,700-line spec for a feature of Enjoy, an app he makes. In about seven hours, by his account, Claude spun up dozens of agents, including adversarial reviewers that checked one another’s work, and built the feature. “Everything just worked,” he said. Bitar now argues that automated tests and agent reviews, not human code review, are the checks that matter.

Mathematics: AI proofs arrive faster than mathematicians can check them

In mathematics, AI is now producing research results faster than experts can check them, and the field is split over whether that counts as progress. On Tuesday, OpenAI posted more than 700 AI-generated math papers to GitHub, saying an unreleased model produced them at an average of about three hours of ChatGPT Pro computing per result. The papers are what survived heavy filtering. OpenAI says it started with about 4,000 problems, and 372 result families made the cut. Roughly 42% of the main results come with proofs a computer can check. A day later, OpenAI withdrew three of the papers after finding a sign error. OpenAI isn’t alone. In May, Google DeepMind researchers reported that an agent pairing a language model with the Lean proof checker resolved 9 of 353 open Erdős problems at a few hundred dollars each. On Oct. 1, Matthew Schwartz, a visiting researcher at Anthropic, described 36 manuscripts in 18 fields produced with the company’s Claude models.

Some math experts are challenging how the results were produced, released and evaluated. “Mathematicians did not ask for this work to be done,” the Association for Human Mathematics said in a statement. The group rejected OpenAI’s assertion that the release advances the field and urged “mathematicians and the public to view the value of this publication model with due skepticism.” On Sept. 29, the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), nine researchers hosted by the Institute for Advanced Study in Princeton, had asked frontier labs “to stop testing advanced mathematical problems on proprietary models” and to formalize their proofs as far as possible. Asked about OpenAI’s release, the group said it is “ultimately up to the mathematical community to assess” whether its recommendations were followed.

Fields Medalist Terence Tao warned of the costs to math research. Mass-produced AI solutions are arriving without the talks, collaboration and explanation that normally help mathematicians understand a breakthrough, he wrote, and a solved problem cannot simply be made open again: “even the mere knowledge that a solution exists ‘contaminates’ efforts by both humans and AI to find alternate routes to the problem that reveal additional insights.” Solutions to open problems, he wrote, are now being “harvested at large scale in an unsustainable fashion.”

Not everyone sees a loss. “My view is that this is great for mathematics,” Dan Litt, a mathematician at the University of Toronto, told Fortune. Michael Douglas, a Harvard theoretical physicist, told Scientific American that the release marked the start of “a new era of math and mathematical physics.”

Praise came from a rival lab, too. Levent Alpöge, a mathematician at Anthropic, called the release “the most significant moment in mathematical history,” according to The Rundown.

Protein structure prediction: Predicted structures become routine

For decades, learning what a protein looks like meant coaxing it into a crystal and firing X-rays at it, or freezing it and imaging it under an electron microscope. Those experiments built the Protein Data Bank, the public archive of solved structures, over half a century. AlphaFold changed the scale. Google DeepMind’s model has predicted the structure of “virtually all the 200 million proteins that researchers have identified,” according to the Nobel committee, and by October 2024 more than two million people in 190 countries had used it.

Despite heavy investment, it is still early for AI in pharma, where a drug often takes a decade to go from discovery to approval. One task, though, has already been transformed. Determining a protein’s 3D structure in the lab can take months or years; AI models now make useful structure predictions readily available, though experiments are still needed to confirm structures, interactions and function. The 2024 Nobel Prize in Chemistry went to Demis Hassabis and John Jumper of Google DeepMind for AlphaFold, shared with David Baker for computational protein design. Among roughly 100 biotech and pharma organizations already using AI surveyed in November 2025, 71% reported using protein structure and property models, according to Benchling’s 2026 Biotech AI Report. A June survey of 50 U.S. pharma executives who oversee R&D technology, by Citi Research, put adoption of protein structure prediction and structural biology tools at 78%, the highest of any use.

The predictions are strongest where experimental data already exist. The Protein Data Bank and public binding databases “overwhelmingly describe well-folded proteins bound at their primary sites,” scientists at Seattle-based Talus Bio wrote in a July preprint. That leaves out most of what drug hunters still want to hit: of roughly 20,000 human proteins, 704 are targets of an approved drug and 1,904 more have a potent small-molecule ligand, “leaving 87% with neither,” they wrote. Many of the rest, such as transcription factors, lack stable pockets because parts of them are floppy, or “intrinsically disordered.” In a virtual-screening test on seven such targets, drawn from Talus’s proprietary data, the structure-prediction model Boltz-2 “fell to near-random” at ranking which molecules bind, while Talus’s model, Ptarmigan-1, which works without a 3D structure, held up. Boltz-2 outperformed Ptarmigan-1 on well-folded proteins and on a binding pocket in the cancer target KRAS. The preprint has not been peer reviewed.

Weather forecasting: AI forecasts move into daily operations

Weather forecasting is one of the clearest cases outside software and math of AI moving from research papers into daily operations. The European Centre for Medium-Range Weather Forecasts began running its machine-learning model, AIFS, alongside its physics-based system on Feb. 25, 2025, saying it beat physics-based models on many measures, including tropical cyclone tracks by up to 20%, while using about 1,000 times less energy per forecast. NOAA put AI versions of its global models into operation in December 2025; a 16-day forecast from its new AIGFS model uses 0.3% of the computing of the traditional GFS. During the 2025 Atlantic hurricane season, a Google DeepMind model supplied to the National Hurricane Center showed slightly better short-range track performance than the center’s official forecasts, according to its March verification report. Forecasting has the fast feedback that code and math have: every forecast is checked against the actual weather within days. Storm strength is still harder to call: NOAA said the first version of AIGFS shows “a degradation in tropical cyclone intensity forecasts, which future versions will address.”

Emerging impact

Drug discovery: Models speed target picking, but the verdict comes years later

Beyond structure prediction, machine-learning models are speeding up individual steps of drug discovery. Pharma executives expect the biggest payoff from target identification, where models mine genetic, protein and patient data to rank which proteins are worth drugging. In Citi’s survey, 70% picked it as the step most likely to improve success rates, and Benchling found that 58% of its AI-using organizations already apply AI there. Whether that choice translates into patient benefit may not be clear until clinical trials, often years later.

Literature review: AI reads faster, and scientists spend time checking it

Reading and summarizing the scientific literature has become one of AI’s most common jobs in R&D, though its speed is easier to measure than its quality. Literature extraction is the most common AI use in Benchling’s survey, at 76% of AI-using organizations. More than 600 U.S. and U.K. scientists in a September study by researchers from Google, Google DeepMind and MIT FutureTech reported net savings of about seven hours a week across all their AI uses, alongside substantial time spent verifying AI output: 89% of those who save time spend more than a tenth of it checking, and 46% more than a quarter.

Regulatory writing: AI drafts the paperwork, for drugmakers and the FDA

Regulatory paperwork may be where pharma sees the clearest payoff so far. In an April poll at the Pistoia Alliance’s London conference, which drew about 300 attendees, 54% of respondents named regulatory submissions and reporting teams as seeing the greatest benefit from AI, well ahead of research analysis at 21%. The conference poll was not a representative survey of the industry.

Drugmakers and regulators are both betting on it. Novo Nordisk says AI has cut the time to write clinical study reports, the documents that summarize a trial’s results for regulators, by 90%, with drafts still going to people for review, according to a customer case study published by Anthropic, whose Claude models the company uses. On the other side of the table, the FDA launched a generative AI tool, Elsa, in June 2025 for tasks including clinical protocol reviews, and agency leaders wrote that AI could give a “first-pass” review of applications that can run past 500,000 pages. Early reports by STAT and NBC News found Elsa sometimes gave inaccurate answers.

Chemistry: Robots run the experiments; chemists pick them

In chemistry, AI can now propose and run large experimental campaigns, but scientists still choose which ideas reach the bench. In June, OpenAI and the startup Molecule.one reported that GPT-5.4 generated and ranked thousands of proposals for improving several classes of reactions used in medicinal chemistry. Chemists picked four to pursue, and for one of them, a stubborn Chan–Lam coupling, Molecule.one’s automated lab ran 10,080 reactions, finding that the additive TEMPO improved product formation. OpenAI called the workflow “near-autonomous.” In November 2024, a University of Liverpool team led by Andrew Cooper demonstrated in Nature mobile robots using a rule-based decision system to operate shared synthesis and analysis instruments and select reactions for further study. Both systems worked on problems people defined. Neither shows that AI can handle unfamiliar chemistry or scale a reaction up for manufacturing.

Materials: Candidates pile up faster than labs can test them

Machine-learning models, from the graph neural networks that screen known chemistry to generative models that design new crystal structures, now propose candidate materials faster than labs can make and test them. In 2023, Google DeepMind said its GNoME model had predicted 380,000 stable new crystals; 736 of them had also been independently synthesized by other researchers. Microsoft Research’s MatterGen, published in Nature in January 2025, generates crystal structures aimed at a target property; as a check, collaborators at the Shenzhen Institutes of Advanced Technology synthesized one of its proposals, TaCr2O6. MatterGen’s authors described the material they made as a compositionally disordered version of the predicted structure, and an April paper in Materials Horizons argued that it matches a compound first reported in the early 1970s and already in MatterGen’s training data. This year, a team including Zonglong Zhu reported in Nature a system that pairs machine-learning discovery with an automated fabrication line. It found a molecule that lifted small perovskite solar cells to 27.22% efficiency, and the automated line’s efficiency results were nearly five times as reproducible as manual fabrication. The team also built mini-modules and ran 1,200 hours of stability tests, but commercial-scale production and years of field performance are still unproven.

Genomics: Models predict what DNA changes do, pending lab tests

In genomics, deep-learning models that read raw DNA sequence are getting better at predicting what a DNA change does, but those predictions still have to be tested in the lab. Google DeepMind’s AlphaGenome, published in Nature in January, reads up to a million DNA letters at once and predicts thousands of molecular signals, from gene expression to RNA splicing. In the Nature paper, it matched or beat the best existing models on 25 of 26 tests of how variants affect gene regulation. The company also says the model struggles with regulatory elements more than 100,000 letters away and is not designed or validated for clinical use or for predicting an individual’s genome.

Brain mapping: AI traces brain wiring at a new scale

In neuroscience, neural networks that trace cells through electron-microscope images have made it possible to map brain wiring at a larger scale than before, and the first maps are already yielding rules for how neurons connect. In April 2025, the MICrONS consortium published a 3D map of a cubic millimeter of mouse visual cortex containing more than 200,000 cells and more than 500 million synapses, using AI and machine-learning tools to annotate the neurons and their connections. First, the researchers recorded the activity of tens of thousands of the neurons while the mouse watched videos, including clips from The Matrix, so the wiring could be compared with what the cells did. The consortium has already reported wiring principles from the data. Whether those insights will translate into treatments remains open.

Instrument data: Deep learning takes the first pass at the firehose

In physics, astronomy and the field sciences, deep-learning models trained on past measurements increasingly take the first pass over data that instruments produce faster than people can sort it. At CERN, the CMS experiment reported in February that a machine-learning algorithm had fully reconstructed proton collisions for the first time, sharpening the reconstruction of particle sprays called jets by 10% to 20% in simulated events. DINGO-BNS, a neural network described in Nature in March 2025, inferred a neutron-star merger’s properties from its gravitational waves in about a second once prepared data were supplied, compared with tens of minutes or hours for other full-inference methods, so telescopes can be pointed at the source sooner. The U.S. Geological Survey reported in 2021 that deep-learning models trained on 136,716 earthquakes detected slightly more quakes than its standard processing and located them more accurately. In ecology, a 2018 study could automate identification for 99.3% of 3.2 million camera-trap images at the same 96.6% accuracy as volunteer teams, which the authors estimated would save 8.4 years of human labeling effort. For specific tasks, such as sorting camera-trap images, the shift is already large. Less clear is how widely these tools have changed routine practice beyond the teams that built them, and deciding what the data mean still falls to scientists.

Too early to judge outcomes

Clinical success: AI-discovered drugs pass early trials, but the real test is later

Whether AI-discovered drugs work better in patients is still an open question, though early trial data offer a tentative signal. A 2024 BCG analysis in Drug Discovery Today found that AI-discovered molecules cleared Phase 1 trials 80% to 90% of the time, against a historical 40% to 65%. Phase 2 success was about 40%, roughly in line with industry averages, from a small sample, and a historical comparison can’t isolate how much of the difference AI made.

Wet-lab work: The biology bench resists the speedup

At the biology bench, where experiments are slow and costly, the open question is whether gains from individual automated systems carry over across labs and whole research programs. In the same Pistoia poll, just 1% of respondents cited value in the wet lab.

Stef van Grieken, CEO of the protein-design company Cradle, says the comparison to code is the mistake. “A lot of people talk about AI-developed drugs and think about it the way they think about AI-developed code,” he told R&D World in July. “The code is there and it works or it doesn’t.” Getting a drug, or any biomolecule, to market, by contrast, “involves a long list of things you have to do to get somewhere.” Each experiment on that list costs hundreds to tens of thousands of dollars and takes weeks to months to report back, he said.

The models also lack what a bench scientist carries. They “still have a lot of room to grow in terms of really understanding the nuances of science in the wet lab, and all the tacit knowledge that someone accumulates throughout their PhD and throughout their career,” said Nicholas Larus-Stone, head of AI and AI agents at the lab-software company Benchling.

Even where the tools work, scientists are reluctant to let software make experimental decisions. “Where you struggle, particularly in the wet labs but really across the board, is this concept of agency,” Christian Baber, chief portfolio officer at the Pistoia Alliance, told R&D World in May. “Scientists don’t want to give up agency completely.”

AI developers are investing in physical labs, too. Anthropic has built its own wet lab in the San Francisco Bay Area, Reuters reported in September. “We believe that to do biology, the final test is still and will be for a while in real lab work,” Eric Kauderer-Abrams, the company’s head of life sciences, told the news agency.

The contrast helps explain why progress has come faster on tasks that can be checked formally. Schwartz, who is also a Harvard physicist, made the case in an Oct. 1 essay. Math, he wrote, is “the one part of science where a problem can be stated completely and an answer checked absolutely. But most of science is not like that.” AI “can sit inside that loop,” he wrote, “but it does not collapse the loop to a point.”

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

After tokenmaxxing backlash, Microsoft moves more agent work onto the desk
Starship’s first orbital flight tests key pieces of SpaceX’s million-satellite data center plan 
Under oath, Google confirms three AI agent test escapes as OpenAI, Anthropic and Meta face NYC lawmakers
What AI-assisted engineering taught us about team design 
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE