Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2025 R&D 100 Award Winners
    • 2025 Professional Award Winners
    • 2025 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Inside AutoLabs: PNNL’s self-correcting AI still needs an expert in the loop

By Julia Rock-Torcivia | July 29, 2026

Researchers from Pacific Northwest National Laboratory (PNNL) developed a generative agentic AI that can translate experimental goals into instructions for a laboratory robot. Their paper was published in Scientific Reports.  

Systems engineer Heather Job works with autonomous laboratory robot Big Kahuna. Job helped develop a new AI program that can help researchers seamlessly design experiments for the robot to conduct.
Photograph credit: Andrea Starr | PNNL

Gihan Panapitiya, a data scientist at PNNL and the lead author of the paper, said the researchers had plans to extend AutoLabs to literature review, persistent memory and other autonomous lab equipment besides the automated lab workstation Big Kahuna from Unchained Labs. 

The self-checks catch errors before instructions reach the hardware. The system can spot errors before instructions reach the physical hardware. “We are also developing additional feedback mechanisms that will allow us to continuously refine and introduce new guardrails over time,” Panapitiya, said. Designing experiments for automated laboratory instruments like Big Kahuna traditionally requires equal understanding between scientists and engineers, whose expertise often does not overlap. The team built AutoLabs to close that gap using natural-language dialogue instead of manual translation of intent into code.  

AutoLabs follows a string of LLM-driven lab automation systems such as Carnegie Mellon’s Coscientist, the Chinese University of Hong Kong and Zhejiang Lab’s Chemist-X, and the University of Toronto’s ORGANA. Carnegie Mellon published Coscientist in Nature in December 2023.  

Still a case study  

The researchers tested 20 distinct agent architecture configurations, each run ten times. The agents were tested on five benchmark experiments of increasing complexity, from a simple naphthalene calibration sample to a multi-plate, timed esterification synthesis. The full physical robot was able to fully execute two of the five experiments.  

The paper notes that, in its final form, AutoLabs achieved 100% hardware-instruction-loading success for experiments one through four, with “occasional issues” for experiment five due to multi-plate complexity.  

This evaluation was confined to a single lab and a single instrument. The paper described its results as a “case study” on Big Kahuna. The authors note that the underlying architecture could theoretically extend to other liquid handling platforms.  

How AutoLabs checks its own work  

AutoLabs has a self-correction mechanism that “identifies and fixes errors in a protocol before it is sent to the robot,” Panapitiya said in an email to R&D World.  

“We pass the agent-generated protocol to the LLM along with a checklist of specific criteria to verify,” he said. “This self-check process can be repeated multiple times, iterating until the system is confident that the protocol contains no remaining errors.”  

The paper describes guided and unguided self-checks. Guided checks improve procedural correctness, while unguided checks are more effective at catching chemical amount errors, but can introduce inefficiencies such as redundant water-addition steps.  

The self-check process checks unit consistency, minimal-step chemical additions, correct handling of transfers between plates, delay and timing checks, plate array validation, solvent specification review and chemical addition tagging structure.  

“We are also developing additional feedback mechanisms that will allow us to continuously refine and introduce new guardrails over time,” Panapitiya said.  

Guardrails with gaps  

The paper provides a list of the guardrails: self-correction and validation, tool-based stoichiometric calculations grounded in a curated chemical database, prompts asking users to confirm ambiguous details, validation against hardware-specific constraints, human review before execution and logging for post-hoc traceability.  

To strengthen the safeguards, the researchers tested retrieval-augmented generation (RAG) to reduce errors by grounding the understand and refine agent in prior validated procedures, with mixed results. In one configuration, RAG improved outcomes across nearly all experiments, but in another, it made results worse. The paper attributes the inconsistency to retrieved information, sometimes introducing “noise, irrelevant information or conflicting instructions,” rather than clarity.  

When an expert systems engineer worked with AutoLabs, results improved significantly over both the fully automated system and non-expert users. But even the expert missed some errors, such as omitted stir-rate commands, cap values assigned to the wrong plate and an unintended delay step that disabled a timed experiment.  

Non-expert users who answered AutoLabs’ clarifying questions produced less accurate results than letting the system proceed on its own judgement.  

Multiple agents 

AutoLabs was initially developed as a single agent before it was expanded to a multi-agent tool to prevent the AI from losing track of its instructions, Panapitiya said.  

“We observed that as a conversation grows longer, the LLM tends to lose track of the instructions it was originally given,” he said. “To address this limitation, we moved to a multi-agent architecture… Because every sub-agent begins each task from its own instruction set, the risk of the model ‘forgetting’ its guidance is greatly reduced.” 

The final version of AutoLabs features multiple specialized agents, each responsible for a specific task and equipped with its own dedicated set of instructions, Panapitiya said. A supervisor agent delegates tasks to the sub-agents. According to the paper, AutoLabs features five sub-agents: understand and refine, which clarifies the request; chemical calculations, which performs stoichiometry calculations via tool-calling; vial arrangement; processing steps, which is in charge of heating, stirring and timing; and final steps.  

“We are actively working to expand the agent’s capabilities to include literature review, hypothesis generation, persistent memory and continuous self-improvement,” Panapitiya said.  

 

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

Claude Opus 5 outscores Fable 5 on 8 of 13 benchmarks at half the token price
As U.S. power demand surges, Brookhaven taps AWS in aim to drive 20-100x faster power grid decision-making with GridSearch AI
Purple Glowing Spiral Fractal Background Image, Illustration - Vortex repeating spiral patterns, Symmetrical repeating geometric patterns. Abstract design, black background
Recursion says its Norstella-backed real-time simulation can expand trial eligibility by up to 40%
How EcoBOT’s automated plant lab knows where it’s guessing 
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2025 R&D 100 Award Winners
    • 2025 Professional Award Winners
    • 2025 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE