Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

Benchling envisions vibe coding cutting R&D app development time from months to minutes

By Brian Buntz | August 28, 2026

BenchBench-Protocol turns published methods and real-world protocol adaptations into expert-reviewed, rubric-scored tasks used to compare frontier and open-weight models. (Image: Benchling)

San Francisco-based Benchling is betting that AI agents can potentially bring something like vibe coding’s rapid build-test-revise loop and targeted automation to scientific work with the requisite level of reproducibility required for science. The aim is to start with data-heavy tasks and eventually extend toward human-supervised wet-lab workflows.

Nicholas Larus-Stone

Nicholas Larus-Stone

“LLMs can work in the wet lab with the right harness and a human in the loop,” Nicholas Larus-Stone, Benchling’s head of AI and AI agents, told R&D World. Benchling is assembling both sides of that proposition: software grounded in scientists’ existing data and workflows, and benchmarks intended to measure, and eventually improve, the models operating underneath it.

To execute the overarching vision, Benchling is building a stack. On August 10, it announced vibe coding capabilities that let scientists create interactive applications from natural-language descriptions. Three days later, it announced the BenchBench-Protocol, which evaluates whether frontier models can reason through real protocol modifications. And then on August 27, it unveiled Benchling Agents, which adds action-taking capabilities, including registering molecules, processing results and updating workflows.

From prompts to custom R&D apps

In its Aug. 10 announcement, Benchling described an available feature that can turn a natural-language request into an interactive candidate dashboard combining efficacy, pharmacokinetics, safety and manufacturability data. The company also previewed a more operational use case involving a purpose-built PK sample-collection interface that would guide a scientist through hundreds of samples, scan tube barcodes, flag issues and record results in Benchling’s underlying data model.

Technically adept customers currently build applications like the proposed sample interface through Benchling’s API, according to the company. “Soon, any scientist will be able to create them in minutes instead of months,” Benchling said.

Benchmarking the wet-lab gap

Scientists are approaching such promises with a mix of FOMO and persistent skepticism. A June Nature poll of 1,907 researchers found that 25% used AI models for research daily and 26% used them weekly. While 48% described their feelings toward AI as broadly negative, 59% worried they would fall behind if they avoided the tools.

Meanwhile, a Pistoia Alliance poll and an Active Site/METR randomized trial suggested that generative AI’s practical value remains limited in wet-lab workflows. Pistoia found just 1% of respondents reporting value in the wet lab, while the randomized trial found similar end-to-end completion rates among novices with and without access to an LLM.

“Our focus is very much on what we see our users and customers actually doing in the lab,” Larus-Stone said. BenchBench derives its tasks from protocols that scientists physically executed, giving it a specific focus on protocol modification and troubleshooting.

The broader landscape includes LifeSciBench, LABBench2 and BioMysteryBench, which evaluate different slices of scientific work. Larus-Stone described one end of that landscape as “book learning,” with textbook-style or synthetically generated questions. BenchBench reflects “a slightly different view of what biological benchmarks should be, and we think that’s why the models are worse at this than they are at some of those other benchmarks,” he said.

BenchBench tests models in isolation, apart from the data, tools and guardrails surrounding a deployed agent. “The benchmark task specifically has no harness,” Larus-Stone said. “It’s just asking the LLMs directly, and it’s only one type of question.” Actual lab work spans a much wider range of tasks.

Benchling’s product layer supplies the harness: software that connects a model to relevant data and tools, governs its actions and routes its work through scientist review. Combine the model’s biological reasoning with that infrastructure and a scientist who can recognize “what is obviously wrong or which suggestions are good ideas,” Larus-Stone said, and “you can have effective AI in the lab.”

From pharma data to verifiable rewards

Benchling is also working with “a number of the frontier labs” to improve model performance on bench-related tasks. Larus-Stone said those collaborations have shown that carefully constructed scientific data can make models better at work Benchling and its customers care about. Producing that data remains difficult because useful training and evaluation tasks need verifiable answers. Pharma companies’ experimental histories could provide the raw material.

“Pharma companies have lots of data, and the question is how you go from that data to creating verifiable tasks that can improve model performance,” he said. Benchling has begun exploring that approach with a handful of customers.

BenchBench has also produced a potentially important finding for companies deciding how much of their AI stack to own: open models performed across much of the same range as frontier models. “That was a little surprising to us,” Larus-Stone said. He called it “a good sign for pharma companies and others who want to build their own intelligence stack,” given the level of capability already available in some open models.

The next step would be to turn a company’s historical data into verifiable rewards, or checkable signals that reinforce correct model outputs. Companies could use those rewards to construct evaluations and training sets, then post-train open-weight models under their control. Larus-Stone said the resulting systems could “match or beat the frontier models, likely at a fraction of the cost.”

The BenchBench-Protocol found open-weight models like Kimi K3 scoring competitively against models from frontier labs, though Opus 5 maintained a notable lead.

“That trend is just emerging, but it’s going to be an important one over the next six to 12 months,” he added.

For the time being, human oversight is likely to remain especially important in the lab, where Larus-Stone said each incremental automation run can cost tens of thousands of dollars. In the near term, he sees agents serving as a check before scientists commit those resources. “I’m going to first ask the agent to help me validate that this is the right experiment to run,” he said.

The promise is to give scientists a wider view of the experiment, with agents handling routine work such as ordering reagents or determining when a chromatography column needs changing while the scientist focuses on the most promising experimental path. As those workflows multiply, the juggling act could move up a level, from managing lab tasks to managing agents. Software-engineering systems already use lead agents to coordinate subagents, and the same pattern has reached experimental-science prototypes such as PNNL’s AutoLabs. Asked whether managing numerous agents could itself become a bottleneck without better systems for coordinating their work and reconciling their outputs, Larus-Stone replied, “I would call that an unsolved problem we’re working on.”

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

Startup Architect Labs says its AI-designed chip beats NVIDIA’s Jetson Orin Nano
SpaceX unveils $100B Louisiana spaceport, targets 175 kW NVIDIA AI system for orbit in 2027
NVIDIA and Apple unveil compact AI computers for robots and agents, claiming 2x and 4x generational AI gains
Businessman working with modern computer virtual dashboard analyzing finance sales data and economic growth graph chart and block chain technology.
Anthropic backers eye $2 trillion valuation. Its projected Q2 revenue was $10.9B
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE