
BenchBench-Protocol turns published methods and real-world protocol adaptations into expert-reviewed, rubric-scored tasks used to compare frontier and open-weight models. (Image: Benchling)
San Francisco-based Benchling is betting that AI agents can potentially bring something like vibe coding’s rapid build-test-revise loop and targeted automation to scientific work with the requisite level of reproducibility required for science. The aim is to start with data-heavy tasks and eventually extend toward human-supervised wet-lab workflows.

Nicholas Larus-Stone
“LLMs can work in the wet lab with the right harness and a human in the loop,” Nicholas Larus-Stone, Benchling’s head of AI and AI agents, told R&D World. Benchling is assembling both sides of that proposition: software grounded in scientists’ existing data and workflows, and benchmarks intended to measure, and eventually improve, the models operating underneath it.
To execute the overarching vision, Benchling is building a stack. On August 10, it announced vibe coding capabilities that let scientists create interactive applications from natural-language descriptions. Three days later, it announced the BenchBench-Protocol, which evaluates whether frontier models can reason through real protocol modifications. And then on August 27, it unveiled Benchling Agents, which adds action-taking capabilities, including registering molecules, processing results and updating workflows.
From prompts to custom R&D apps
In its Aug. 10 announcement, Benchling described an available feature that can turn a natural-language request into an interactive candidate dashboard combining efficacy, pharmacokinetics, safety and manufacturability data. The company also previewed a more operational use case involving a purpose-built PK sample-collection interface that would guide a scientist through hundreds of samples, scan tube barcodes, flag issues and record results in Benchling’s underlying data model.
Technically adept customers currently build applications like the proposed sample interface through Benchling’s API, according to the company. “Soon, any scientist will be able to create them in minutes instead of months,” Benchling said.
Benchmarking the wet-lab gap
Scientists are approaching such promises with a mix of FOMO and persistent skepticism. A June Nature poll of 1,907 researchers found that 25% used AI models for research daily and 26% used them weekly. While 48% described their feelings toward AI as broadly negative, 59% worried they would fall behind if they avoided the tools.
Meanwhile, a Pistoia Alliance poll and an Active Site/METR randomized trial suggested that generative AI’s practical value remains limited in wet-lab workflows. Pistoia found just 1% of respondents reporting value in the wet lab, while the randomized trial found similar end-to-end completion rates among novices with and without access to an LLM.
“Our focus is very much on what we see our users and customers actually doing in the lab,” Larus-Stone said. BenchBench derives its tasks from protocols that scientists physically executed, giving it a specific focus on protocol modification and troubleshooting.
The broader landscape includes LifeSciBench, LABBench2 and BioMysteryBench, which evaluate different slices of scientific work. Larus-Stone described one end of that landscape as “book learning,” with textbook-style or synthetically generated questions. BenchBench reflects “a slightly different view of what biological benchmarks should be, and we think that’s why the models are worse at this than they are at some of those other benchmarks,” he said.
BenchBench tests models in isolation, apart from the data, tools and guardrails surrounding a deployed agent. “The benchmark task specifically has no harness,” Larus-Stone said. “It’s just asking the LLMs directly, and it’s only one type of question.” Actual lab work spans a much wider range of tasks.
Benchling’s product layer supplies the harness: software that connects a model to relevant data and tools, governs its actions and routes its work through scientist review. Combine the model’s biological reasoning with that infrastructure and a scientist who can recognize “what is obviously wrong or which suggestions are good ideas,” Larus-Stone said, and “you can have effective AI in the lab.”
From pharma data to verifiable rewards
Benchling is also working with “a number of the frontier labs” to improve model performance on bench-related tasks. Larus-Stone said those collaborations have shown that carefully constructed scientific data can make models better at work Benchling and its customers care about. Producing that data remains difficult because useful training and evaluation tasks need verifiable answers. Pharma companies’ experimental histories could provide the raw material.
“Pharma companies have lots of data, and the question is how you go from that data to creating verifiable tasks that can improve model performance,” he said. Benchling has begun exploring that approach with a handful of customers.
BenchBench has also produced a potentially important finding for companies deciding how much of their AI stack to own: open models performed across much of the same range as frontier models. “That was a little surprising to us,” Larus-Stone said. He called it “a good sign for pharma companies and others who want to build their own intelligence stack,” given the level of capability already available in some open models.
The next step would be to turn a company’s historical data into verifiable rewards, or checkable signals that reinforce correct model outputs. Companies could use those rewards to construct evaluations and training sets, then post-train open-weight models under their control. Larus-Stone said the resulting systems could “match or beat the frontier models, likely at a fraction of the cost.”

The BenchBench-Protocol found open-weight models like Kimi K3 scoring competitively against models from frontier labs, though Opus 5 maintained a notable lead.
“That trend is just emerging, but it’s going to be an important one over the next six to 12 months,” he added.
For the time being, human oversight is likely to remain especially important in the lab, where Larus-Stone said each incremental automation run can cost tens of thousands of dollars. In the near term, he sees agents serving as a check before scientists commit those resources. “I’m going to first ask the agent to help me validate that this is the right experiment to run,” he said.
The promise is to give scientists a wider view of the experiment, with agents handling routine work such as ordering reagents or determining when a chromatography column needs changing while the scientist focuses on the most promising experimental path. As those workflows multiply, the juggling act could move up a level, from managing lab tasks to managing agents. Software-engineering systems already use lead agents to coordinate subagents, and the same pattern has reached experimental-science prototypes such as PNNL’s AutoLabs. Asked whether managing numerous agents could itself become a bottleneck without better systems for coordinating their work and reconciling their outputs, Larus-Stone replied, “I would call that an unsolved problem we’re working on.”




Tell Us What You Think!
You must be logged in to post a comment.