Saved in:
Bibliographic Details
Main Authors: Gupta, Rushil, Hartford, Jason, Liu, Bang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.21403
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916970674257920
author Gupta, Rushil
Hartford, Jason
Liu, Bang
author_facet Gupta, Rushil
Hartford, Jason
Liu, Bang
contents Large language models (LLMs) have recently been proposed as general-purpose agents for experimental design, with claims that they can perform in-context experimental design. We evaluate this hypothesis using both open- and closed-source instruction-tuned LLMs applied to genetic perturbation and molecular property discovery tasks. We find that LLM-based agents show no sensitivity to experimental feedback: replacing true outcomes with randomly permuted labels has no impact on performance. Across benchmarks, classical methods such as linear bandits and Gaussian process optimization consistently outperform LLM agents. We further propose a simple hybrid method, LLM-guided Nearest Neighbour (LLMNN) sampling, that combines LLM prior knowledge with nearest-neighbor sampling to guide the design of experiments. LLMNN achieves competitive or superior performance across domains without requiring significant in-context adaptation. These results suggest that current open- and closed-source LLMs do not perform in-context experimental design in practice and highlight the need for hybrid frameworks that decouple prior-based reasoning from batch acquisition with updated posteriors.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21403
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
Gupta, Rushil
Hartford, Jason
Liu, Bang
Machine Learning
Computation and Language
Large language models (LLMs) have recently been proposed as general-purpose agents for experimental design, with claims that they can perform in-context experimental design. We evaluate this hypothesis using both open- and closed-source instruction-tuned LLMs applied to genetic perturbation and molecular property discovery tasks. We find that LLM-based agents show no sensitivity to experimental feedback: replacing true outcomes with randomly permuted labels has no impact on performance. Across benchmarks, classical methods such as linear bandits and Gaussian process optimization consistently outperform LLM agents. We further propose a simple hybrid method, LLM-guided Nearest Neighbour (LLMNN) sampling, that combines LLM prior knowledge with nearest-neighbor sampling to guide the design of experiments. LLMNN achieves competitive or superior performance across domains without requiring significant in-context adaptation. These results suggest that current open- and closed-source LLMs do not perform in-context experimental design in practice and highlight the need for hybrid frameworks that decouple prior-based reasoning from batch acquisition with updated posteriors.
title LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.21403