Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guan, Zihan, Datta, Rituparna, Hu, Mengxuan, Liu, Shunshun, Zhang, Aiying, Balachandran, Prasanna, Li, Sheng, Vullikanti, Anil
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913175427874816
author Guan, Zihan
Datta, Rituparna
Hu, Mengxuan
Liu, Shunshun
Zhang, Aiying
Balachandran, Prasanna
Li, Sheng
Vullikanti, Anil
author_facet Guan, Zihan
Datta, Rituparna
Hu, Mengxuan
Liu, Shunshun
Zhang, Aiying
Balachandran, Prasanna
Li, Sheng
Vullikanti, Anil
contents Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on simplified settings and fail to capture the complexity of real-world scientific modeling. In practice, such modeling often involves neural-integrated formulations, where a mechanistic model component and a neural network component are jointly constructed, leading to a significantly more complex search space. Motivated by this gap, we introduce the Neural-Integrated Mechanistic Modeling (NIMM) benchmark, which evaluates LLM-generated neural-integrated mechanistic models across three scientific domains. Experiments on NIMM reveal that existing LLM-based approaches struggle to effectively explore this complex space, resulting in limited search stability and solution quality. To address this challenge, we propose NIMMGen, a tree-guided agentic framework that enables diversified exploration via branch-level search and improves solutions through atomic model refinement. Extensive experiments demonstrate that NIMMGen achieves state-of-the-art performance on NIMM, significantly improving search stability and solution quality.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18008
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
Guan, Zihan
Datta, Rituparna
Hu, Mengxuan
Liu, Shunshun
Zhang, Aiying
Balachandran, Prasanna
Li, Sheng
Vullikanti, Anil
Machine Learning
Artificial Intelligence
Computation and Language
Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on simplified settings and fail to capture the complexity of real-world scientific modeling. In practice, such modeling often involves neural-integrated formulations, where a mechanistic model component and a neural network component are jointly constructed, leading to a significantly more complex search space. Motivated by this gap, we introduce the Neural-Integrated Mechanistic Modeling (NIMM) benchmark, which evaluates LLM-generated neural-integrated mechanistic models across three scientific domains. Experiments on NIMM reveal that existing LLM-based approaches struggle to effectively explore this complex space, resulting in limited search stability and solution quality. To address this challenge, we propose NIMMGen, a tree-guided agentic framework that enables diversified exploration via branch-level search and improves solutions through atomic model refinement. Extensive experiments demonstrate that NIMMGen achieves state-of-the-art performance on NIMM, significantly improving search stability and solution quality.
title Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.18008