SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
Fuente:
arXiv
Saved in:
| Main Authors: | Somasekharan, Nithin, Hassan, Youssef, Lin, Shiyao, Panapitiya, Gihan, Emami, Patrick, Acharya, Anurag, Horawalavithana, Sameera, Pan, Shaowu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Kolmogorov Barrier: A Learnable Weighted Hybrid Autoencoder for Model Order Reduction
by: Somasekharan, Nithin, et al.
Published: (2024)
by: Somasekharan, Nithin, et al.
Published: (2024)
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
by: Somasekharan, Nithin, et al.
Published: (2025)
by: Somasekharan, Nithin, et al.
Published: (2025)
Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery
by: Chintalapati, Renuka, et al.
Published: (2026)
by: Chintalapati, Renuka, et al.
Published: (2026)
Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
by: Yue, Ling, et al.
Published: (2025)
by: Yue, Ling, et al.
Published: (2025)
AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
by: Panapitiya, Gihan, et al.
Published: (2025)
by: Panapitiya, Gihan, et al.
Published: (2025)
UniFoil: A Universal Dataset of Airfoils in Transitional and Turbulent Regimes for Subsonic and Transonic Flows
by: Kanchi, Rohit Sunil, et al.
Published: (2025)
by: Kanchi, Rohit Sunil, et al.
Published: (2025)
Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
by: Chaturvedi, Sarthak, et al.
Published: (2025)
by: Chaturvedi, Sarthak, et al.
Published: (2025)
AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents
by: Somasekharan, Nithin, et al.
Published: (2026)
by: Somasekharan, Nithin, et al.
Published: (2026)
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain
by: Meyur, Rounak, et al.
Published: (2024)
by: Meyur, Rounak, et al.
Published: (2024)
Foam-Agent: Towards Automated Intelligent CFD Workflows
by: Yue, Ling, et al.
Published: (2025)
by: Yue, Ling, et al.
Published: (2025)
Reward Design for Physical Reasoning in Vision-Language Models
by: Lilienthal, Derek, et al.
Published: (2026)
by: Lilienthal, Derek, et al.
Published: (2026)
FragNet: A Graph Neural Network for Molecular Property Prediction with Four Levels of Interpretability
by: Panapitiya, Gihan, et al.
Published: (2024)
by: Panapitiya, Gihan, et al.
Published: (2024)
Benchmarking LLMs for Environmental Review and Permitting
by: Meyur, Rounak, et al.
Published: (2024)
by: Meyur, Rounak, et al.
Published: (2024)
Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA
by: Su, Jinyan, et al.
Published: (2026)
by: Su, Jinyan, et al.
Published: (2026)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
by: Luo, Sichun, et al.
Published: (2025)
by: Luo, Sichun, et al.
Published: (2025)
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
by: Horawalavithana, Sameera, et al.
Published: (2023)
by: Horawalavithana, Sameera, et al.
Published: (2023)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
by: Horawalavithana, Sameera, et al.
Published: (2026)
by: Horawalavithana, Sameera, et al.
Published: (2026)
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
by: Ramezan, Kimia, et al.
Published: (2025)
by: Ramezan, Kimia, et al.
Published: (2025)
Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
by: Yunusov, Sarfaroz, et al.
Published: (2025)
by: Yunusov, Sarfaroz, et al.
Published: (2025)
Development and Formulation of Herbal Products- A Veterinary Herbal Care Perspective
by: Sameera A*
Published: (2026)
by: Sameera A*
Published: (2026)
Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems
by: Sarker, Titu Ranjan, et al.
Published: (2026)
by: Sarker, Titu Ranjan, et al.
Published: (2026)
MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge
by: Cheung, Jerry Junyang, et al.
Published: (2025)
by: Cheung, Jerry Junyang, et al.
Published: (2025)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
by: Fouda, Aya E., et al.
Published: (2025)
by: Fouda, Aya E., et al.
Published: (2025)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
by: He, Yun, et al.
Published: (2024)
by: He, Yun, et al.
Published: (2024)
A Librarian-Teacher Collaboration: Integrating Information Literacy and Technology in the K-12 Classroom
by: Mohamad, Gihan
Published: (2017)
by: Mohamad, Gihan
Published: (2017)
Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing
by: Raab, Reilly, et al.
Published: (2025)
by: Raab, Reilly, et al.
Published: (2025)
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
by: Munikoti, Sai, et al.
Published: (2026)
by: Munikoti, Sai, et al.
Published: (2026)
A Cloud-based Multi-Agentic Workflow for Science
by: Acharya, Anurag, et al.
Published: (2026)
by: Acharya, Anurag, et al.
Published: (2026)
Phytoextracted Formulations for Periodontitis: Formulation Strategies, Therapeutic Mechanisms, and Clinical Applications
by: Anuj Kumar, et al.
Published: (2026)
by: Anuj Kumar, et al.
Published: (2026)
Instantaneous Complex Phase and Frequency: Conceptual Clarification and Equivalence between Formulations
by: García-Veloso, César, et al.
Published: (2025)
by: García-Veloso, César, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities
by: Munikoti, Sai, et al.
Published: (2024)
by: Munikoti, Sai, et al.
Published: (2024)
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
by: Cheng, Xiang, et al.
Published: (2025)
by: Cheng, Xiang, et al.
Published: (2025)
BuildingsBench: A Large-Scale Dataset of 900K Buildings and Benchmark for Short-Term Load Forecasting
by: Emami, Patrick, et al.
Published: (2023)
by: Emami, Patrick, et al.
Published: (2023)
On the lifting and reconstruction of nonlinear systems with multiple invariant sets
by: Pan, Shaowu, et al.
Published: (2023)
by: Pan, Shaowu, et al.
Published: (2023)
Robust-Multi-Task Gradient Boosting
by: Emami, Seyedsaman, et al.
Published: (2025)
by: Emami, Seyedsaman, et al.
Published: (2025)
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
by: Ajjour, Yamen, et al.
Published: (2026)
by: Ajjour, Yamen, et al.
Published: (2026)
Similar Items
-
Beyond the Kolmogorov Barrier: A Learnable Weighted Hybrid Autoencoder for Model Order Reduction
by: Somasekharan, Nithin, et al.
Published: (2024) -
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
by: Somasekharan, Nithin, et al.
Published: (2025) -
Evaluating Memory Condensation Strategies for Coding Agents in Data-Driven Scientific Discovery
by: Chintalapati, Renuka, et al.
Published: (2026) -
Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
by: Yue, Ling, et al.
Published: (2025) -
AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
by: Panapitiya, Gihan, et al.
Published: (2025)