BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Tian, Gao, Tianrun, Deng, Wenhao, Wei, Long, Qian, Xiaowei, Yu, Chenglei, Wu, Tailin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation Prediction
by: Yu, Chenglei, et al.
Published: (2026)
by: Yu, Chenglei, et al.
Published: (2026)
On the Guidance of Flow Matching
by: Feng, Ruiqi, et al.
Published: (2025)
by: Feng, Ruiqi, et al.
Published: (2025)
GenCP: Towards Generative Modeling Paradigm of Coupled Physics
by: Gao, Tianrun, et al.
Published: (2026)
by: Gao, Tianrun, et al.
Published: (2026)
RealPDEBench: A Benchmark for Complex Physical Systems with Real-World Data
by: Hu, Peiyan, et al.
Published: (2026)
by: Hu, Peiyan, et al.
Published: (2026)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
EqCollide: Equivariant and Collision-Aware Deformable Objects Neural Simulator
by: Chen, Qianyi, et al.
Published: (2025)
by: Chen, Qianyi, et al.
Published: (2025)
PACE: Geometry-Aware Bridge Transport for Single-Cell Trajectory Inference
by: Yu, Chenglei, et al.
Published: (2026)
by: Yu, Chenglei, et al.
Published: (2026)
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
by: Luo, Haipeng, et al.
Published: (2024)
by: Luo, Haipeng, et al.
Published: (2024)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
by: Zhou, Ruiwen, et al.
Published: (2024)
by: Zhou, Ruiwen, et al.
Published: (2024)
FLDmamba: Integrating Fourier and Laplace Transform Decomposition with Mamba for Enhanced Time Series Prediction
by: Zhang, Qianru, et al.
Published: (2025)
by: Zhang, Qianru, et al.
Published: (2025)
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025)
by: Cai, Mengzhang, et al.
Published: (2025)
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
by: Jiang, Lifan, et al.
Published: (2026)
by: Jiang, Lifan, et al.
Published: (2026)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
by: Liu, Peigen, et al.
Published: (2026)
by: Liu, Peigen, et al.
Published: (2026)
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
by: Wu, Jinge, et al.
Published: (2026)
by: Wu, Jinge, et al.
Published: (2026)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
by: Yuan, Jiarui, et al.
Published: (2026)
by: Yuan, Jiarui, et al.
Published: (2026)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
BABE: Biology Arena BEnchmark
by: Zhou, Junting, et al.
Published: (2026)
by: Zhou, Junting, et al.
Published: (2026)
Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents
by: Li, Jinyang, et al.
Published: (2024)
by: Li, Jinyang, et al.
Published: (2024)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
by: Zhang, Danyang, et al.
Published: (2023)
by: Zhang, Danyang, et al.
Published: (2023)
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
by: Wan, Haiyuan, et al.
Published: (2025)
by: Wan, Haiyuan, et al.
Published: (2025)
SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
by: Xia, Wei, et al.
Published: (2025)
by: Xia, Wei, et al.
Published: (2025)
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
by: Jia, Hangyi, et al.
Published: (2025)
by: Jia, Hangyi, et al.
Published: (2025)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
One step further with Monte-Carlo sampler to guide diffusion better
by: Ren, Minsi, et al.
Published: (2026)
by: Ren, Minsi, et al.
Published: (2026)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
by: Myrzakhan, Aidar, et al.
Published: (2024)
by: Myrzakhan, Aidar, et al.
Published: (2024)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
by: Zhu, Yakun, et al.
Published: (2025)
by: Zhu, Yakun, et al.
Published: (2025)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
by: Sun, Zhongxiang, et al.
Published: (2026)
by: Sun, Zhongxiang, et al.
Published: (2026)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
by: Ji, Haonian, et al.
Published: (2026)
by: Ji, Haonian, et al.
Published: (2026)
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
by: Chen, Yimeng, et al.
Published: (2025)
by: Chen, Yimeng, et al.
Published: (2025)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
by: Wang, Daoyu, et al.
Published: (2025)
by: Wang, Daoyu, et al.
Published: (2025)
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
by: Wang, Kevin, et al.
Published: (2026)
by: Wang, Kevin, et al.
Published: (2026)
Similar Items
-
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025) -
scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation Prediction
by: Yu, Chenglei, et al.
Published: (2026) -
On the Guidance of Flow Matching
by: Feng, Ruiqi, et al.
Published: (2025) -
GenCP: Towards Generative Modeling Paradigm of Coupled Physics
by: Gao, Tianrun, et al.
Published: (2026) -
RealPDEBench: A Benchmark for Complex Physical Systems with Real-World Data
by: Hu, Peiyan, et al.
Published: (2026)