PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Fuente:
arXiv
Saved in:
| Main Authors: | Miao, Tingjia, Jin, Wenkai, Zhang, Muhua, Tan, Jinxin, Hu, Yuelin, Guo, Tu, Zhang, Jiejun, Wang, Yuhan, Li, Wenbo, Gao, Yinuo, Chen, Shuo, Jiang, Weiqi, Hu, Yayun, Lei, Zixing, Pang, Xianghe, Liu, Zexi, Zhang, Yuzhi, Zhang, Linfeng, Chen, Kun, Wang, Wei, E, Weinan, Chen, Siheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
by: Miao, Tingjia, et al.
Published: (2025)
by: Miao, Tingjia, et al.
Published: (2025)
Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
Towards Self-Evolving Agentic Literature Retrieval
by: Du, Yuwen, et al.
Published: (2026)
by: Du, Yuwen, et al.
Published: (2026)
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
by: Liu, Yuzhi, et al.
Published: (2026)
by: Liu, Yuzhi, et al.
Published: (2026)
SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
by: Chai, Jingyi, et al.
Published: (2025)
by: Chai, Jingyi, et al.
Published: (2025)
Pragmatic Communication in Multi-Agent Collaborative Perception
by: Hu, Yue, et al.
Published: (2024)
by: Hu, Yue, et al.
Published: (2024)
DiPRL: Learning Discrete Programmatic Policies via Architecture Entropy Regularization
by: Hu, Chengpeng, et al.
Published: (2026)
by: Hu, Chengpeng, et al.
Published: (2026)
Uni-ELF: A Multi-Level Representation Learning Framework for Electrolyte Formulation Design
by: Zeng, Boshen, et al.
Published: (2024)
by: Zeng, Boshen, et al.
Published: (2024)
Selective Attention in Early Word Learning: An Eye‐Tracking Study on Viewing Naturalistic Egocentric Scenes
by: Yayun Zhang, et al.
Published: (2025)
by: Yayun Zhang, et al.
Published: (2025)
Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
by: Pang, Xianghe, et al.
Published: (2024)
by: Pang, Xianghe, et al.
Published: (2024)
Machine-Learning-Based Interatomic Potentials for Group IIB to VIA Semiconductors: Towards a Universal Model
by: Liu, Jianchuan, et al.
Published: (2023)
by: Liu, Jianchuan, et al.
Published: (2023)
Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving
by: Ren, Shunli, et al.
Published: (2023)
by: Ren, Shunli, et al.
Published: (2023)
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
by: Du, Yaxin, et al.
Published: (2025)
by: Du, Yaxin, et al.
Published: (2025)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
by: Tang, Shuo, et al.
Published: (2024)
by: Tang, Shuo, et al.
Published: (2024)
DataMaster: Data-Centric Autonomous AI Research
by: Du, Yaxin, et al.
Published: (2026)
by: Du, Yaxin, et al.
Published: (2026)
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
by: Chen, Menglan, et al.
Published: (2025)
by: Chen, Menglan, et al.
Published: (2025)
Rate-Distortion Optimized Communication for Collaborative Perception
by: Liu, Genjia, et al.
Published: (2025)
by: Liu, Genjia, et al.
Published: (2025)
Temporal Knowledge Graph Reasoning Based on Dynamic Fusion Representation Learning
by: Hongwei Chen, et al.
Published: (2024)
by: Hongwei Chen, et al.
Published: (2024)
Malicious Agent Detection for Robust Multi-Agent Collaborative Perception
by: Zhao, Yangheng, et al.
Published: (2023)
by: Zhao, Yangheng, et al.
Published: (2023)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025)
by: Cai, Muzhen, et al.
Published: (2025)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024)
by: Yin, Sheng, et al.
Published: (2024)
Research on Feature Extraction Data Processing System For MRI of Brain Diseases Based on Computer Deep Learning
by: Xiao, Lingxi, et al.
Published: (2024)
by: Xiao, Lingxi, et al.
Published: (2024)
Self-Localized Collaborative Perception
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
Seeking Meaning: Incorporating Linguistic Information in Cross‐Situational Verb Learning
by: Chi‐hsin Chen, et al.
Published: (2025)
by: Chi‐hsin Chen, et al.
Published: (2025)
AgenticRS-EnsNAS: Ensemble-Decoupled Self-Evolving Architecture Search
by: Chen, Yun, et al.
Published: (2026)
by: Chen, Yun, et al.
Published: (2026)
DGenCTR: Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion
by: Zhang, Moyu, et al.
Published: (2025)
by: Zhang, Moyu, et al.
Published: (2025)
37‐2: Invited Paper: Intermolecular Charge Transfer for High‐Performance Organic Light‐Emitting Diodes
by: Zhen Zhang, et al.
Published: (2024)
by: Zhen Zhang, et al.
Published: (2024)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
by: Cai, Yuzhu, et al.
Published: (2026)
by: Cai, Yuzhu, et al.
Published: (2026)
SketchPlay: Intuitive Creation of Physically Realistic VR Content with Gesture-Driven Sketching
by: Zhang, Xiangwen, et al.
Published: (2025)
by: Zhang, Xiangwen, et al.
Published: (2025)
A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application
by: Chen, Shuaihang, et al.
Published: (2024)
by: Chen, Shuaihang, et al.
Published: (2024)
Computational Multi-Agents Society Experiments: Social Modeling Framework Based on Generative Agents
by: Zhang, Hanzhong, et al.
Published: (2025)
by: Zhang, Hanzhong, et al.
Published: (2025)
Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
by: Zhang, Junyi, et al.
Published: (2025)
by: Zhang, Junyi, et al.
Published: (2025)
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
by: Shi, Liangzhi, et al.
Published: (2026)
by: Shi, Liangzhi, et al.
Published: (2026)
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation
by: Zhou, Zibo, et al.
Published: (2025)
by: Zhou, Zibo, et al.
Published: (2025)
Hypergraph Transformer for Semi-Supervised Classification
by: Liu, Zexi, et al.
Published: (2023)
by: Liu, Zexi, et al.
Published: (2023)
Training-Free Message Passing for Learning on Hypergraphs
by: Tang, Bohan, et al.
Published: (2024)
by: Tang, Bohan, et al.
Published: (2024)
Similar Items
-
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
by: Miao, Tingjia, et al.
Published: (2025) -
Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
by: Zhu, Xinyu, et al.
Published: (2026) -
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
by: Zhu, Xinyu, et al.
Published: (2026) -
Towards Self-Evolving Agentic Literature Retrieval
by: Du, Yuwen, et al.
Published: (2026) -
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
by: Lei, Zixing, et al.
Published: (2026)