MDGYM: Benchmarking AI Agents on Molecular Simulations
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumar, Vinay, Rajput, Satyendra, Mausam, Krishnan, N. M. Anoop |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
por: Bisht, Harshit, et al.
Publicado: (2026)
por: Bisht, Harshit, et al.
Publicado: (2026)
Leveraging Text Guidance for Enhancing Demographic Fairness in Gender Classification
por: Krishnan, Anoop
Publicado: (2025)
por: Krishnan, Anoop
Publicado: (2025)
Are LLMs Ready for Real-World Materials Discovery?
por: Miret, Santiago, et al.
Publicado: (2024)
por: Miret, Santiago, et al.
Publicado: (2024)
Beyond Context Sharing: A Unified Agent Communication Protocol (ACP) for Secure, Federated, and Autonomous Agent-to-Agent (A2A) Orchestration
por: Krishnan, Naveen Kumar
Publicado: (2026)
por: Krishnan, Naveen Kumar
Publicado: (2026)
A Benchmark for Procedural Memory Retrieval in Language Agents
por: Kohar, Ishant, et al.
Publicado: (2025)
por: Kohar, Ishant, et al.
Publicado: (2025)
MarketBench: Evaluating AI Agents as Market Participants
por: Fradkin, Andrey, et al.
Publicado: (2026)
por: Fradkin, Andrey, et al.
Publicado: (2026)
Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
por: Vinay, Vaishali
Publicado: (2025)
por: Vinay, Vaishali
Publicado: (2025)
Efficient Benchmarking of AI Agents
por: Ndzomga, Franck
Publicado: (2026)
por: Ndzomga, Franck
Publicado: (2026)
AxOMaP: Designing FPGA-based Approximate Arithmetic Operators using Mathematical Programming
por: Sahoo, Siva Satyendra, et al.
Publicado: (2023)
por: Sahoo, Siva Satyendra, et al.
Publicado: (2023)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
por: Wang, Jianghui, et al.
Publicado: (2025)
por: Wang, Jianghui, et al.
Publicado: (2025)
BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
por: Metcalf, Sara, et al.
Publicado: (2026)
por: Metcalf, Sara, et al.
Publicado: (2026)
PLACO: A Multi-Stage Framework for Cost-Effective Performance in Human-AI Teams
por: Mallela, Pranavkumar, et al.
Publicado: (2026)
por: Mallela, Pranavkumar, et al.
Publicado: (2026)
CoNO: Complex Neural Operator for Continous Dynamical Physical Systems
por: Tiwari, Karn, et al.
Publicado: (2024)
por: Tiwari, Karn, et al.
Publicado: (2024)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
por: Kim, Myeongsoo, et al.
Publicado: (2025)
por: Kim, Myeongsoo, et al.
Publicado: (2025)
Iterative Repair with Weak Verifiers for Few-shot Transfer in KBQA with Unanswerability
por: Sawhney, Riya, et al.
Publicado: (2024)
por: Sawhney, Riya, et al.
Publicado: (2024)
GTM: Simulating the World of Tools for AI Agents
por: Ren, Zhenzhen, et al.
Publicado: (2025)
por: Ren, Zhenzhen, et al.
Publicado: (2025)
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
por: Garg, Divyansh, et al.
Publicado: (2025)
por: Garg, Divyansh, et al.
Publicado: (2025)
Tu(r)ning AI Green: Exploring Energy Efficiency Cascading with Orthogonal Optimizations
por: Rajput, Saurabhsingh, et al.
Publicado: (2025)
por: Rajput, Saurabhsingh, et al.
Publicado: (2025)
OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
por: Davydova, Mariya, et al.
Publicado: (2025)
por: Davydova, Mariya, et al.
Publicado: (2025)
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
por: He, Muyu, et al.
Publicado: (2025)
por: He, Muyu, et al.
Publicado: (2025)
Sustainable Materials Discovery in the Era of Artificial Intelligence
por: Mannan, Sajid, et al.
Publicado: (2026)
por: Mannan, Sajid, et al.
Publicado: (2026)
AI scientists produce results without reasoning scientifically
por: Ríos-García, Martiño, et al.
Publicado: (2026)
por: Ríos-García, Martiño, et al.
Publicado: (2026)
Autonomous Microscopy Experiments through Large Language Model Agents
por: Mandal, Indrajeet, et al.
Publicado: (2024)
por: Mandal, Indrajeet, et al.
Publicado: (2024)
MatSKRAFT: A framework for large-scale materials knowledge extraction from scientific tables
por: Hira, Kausik, et al.
Publicado: (2025)
por: Hira, Kausik, et al.
Publicado: (2025)
Advancing Multi-Agent Systems Through Model Context Protocol: Architecture, Implementation, and Applications
por: Krishnan, Naveen
Publicado: (2025)
por: Krishnan, Naveen
Publicado: (2025)
Beyond Component Strength: Synergistic Integration and Adaptive Calibration in Multi-Agent RAG Systems
por: Krishnan, Jithin
Publicado: (2025)
por: Krishnan, Jithin
Publicado: (2025)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
por: Ebrahimi, Amir M., et al.
Publicado: (2026)
por: Ebrahimi, Amir M., et al.
Publicado: (2026)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
por: Li, Miles Q., et al.
Publicado: (2026)
por: Li, Miles Q., et al.
Publicado: (2026)
On the Equivalence of Regression and Classification
por: Jayadeva, et al.
Publicado: (2025)
por: Jayadeva, et al.
Publicado: (2025)
A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
por: Li, Miles Q., et al.
Publicado: (2025)
por: Li, Miles Q., et al.
Publicado: (2025)
Energy & Force Regression on DFT Trajectories is Not Enough for Universal Machine Learning Interatomic Potentials
por: Miret, Santiago, et al.
Publicado: (2025)
por: Miret, Santiago, et al.
Publicado: (2025)
Simple Augmentations of Logical Rules for Neuro-Symbolic Knowledge Graph Completion
por: Nandi, Ananjan, et al.
Publicado: (2024)
por: Nandi, Ananjan, et al.
Publicado: (2024)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
por: Mittal, Chinmay, et al.
Publicado: (2024)
por: Mittal, Chinmay, et al.
Publicado: (2024)
Agent-based Simulation with Netlogo to Evaluate AmI Scenarios
por: Carbo, J., et al.
Publicado: (2024)
por: Carbo, J., et al.
Publicado: (2024)
AI Agents: Evolution, Architecture, and Real-World Applications
por: Krishnan, Naveen
Publicado: (2025)
por: Krishnan, Naveen
Publicado: (2025)
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
por: Mantravadi, Ananya, et al.
Publicado: (2026)
por: Mantravadi, Ananya, et al.
Publicado: (2026)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
por: He, Muyu, et al.
Publicado: (2026)
por: He, Muyu, et al.
Publicado: (2026)
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
por: Liu, Xuan, et al.
Publicado: (2026)
por: Liu, Xuan, et al.
Publicado: (2026)
From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
por: Wang, Yuhang, et al.
Publicado: (2026)
por: Wang, Yuhang, et al.
Publicado: (2026)
Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands
por: Seth, Pratinav, et al.
Publicado: (2026)
por: Seth, Pratinav, et al.
Publicado: (2026)
Ejemplares similares
-
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
por: Bisht, Harshit, et al.
Publicado: (2026) -
Leveraging Text Guidance for Enhancing Demographic Fairness in Gender Classification
por: Krishnan, Anoop
Publicado: (2025) -
Are LLMs Ready for Real-World Materials Discovery?
por: Miret, Santiago, et al.
Publicado: (2024) -
Beyond Context Sharing: A Unified Agent Communication Protocol (ACP) for Secure, Federated, and Autonomous Agent-to-Agent (A2A) Orchestration
por: Krishnan, Naveen Kumar
Publicado: (2026) -
A Benchmark for Procedural Memory Retrieval in Language Agents
por: Kohar, Ishant, et al.
Publicado: (2025)