EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Mehta, Sushant, Ritchie, Logan, Garre, Suhaas, Niebres, Ian, Heiner, Nick, Chen, Edwin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
by: Ritchie, Logan, et al.
Published: (2026)
by: Ritchie, Logan, et al.
Published: (2026)
Riemann-Bench: A Benchmark for Moonshot Mathematics
by: Garre, Suhaas, et al.
Published: (2026)
by: Garre, Suhaas, et al.
Published: (2026)
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
by: Mehta, Sushant
Published: (2025)
by: Mehta, Sushant
Published: (2025)
MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations
by: Mehta, Sushant
Published: (2026)
by: Mehta, Sushant
Published: (2026)
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
by: Mehta, Sushant
Published: (2025)
by: Mehta, Sushant
Published: (2025)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift
by: Mehta, Sushant
Published: (2025)
by: Mehta, Sushant
Published: (2025)
HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems
by: Mehta, Sushant
Published: (2026)
by: Mehta, Sushant
Published: (2026)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
by: Radmehr, Bahar, et al.
Published: (2024)
by: Radmehr, Bahar, et al.
Published: (2024)
TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents
by: Zhu, Kaijie, et al.
Published: (2026)
by: Zhu, Kaijie, et al.
Published: (2026)
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
Scalable Environments Drive Generalizable Agents
by: Zhang, Jiayi, et al.
Published: (2026)
by: Zhang, Jiayi, et al.
Published: (2026)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
by: Mo, Ying, et al.
Published: (2026)
by: Mo, Ying, et al.
Published: (2026)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
by: Singla, Pratham, et al.
Published: (2025)
by: Singla, Pratham, et al.
Published: (2025)
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
by: Chandwani, Abhishek, et al.
Published: (2026)
by: Chandwani, Abhishek, et al.
Published: (2026)
AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
Quantifying the Privacy Implications of High-Fidelity Synthetic Network Traffic
by: Tran, Van, et al.
Published: (2025)
by: Tran, Van, et al.
Published: (2025)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
by: Shi, Wentao, et al.
Published: (2026)
by: Shi, Wentao, et al.
Published: (2026)
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
by: Liu, Ruoqi, et al.
Published: (2026)
by: Liu, Ruoqi, et al.
Published: (2026)
AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
by: Kleczek, Darek, et al.
Published: (2026)
by: Kleczek, Darek, et al.
Published: (2026)
JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
by: Rutherford, Alexander, et al.
Published: (2023)
by: Rutherford, Alexander, et al.
Published: (2023)
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
by: Han, Yi, et al.
Published: (2026)
by: Han, Yi, et al.
Published: (2026)
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
by: Gao, Shenyuan, et al.
Published: (2024)
by: Gao, Shenyuan, et al.
Published: (2024)
From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation
by: Lu, Jasper, et al.
Published: (2026)
by: Lu, Jasper, et al.
Published: (2026)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
by: Li, Kuan, et al.
Published: (2026)
by: Li, Kuan, et al.
Published: (2026)
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
by: Bhattarai, Manish, et al.
Published: (2026)
by: Bhattarai, Manish, et al.
Published: (2026)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
by: Li, Wanli, et al.
Published: (2026)
by: Li, Wanli, et al.
Published: (2026)
CCrepairBench: A High-Fidelity Benchmark and Reinforcement Learning Framework for C++ Compilation Repair
by: Sun, Weixuan, et al.
Published: (2025)
by: Sun, Weixuan, et al.
Published: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Dingtalk DeepResearch: A Unified Multi Agent Framework for Adaptive Intelligence in Enterprise Environments
by: Chen, Mengyuan, et al.
Published: (2025)
by: Chen, Mengyuan, et al.
Published: (2025)
When Agents Disagree With Themselves: Measuring Behavioral Consistency in LLM-Based Agents
by: Mehta, Aman
Published: (2026)
by: Mehta, Aman
Published: (2026)
CanaryBench: Stress Testing Privacy Leakage in Cluster-Level Conversation Summaries
by: Mehta, Deep
Published: (2026)
by: Mehta, Deep
Published: (2026)
TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control
by: Lai, Siqi, et al.
Published: (2026)
by: Lai, Siqi, et al.
Published: (2026)
Similar Items
-
The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
by: Ritchie, Logan, et al.
Published: (2026) -
Riemann-Bench: A Benchmark for Moonshot Mathematics
by: Garre, Suhaas, et al.
Published: (2026) -
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
by: Mehta, Sushant
Published: (2025) -
MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations
by: Mehta, Sushant
Published: (2026) -
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
by: Mehta, Sushant
Published: (2025)