The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ritchie, Logan, Mehta, Sushant, Heiner, Nick, Yu, Mason, Chen, Edwin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
par: Mehta, Sushant, et autres
Publié: (2026)
par: Mehta, Sushant, et autres
Publié: (2026)
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
par: Mehta, Sushant
Publié: (2025)
par: Mehta, Sushant
Publié: (2025)
HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems
par: Mehta, Sushant
Publié: (2026)
par: Mehta, Sushant
Publié: (2026)
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
par: Mehta, Sushant
Publié: (2025)
par: Mehta, Sushant
Publié: (2025)
Riemann-Bench: A Benchmark for Moonshot Mathematics
par: Garre, Suhaas, et autres
Publié: (2026)
par: Garre, Suhaas, et autres
Publié: (2026)
MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations
par: Mehta, Sushant
Publié: (2026)
par: Mehta, Sushant
Publié: (2026)
When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift
par: Mehta, Sushant
Publié: (2025)
par: Mehta, Sushant
Publié: (2025)
Automatic Environment Shaping is the Next Frontier in RL
par: Park, Younghyo, et autres
Publié: (2024)
par: Park, Younghyo, et autres
Publié: (2024)
Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
par: Sharma, Raghav, et autres
Publié: (2025)
par: Sharma, Raghav, et autres
Publié: (2025)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
par: Mehta, Sushant, et autres
Publié: (2025)
par: Mehta, Sushant, et autres
Publié: (2025)
Open-World Evaluations for Measuring Frontier AI Capabilities
par: Kapoor, Sayash, et autres
Publié: (2026)
par: Kapoor, Sayash, et autres
Publié: (2026)
Reproducibility: The New Frontier in AI Governance
par: Mason-Williams, Israel, et autres
Publié: (2025)
par: Mason-Williams, Israel, et autres
Publié: (2025)
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
par: Mehta, Sushant, et autres
Publié: (2025)
par: Mehta, Sushant, et autres
Publié: (2025)
Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier
par: Henry, Jazmia
Publié: (2026)
par: Henry, Jazmia
Publié: (2026)
Frontier Models are Capable of In-context Scheming
par: Meinke, Alexander, et autres
Publié: (2024)
par: Meinke, Alexander, et autres
Publié: (2024)
Echo-N1: Affective RL Frontier
par: Zhang, Naifan, et autres
Publié: (2025)
par: Zhang, Naifan, et autres
Publié: (2025)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
par: Zhu, Pengyu, et autres
Publié: (2026)
par: Zhu, Pengyu, et autres
Publié: (2026)
CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
par: Mai, Shinji, et autres
Publié: (2025)
par: Mai, Shinji, et autres
Publié: (2025)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
par: Wei, Qianshan, et autres
Publié: (2026)
par: Wei, Qianshan, et autres
Publié: (2026)
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
par: Chu, Meng, et autres
Publié: (2026)
par: Chu, Meng, et autres
Publié: (2026)
Forecasting Frontier Language Model Agent Capabilities
par: Pimpale, Govind, et autres
Publié: (2025)
par: Pimpale, Govind, et autres
Publié: (2025)
Latent Multi-Head Attention for Small Language Models
par: Mehta, Sushant, et autres
Publié: (2025)
par: Mehta, Sushant, et autres
Publié: (2025)
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
par: McGinness, Lachlan, et autres
Publié: (2025)
par: McGinness, Lachlan, et autres
Publié: (2025)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
par: Sirdeshmukh, Ved, et autres
Publié: (2025)
par: Sirdeshmukh, Ved, et autres
Publié: (2025)
Jailbroken Frontier Models Retain Their Capabilities
par: Zhu, Daniel, et autres
Publié: (2026)
par: Zhu, Daniel, et autres
Publié: (2026)
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
par: Comanici, Gheorghe, et autres
Publié: (2025)
par: Comanici, Gheorghe, et autres
Publié: (2025)
The Missing Memory Hierarchy: Demand Paging for LLM Context Windows
par: Mason, Tony
Publié: (2026)
par: Mason, Tony
Publié: (2026)
Skill Reuse as Compression in Agentic RL
par: Xu, Zhikun, et autres
Publié: (2026)
par: Xu, Zhikun, et autres
Publié: (2026)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
par: Pignatelli, Eduardo, et autres
Publié: (2024)
par: Pignatelli, Eduardo, et autres
Publié: (2024)
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
par: Zhang, Tao, et autres
Publié: (2025)
par: Zhang, Tao, et autres
Publié: (2025)
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
par: Zhuang, Tianyi, et autres
Publié: (2025)
par: Zhuang, Tianyi, et autres
Publié: (2025)
PyVision-RL: Forging Open Agentic Vision Models via RL
par: Zhao, Shitian, et autres
Publié: (2026)
par: Zhao, Shitian, et autres
Publié: (2026)
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
par: Chi, Banghao, et autres
Publié: (2026)
par: Chi, Banghao, et autres
Publié: (2026)
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
par: Dong, Yihong, et autres
Publié: (2026)
par: Dong, Yihong, et autres
Publié: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
par: Chen, Wanyi, et autres
Publié: (2026)
par: Chen, Wanyi, et autres
Publié: (2026)
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
par: Hu, Tianyi, et autres
Publié: (2026)
par: Hu, Tianyi, et autres
Publié: (2026)
TRACE: Capability-Targeted Agentic Training
par: Kang, Hangoo, et autres
Publié: (2026)
par: Kang, Hangoo, et autres
Publié: (2026)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
par: Zhang, Jiazheng, et autres
Publié: (2026)
par: Zhang, Jiazheng, et autres
Publié: (2026)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
par: Khoo, Shaun, et autres
Publié: (2025)
par: Khoo, Shaun, et autres
Publié: (2025)
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
par: Shi, Yucheng, et autres
Publié: (2026)
par: Shi, Yucheng, et autres
Publié: (2026)
Documents similaires
-
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
par: Mehta, Sushant, et autres
Publié: (2026) -
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
par: Mehta, Sushant
Publié: (2025) -
HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems
par: Mehta, Sushant
Publié: (2026) -
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
par: Mehta, Sushant
Publié: (2025) -
Riemann-Bench: A Benchmark for Moonshot Mathematics
par: Garre, Suhaas, et autres
Publié: (2026)