The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
Fuente:
arXiv
Guardado en:
| Autores principales: | Ritchie, Logan, Mehta, Sushant, Heiner, Nick, Yu, Mason, Chen, Edwin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
por: Mehta, Sushant, et al.
Publicado: (2026)
por: Mehta, Sushant, et al.
Publicado: (2026)
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
por: Mehta, Sushant
Publicado: (2025)
por: Mehta, Sushant
Publicado: (2025)
HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems
por: Mehta, Sushant
Publicado: (2026)
por: Mehta, Sushant
Publicado: (2026)
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
por: Mehta, Sushant
Publicado: (2025)
por: Mehta, Sushant
Publicado: (2025)
Riemann-Bench: A Benchmark for Moonshot Mathematics
por: Garre, Suhaas, et al.
Publicado: (2026)
por: Garre, Suhaas, et al.
Publicado: (2026)
MATRAG: Multi-Agent Transparent Retrieval-Augmented Generation for Explainable Recommendations
por: Mehta, Sushant
Publicado: (2026)
por: Mehta, Sushant
Publicado: (2026)
When Are Learning Biases Equivalent? A Unifying Framework for Fairness, Robustness, and Distribution Shift
por: Mehta, Sushant
Publicado: (2025)
por: Mehta, Sushant
Publicado: (2025)
Automatic Environment Shaping is the Next Frontier in RL
por: Park, Younghyo, et al.
Publicado: (2024)
por: Park, Younghyo, et al.
Publicado: (2024)
Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
por: Sharma, Raghav, et al.
Publicado: (2025)
por: Sharma, Raghav, et al.
Publicado: (2025)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
por: Mehta, Sushant, et al.
Publicado: (2025)
por: Mehta, Sushant, et al.
Publicado: (2025)
Open-World Evaluations for Measuring Frontier AI Capabilities
por: Kapoor, Sayash, et al.
Publicado: (2026)
por: Kapoor, Sayash, et al.
Publicado: (2026)
Reproducibility: The New Frontier in AI Governance
por: Mason-Williams, Israel, et al.
Publicado: (2025)
por: Mason-Williams, Israel, et al.
Publicado: (2025)
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
por: Mehta, Sushant, et al.
Publicado: (2025)
por: Mehta, Sushant, et al.
Publicado: (2025)
Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier
por: Henry, Jazmia
Publicado: (2026)
por: Henry, Jazmia
Publicado: (2026)
Frontier Models are Capable of In-context Scheming
por: Meinke, Alexander, et al.
Publicado: (2024)
por: Meinke, Alexander, et al.
Publicado: (2024)
Echo-N1: Affective RL Frontier
por: Zhang, Naifan, et al.
Publicado: (2025)
por: Zhang, Naifan, et al.
Publicado: (2025)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
por: Zhu, Pengyu, et al.
Publicado: (2026)
por: Zhu, Pengyu, et al.
Publicado: (2026)
CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
por: Mai, Shinji, et al.
Publicado: (2025)
por: Mai, Shinji, et al.
Publicado: (2025)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
por: Wei, Qianshan, et al.
Publicado: (2026)
por: Wei, Qianshan, et al.
Publicado: (2026)
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
por: Chu, Meng, et al.
Publicado: (2026)
por: Chu, Meng, et al.
Publicado: (2026)
Forecasting Frontier Language Model Agent Capabilities
por: Pimpale, Govind, et al.
Publicado: (2025)
por: Pimpale, Govind, et al.
Publicado: (2025)
Latent Multi-Head Attention for Small Language Models
por: Mehta, Sushant, et al.
Publicado: (2025)
por: Mehta, Sushant, et al.
Publicado: (2025)
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
por: McGinness, Lachlan, et al.
Publicado: (2025)
por: McGinness, Lachlan, et al.
Publicado: (2025)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
por: Sirdeshmukh, Ved, et al.
Publicado: (2025)
por: Sirdeshmukh, Ved, et al.
Publicado: (2025)
Jailbroken Frontier Models Retain Their Capabilities
por: Zhu, Daniel, et al.
Publicado: (2026)
por: Zhu, Daniel, et al.
Publicado: (2026)
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
por: Comanici, Gheorghe, et al.
Publicado: (2025)
por: Comanici, Gheorghe, et al.
Publicado: (2025)
The Missing Memory Hierarchy: Demand Paging for LLM Context Windows
por: Mason, Tony
Publicado: (2026)
por: Mason, Tony
Publicado: (2026)
Skill Reuse as Compression in Agentic RL
por: Xu, Zhikun, et al.
Publicado: (2026)
por: Xu, Zhikun, et al.
Publicado: (2026)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
por: Zhang, Tao, et al.
Publicado: (2025)
por: Zhang, Tao, et al.
Publicado: (2025)
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
por: Zhuang, Tianyi, et al.
Publicado: (2025)
por: Zhuang, Tianyi, et al.
Publicado: (2025)
PyVision-RL: Forging Open Agentic Vision Models via RL
por: Zhao, Shitian, et al.
Publicado: (2026)
por: Zhao, Shitian, et al.
Publicado: (2026)
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
por: Chi, Banghao, et al.
Publicado: (2026)
por: Chi, Banghao, et al.
Publicado: (2026)
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
por: Dong, Yihong, et al.
Publicado: (2026)
por: Dong, Yihong, et al.
Publicado: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
por: Chen, Wanyi, et al.
Publicado: (2026)
por: Chen, Wanyi, et al.
Publicado: (2026)
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
por: Hu, Tianyi, et al.
Publicado: (2026)
por: Hu, Tianyi, et al.
Publicado: (2026)
TRACE: Capability-Targeted Agentic Training
por: Kang, Hangoo, et al.
Publicado: (2026)
por: Kang, Hangoo, et al.
Publicado: (2026)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
por: Zhang, Jiazheng, et al.
Publicado: (2026)
por: Zhang, Jiazheng, et al.
Publicado: (2026)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
por: Khoo, Shaun, et al.
Publicado: (2025)
por: Khoo, Shaun, et al.
Publicado: (2025)
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
por: Shi, Yucheng, et al.
Publicado: (2026)
por: Shi, Yucheng, et al.
Publicado: (2026)
Ejemplares similares
-
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
por: Mehta, Sushant, et al.
Publicado: (2026) -
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
por: Mehta, Sushant
Publicado: (2025) -
HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems
por: Mehta, Sushant
Publicado: (2026) -
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
por: Mehta, Sushant
Publicado: (2025) -
Riemann-Bench: A Benchmark for Moonshot Mathematics
por: Garre, Suhaas, et al.
Publicado: (2026)