TRACE: Capability-Targeted Agentic Training
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Hangoo, Suresh, Tarun, Saad-Falcon, Jon, Mirhoseini, Azalia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRAP: Targeted Redirecting of Agentic Preferences
by: Kang, Hangoo, et al.
Published: (2025)
by: Kang, Hangoo, et al.
Published: (2025)
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
by: Goldie, Anna, et al.
Published: (2024)
by: Goldie, Anna, et al.
Published: (2024)
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
by: Winston, Caleb, et al.
Published: (2026)
by: Winston, Caleb, et al.
Published: (2026)
Archon: An Architecture Search Framework for Inference-Time Techniques
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
by: Talaei, Shayan, et al.
Published: (2024)
by: Talaei, Shayan, et al.
Published: (2024)
SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
by: Biju, Emil, et al.
Published: (2025)
by: Biju, Emil, et al.
Published: (2025)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
by: Goldie, Anna, et al.
Published: (2025)
by: Goldie, Anna, et al.
Published: (2025)
OpenJarvis: Personal AI, On Personal Devices
by: Saad-Falcon, Jon, et al.
Published: (2026)
by: Saad-Falcon, Jon, et al.
Published: (2026)
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
by: Kwok, Jacky, et al.
Published: (2026)
by: Kwok, Jacky, et al.
Published: (2026)
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
by: Brown, Bradley, et al.
Published: (2024)
by: Brown, Bradley, et al.
Published: (2024)
TRACE: A Conversational Framework for Sustainable Tourism Recommendation with Agentic Counterfactual Explanations
by: Banerjee, Ashmi, et al.
Published: (2026)
by: Banerjee, Ashmi, et al.
Published: (2026)
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
by: Kwok, Jacky, et al.
Published: (2025)
by: Kwok, Jacky, et al.
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
by: Vega, Jason, et al.
Published: (2024)
by: Vega, Jason, et al.
Published: (2024)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
TRACE: Learning to Compute on Circuit Graphs
by: Zheng, Ziyang, et al.
Published: (2025)
by: Zheng, Ziyang, et al.
Published: (2025)
Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL
by: Pourreza, Mohammadreza, et al.
Published: (2025)
by: Pourreza, Mohammadreza, et al.
Published: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
How Do Large Language Monkeys Get Their Power (Laws)?
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
by: Wei, Qianshan, et al.
Published: (2026)
by: Wei, Qianshan, et al.
Published: (2026)
Proper Scoring Rules for Agentic Uncertainty Quantification
by: Raghu, Suresh, et al.
Published: (2026)
by: Raghu, Suresh, et al.
Published: (2026)
Learning a Pessimistic Reward Model in RLHF
by: Xu, Yinglun, et al.
Published: (2025)
by: Xu, Yinglun, et al.
Published: (2025)
SynCode: LLM Generation with Grammar Augmentation
by: Ugare, Shubham, et al.
Published: (2024)
by: Ugare, Shubham, et al.
Published: (2024)
TRACE: Textual Reasoning for Affordance Coordinate Extraction
by: Park, Sangyun, et al.
Published: (2025)
by: Park, Sangyun, et al.
Published: (2025)
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
by: Chu, Meng, et al.
Published: (2026)
by: Chu, Meng, et al.
Published: (2026)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
by: Zhu, Pengyu, et al.
Published: (2026)
by: Zhu, Pengyu, et al.
Published: (2026)
Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems
by: Babu, Rahul Suresh, et al.
Published: (2026)
by: Babu, Rahul Suresh, et al.
Published: (2026)
TRACE: A Metrologically-Grounded Engineering Framework for Trustworthy Agentic AI Systems in Operationally Critical Domains
by: Zabolotnii, Serhii
Published: (2026)
by: Zabolotnii, Serhii
Published: (2026)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
AgenticRAG: Agentic Retrieval for Enterprise Knowledge Bases
by: Suresh, Susheel, et al.
Published: (2026)
by: Suresh, Susheel, et al.
Published: (2026)
Cartridges: Lightweight and general-purpose long context representations via self-study
by: Eyuboglu, Sabri, et al.
Published: (2025)
by: Eyuboglu, Sabri, et al.
Published: (2025)
Analysis of the Memorization and Generalization Capabilities of AI Agents: Are Continual Learners Robust?
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
by: Hong, Zhepei, et al.
Published: (2026)
by: Hong, Zhepei, et al.
Published: (2026)
The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
by: Ritchie, Logan, et al.
Published: (2026)
by: Ritchie, Logan, et al.
Published: (2026)
TRACE: Transparent Web Reliability Assessment with Contextual Explanations
by: Chandra, Joydeep, et al.
Published: (2025)
by: Chandra, Joydeep, et al.
Published: (2025)
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
by: Costello, Caia, et al.
Published: (2025)
by: Costello, Caia, et al.
Published: (2025)
TRACE-CS: A Hybrid Logic-LLM System for Explainable Course Scheduling
by: Vasileiou, Stylianos Loukas, et al.
Published: (2024)
by: Vasileiou, Stylianos Loukas, et al.
Published: (2024)
Similar Items
-
TRAP: Targeted Redirecting of Agentic Preferences
by: Kang, Hangoo, et al.
Published: (2025) -
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
by: Goldie, Anna, et al.
Published: (2024) -
ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025) -
On the Role of Temperature Sampling in Test-Time Scaling
by: Wu, Yuheng, et al.
Published: (2025) -
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
by: Winston, Caleb, et al.
Published: (2026)