Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cai, Yuandao, Zhu, Yuzhang, Gao, Liyou, Tang, Wensheng, Qin, Shengchao |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ConCovUp: Effective Agent-Based Test Driver Generation for Concurrency Testing
par: Cai, Yuandao, et autres
Publié: (2026)
par: Cai, Yuandao, et autres
Publié: (2026)
FROAV: A Framework for RAG Observation and Agent Verification -- Lowering the Barrier to LLM Agent Research
par: Lin, Tzu-Hsuan, et autres
Publié: (2026)
par: Lin, Tzu-Hsuan, et autres
Publié: (2026)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
par: Wang, Peiran, et autres
Publié: (2025)
par: Wang, Peiran, et autres
Publié: (2025)
Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
par: Cai, Yuandao, et autres
Publié: (2026)
par: Cai, Yuandao, et autres
Publié: (2026)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
par: Kuntz, Thomas, et autres
Publié: (2025)
par: Kuntz, Thomas, et autres
Publié: (2025)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
par: Han, Xiaoke, et autres
Publié: (2025)
par: Han, Xiaoke, et autres
Publié: (2025)
Your Code Agent Can Grow Alongside You with Structured Memory
par: Deng, Yi-Xuan, et autres
Publié: (2026)
par: Deng, Yi-Xuan, et autres
Publié: (2026)
MooseAgent: A LLM Based Multi-agent Framework for Automating Moose Simulation
par: Zhang, Tao, et autres
Publié: (2025)
par: Zhang, Tao, et autres
Publié: (2025)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
par: Manglik, Akshay, et autres
Publié: (2026)
par: Manglik, Akshay, et autres
Publié: (2026)
Boosting Path-Sensitive Value Flow Analysis via Removal of Redundant Summaries
par: Wang, Yongchao, et autres
Publié: (2025)
par: Wang, Yongchao, et autres
Publié: (2025)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
par: Jasper, Surya, et autres
Publié: (2025)
par: Jasper, Surya, et autres
Publié: (2025)
Measuring Agents in Production
par: Pan, Melissa Z., et autres
Publié: (2025)
par: Pan, Melissa Z., et autres
Publié: (2025)
AgentForge: A Flexible Low-Code Platform for Reinforcement Learning Agent Design
par: Junior, Francisco Erivaldo Fernandes, et autres
Publié: (2024)
par: Junior, Francisco Erivaldo Fernandes, et autres
Publié: (2024)
The Dual-State Architecture for Reliable LLM Agents
par: Thompson, Matthew
Publié: (2025)
par: Thompson, Matthew
Publié: (2025)
Exploring LLM-based Agents for Root Cause Analysis
par: Roy, Devjeet, et autres
Publié: (2024)
par: Roy, Devjeet, et autres
Publié: (2024)
Bootstrapping Coding Agents: The Specification Is the Program
par: Monperrus, Martin
Publié: (2026)
par: Monperrus, Martin
Publié: (2026)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
par: Kim, Naryeong, et autres
Publié: (2026)
par: Kim, Naryeong, et autres
Publié: (2026)
ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories
par: Ibrahimzada, Ali Reza, et autres
Publié: (2026)
par: Ibrahimzada, Ali Reza, et autres
Publié: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
par: Rank, Ben, et autres
Publié: (2026)
par: Rank, Ben, et autres
Publié: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
par: Yu, Shasha, et autres
Publié: (2026)
par: Yu, Shasha, et autres
Publié: (2026)
Agint: Agentic Graph Compilation for Software Engineering Agents
par: Chivukula, Abhi, et autres
Publié: (2025)
par: Chivukula, Abhi, et autres
Publié: (2025)
daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently
par: Jiang, Mohan, et autres
Publié: (2026)
par: Jiang, Mohan, et autres
Publié: (2026)
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
par: Golubev, Alexander, et autres
Publié: (2025)
par: Golubev, Alexander, et autres
Publié: (2025)
Data Requirement Goal Modeling for Machine Learning Systems
par: Yamani, Asma, et autres
Publié: (2025)
par: Yamani, Asma, et autres
Publié: (2025)
Can Coding Agents Be General Agents?
par: Ivanov, Maksim, et autres
Publié: (2026)
par: Ivanov, Maksim, et autres
Publié: (2026)
CA2: Code-Aware Agent for Automated Game Testing
par: Adaikkappan, Valliappan Chidambaram, et autres
Publié: (2026)
par: Adaikkappan, Valliappan Chidambaram, et autres
Publié: (2026)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
par: Raghavendra, Mohit, et autres
Publié: (2026)
par: Raghavendra, Mohit, et autres
Publié: (2026)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
par: Yoon, Juyeon, et autres
Publié: (2025)
par: Yoon, Juyeon, et autres
Publié: (2025)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
par: Xi, Haoran, et autres
Publié: (2025)
par: Xi, Haoran, et autres
Publié: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
par: Xiao, Yijia, et autres
Publié: (2025)
par: Xiao, Yijia, et autres
Publié: (2025)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
par: Bula, Timothy, et autres
Publié: (2025)
par: Bula, Timothy, et autres
Publié: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
par: Aggarwal, Pranjal, et autres
Publié: (2025)
par: Aggarwal, Pranjal, et autres
Publié: (2025)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
par: Gao, Pengfei, et autres
Publié: (2025)
par: Gao, Pengfei, et autres
Publié: (2025)
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
par: Jin, Yiyang, et autres
Publié: (2025)
par: Jin, Yiyang, et autres
Publié: (2025)
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
par: Wang, Shuaiqi, et autres
Publié: (2026)
par: Wang, Shuaiqi, et autres
Publié: (2026)
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
par: Dhulipala, Hridya, et autres
Publié: (2025)
par: Dhulipala, Hridya, et autres
Publié: (2025)
Concept-Guided LLM Agents for Human-AI Safety Codesign
par: Geissler, Florian, et autres
Publié: (2024)
par: Geissler, Florian, et autres
Publié: (2024)
Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance
par: Dimidov, Valeriu, et autres
Publié: (2025)
par: Dimidov, Valeriu, et autres
Publié: (2025)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
par: Wu, Jie JW, et autres
Publié: (2025)
par: Wu, Jie JW, et autres
Publié: (2025)
Follow Your Nose -- Which Code Smells are Worth Chasing?
par: Amit, Idan, et autres
Publié: (2021)
par: Amit, Idan, et autres
Publié: (2021)
Documents similaires
-
ConCovUp: Effective Agent-Based Test Driver Generation for Concurrency Testing
par: Cai, Yuandao, et autres
Publié: (2026) -
FROAV: A Framework for RAG Observation and Agent Verification -- Lowering the Barrier to LLM Agent Research
par: Lin, Tzu-Hsuan, et autres
Publié: (2026) -
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
par: Wang, Peiran, et autres
Publié: (2025) -
Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents
par: Cai, Yuandao, et autres
Publié: (2026) -
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
par: Kuntz, Thomas, et autres
Publié: (2025)