ProgAgent:A Continual RL Agent with Progress-Aware Rewards
Fuente:
arXiv
Guardado en:
| Autores principales: | Tan, Jinzhou, Adineera, Gabriel, Kim, Jinoh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ResWM: Residual-Action World Model for Visual RL
por: Zhang, Jseen, et al.
Publicado: (2026)
por: Zhang, Jseen, et al.
Publicado: (2026)
ProgRM: Build Better GUI Agents with Progress Rewards
por: Zhang, Danyang, et al.
Publicado: (2025)
por: Zhang, Danyang, et al.
Publicado: (2025)
ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive Training
por: Wang, Hui-Po, et al.
Publicado: (2021)
por: Wang, Hui-Po, et al.
Publicado: (2021)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
por: Salimi, Moein, et al.
Publicado: (2026)
por: Salimi, Moein, et al.
Publicado: (2026)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
por: Muslimani, Calarina, et al.
Publicado: (2025)
por: Muslimani, Calarina, et al.
Publicado: (2025)
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
por: Vora, Manav, et al.
Publicado: (2024)
por: Vora, Manav, et al.
Publicado: (2024)
Meta-RL Induces Exploration in Language Agents
por: Jiang, Yulun, et al.
Publicado: (2025)
por: Jiang, Yulun, et al.
Publicado: (2025)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
por: Duan, Xintong, et al.
Publicado: (2025)
por: Duan, Xintong, et al.
Publicado: (2025)
RDAR: Reward-Driven Agent Relevance Estimation for Autonomous Driving
por: Bosio, Carlo, et al.
Publicado: (2025)
por: Bosio, Carlo, et al.
Publicado: (2025)
RF-Agent: Automated Reward Function Design via Language Agent Tree Search
por: Gao, Ning, et al.
Publicado: (2026)
por: Gao, Ning, et al.
Publicado: (2026)
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
por: Zhang, Qiang, et al.
Publicado: (2026)
por: Zhang, Qiang, et al.
Publicado: (2026)
ProgCo: Program Helps Self-Correction of Large Language Models
por: Song, Xiaoshuai, et al.
Publicado: (2025)
por: Song, Xiaoshuai, et al.
Publicado: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
por: Xia, Yu, et al.
Publicado: (2025)
por: Xia, Yu, et al.
Publicado: (2025)
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
por: Cherepanov, Egor, et al.
Publicado: (2024)
por: Cherepanov, Egor, et al.
Publicado: (2024)
Mind the Model, Not the Agent: The Primacy Bias in Model-based RL
por: Qiao, Zhongjian, et al.
Publicado: (2023)
por: Qiao, Zhongjian, et al.
Publicado: (2023)
Proto Successor Measure: Representing the Behavior Space of an RL Agent
por: Agarwal, Siddhant, et al.
Publicado: (2024)
por: Agarwal, Siddhant, et al.
Publicado: (2024)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
por: Feng, Xiao, et al.
Publicado: (2026)
por: Feng, Xiao, et al.
Publicado: (2026)
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
por: Koh, Woosung, et al.
Publicado: (2024)
por: Koh, Woosung, et al.
Publicado: (2024)
Analysis of the Memorization and Generalization Capabilities of AI Agents: Are Continual Learners Robust?
por: Kim, Minsu, et al.
Publicado: (2023)
por: Kim, Minsu, et al.
Publicado: (2023)
Reward Sharpness-Aware Fine-Tuning for Diffusion Models
por: Kim, Kwanyoung, et al.
Publicado: (2026)
por: Kim, Kwanyoung, et al.
Publicado: (2026)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
por: Xu, Yifan, et al.
Publicado: (2025)
por: Xu, Yifan, et al.
Publicado: (2025)
Which Experiences Are Influential for RL Agents? Efficiently Estimating The Influence of Experiences
por: Hiraoka, Takuya, et al.
Publicado: (2024)
por: Hiraoka, Takuya, et al.
Publicado: (2024)
Process Reward Models for LLM Agents: Practical Framework and Directions
por: Choudhury, Sanjiban
Publicado: (2025)
por: Choudhury, Sanjiban
Publicado: (2025)
Online Continual Learning For Interactive Instruction Following Agents
por: Kim, Byeonghwi, et al.
Publicado: (2024)
por: Kim, Byeonghwi, et al.
Publicado: (2024)
Adaptive Milestone Reward for GUI Agents
por: Zheng, Congmin, et al.
Publicado: (2026)
por: Zheng, Congmin, et al.
Publicado: (2026)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
por: Nie, Yuzhou, et al.
Publicado: (2024)
por: Nie, Yuzhou, et al.
Publicado: (2024)
ProgVLA: Progress-Aware Robot Manipulation Skill Learning
por: Kim, Seungsu, et al.
Publicado: (2026)
por: Kim, Seungsu, et al.
Publicado: (2026)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
por: Lù, Xing Han, et al.
Publicado: (2025)
por: Lù, Xing Han, et al.
Publicado: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
por: Jin, Can, et al.
Publicado: (2025)
por: Jin, Can, et al.
Publicado: (2025)
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
por: Mehta, Sushant, et al.
Publicado: (2026)
por: Mehta, Sushant, et al.
Publicado: (2026)
Plan Before You Trade: Inference-Time Optimization for RL Trading Agents
por: Go, Eun, et al.
Publicado: (2026)
por: Go, Eun, et al.
Publicado: (2026)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
por: Schmied, Thomas, et al.
Publicado: (2025)
por: Schmied, Thomas, et al.
Publicado: (2025)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
por: Thaman, Kunvar
Publicado: (2026)
por: Thaman, Kunvar
Publicado: (2026)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
por: Wu, Boyuan
Publicado: (2025)
por: Wu, Boyuan
Publicado: (2025)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
por: Kim, Yoonjeon, et al.
Publicado: (2025)
por: Kim, Yoonjeon, et al.
Publicado: (2025)
AgentGA: Evolving Code Solutions in Agent-Seed Space
por: Tan, David Y. Y., et al.
Publicado: (2026)
por: Tan, David Y. Y., et al.
Publicado: (2026)
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
por: Ning, Yansong, et al.
Publicado: (2026)
por: Ning, Yansong, et al.
Publicado: (2026)
Spatially-Aware Transformer for Embodied Agents
por: Cho, Junmo, et al.
Publicado: (2024)
por: Cho, Junmo, et al.
Publicado: (2024)
FairAgent: Democratizing Fairness-Aware Machine Learning with LLM-Powered Agents
por: Dai, Yucong, et al.
Publicado: (2025)
por: Dai, Yucong, et al.
Publicado: (2025)
Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
por: Kim, Sung-Hyun, et al.
Publicado: (2025)
por: Kim, Sung-Hyun, et al.
Publicado: (2025)
Ejemplares similares
-
ResWM: Residual-Action World Model for Visual RL
por: Zhang, Jseen, et al.
Publicado: (2026) -
ProgRM: Build Better GUI Agents with Progress Rewards
por: Zhang, Danyang, et al.
Publicado: (2025) -
ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive Training
por: Wang, Hui-Po, et al.
Publicado: (2021) -
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
por: Salimi, Moein, et al.
Publicado: (2026) -
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
por: Muslimani, Calarina, et al.
Publicado: (2025)