Co-Evolution of Policy and Internal Reward for Language Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xinyu, Wu, Hanwei, Song, Jingwei, Zhang, Shuyuan, Zhang, Jiayi, Kong, Fanqi, Kwok, Tung Sum Thomas, Chang, Xiao-Wen, Luo, Yuyu, Wu, Chenglin, Liu, Bang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
por: Kong, Fanqi, et al.
Publicado: (2026)
por: Kong, Fanqi, et al.
Publicado: (2026)
Scalable Environments Drive Generalizable Agents
por: Zhang, Jiayi, et al.
Publicado: (2026)
por: Zhang, Jiayi, et al.
Publicado: (2026)
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
por: Zhang, Jiayi, et al.
Publicado: (2025)
por: Zhang, Jiayi, et al.
Publicado: (2025)
Enhancing Table Reasoning with Deterministic Table-State Rewards
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)
From Table to Cell: Attention for Better Reasoning with TABALIGN
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)
Harnessing Agentic Evolution
por: Zhang, Jiayi, et al.
Publicado: (2026)
por: Zhang, Jiayi, et al.
Publicado: (2026)
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
por: Ruan, Jianhao, et al.
Publicado: (2026)
por: Ruan, Jianhao, et al.
Publicado: (2026)
TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)
Atom of Thoughts for Markov LLM Test-Time Scaling
por: Teng, Fengwei, et al.
Publicado: (2025)
por: Teng, Fengwei, et al.
Publicado: (2025)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
por: Xie, Yupeng, et al.
Publicado: (2025)
por: Xie, Yupeng, et al.
Publicado: (2025)
Mutual-Taught for Co-adapting Policy and Reward Models
por: Shi, Tianyuan, et al.
Publicado: (2025)
por: Shi, Tianyuan, et al.
Publicado: (2025)
Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2025)
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2025)
DEREC-SIMPRO: unlock Language Model benefits to advance Synthesis in Data Clean Room
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2024)
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2024)
GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2025)
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2025)
ReCode: Unify Plan and Action for Universal Granularity Control
por: Yu, Zhaoyang, et al.
Publicado: (2025)
por: Yu, Zhaoyang, et al.
Publicado: (2025)
Constructing the GenAI Literacy Model for Pre‐Service Second Language Teachers: A Behavioural Event Interview Approach
por: Hanwei Wu, et al.
Publicado: (2026)
por: Hanwei Wu, et al.
Publicado: (2026)
SEAL: Synergistic Co-Evolution of Agents and Learning Environments
por: Hu, Yihao, et al.
Publicado: (2026)
por: Hu, Yihao, et al.
Publicado: (2026)
SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation
por: Wang, Xinyu, et al.
Publicado: (2026)
por: Wang, Xinyu, et al.
Publicado: (2026)
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution
por: Huang, Changxin, et al.
Publicado: (2024)
por: Huang, Changxin, et al.
Publicado: (2024)
AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning
por: Pan, Zhenyu, et al.
Publicado: (2025)
por: Pan, Zhenyu, et al.
Publicado: (2025)
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
por: Zhang, Yanfei, et al.
Publicado: (2026)
por: Zhang, Yanfei, et al.
Publicado: (2026)
Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety
por: Pan, Zhenyu, et al.
Publicado: (2025)
por: Pan, Zhenyu, et al.
Publicado: (2025)
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
por: Li, Bolian, et al.
Publicado: (2025)
por: Li, Bolian, et al.
Publicado: (2025)
Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
por: Kong, Fanqi, et al.
Publicado: (2025)
por: Kong, Fanqi, et al.
Publicado: (2025)
Self-Evolution Fine-Tuning for Policy Optimization
por: Chen, Ruijun, et al.
Publicado: (2024)
por: Chen, Ruijun, et al.
Publicado: (2024)
MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward
por: Chen, Kesheng, et al.
Publicado: (2025)
por: Chen, Kesheng, et al.
Publicado: (2025)
Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
por: Wu, Yifan, et al.
Publicado: (2025)
por: Wu, Yifan, et al.
Publicado: (2025)
SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
por: Chi, Yizhou, et al.
Publicado: (2024)
por: Chi, Yizhou, et al.
Publicado: (2024)
ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
por: Zhang, Yiwen, et al.
Publicado: (2025)
por: Zhang, Yiwen, et al.
Publicado: (2025)
Trajectory-Oriented Policy Optimization with Sparse Rewards
por: Wang, Guojian, et al.
Publicado: (2024)
por: Wang, Guojian, et al.
Publicado: (2024)
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
por: Wu, Yifan, et al.
Publicado: (2026)
por: Wu, Yifan, et al.
Publicado: (2026)
Improving Context Fidelity via Native Retrieval-Augmented Reasoning
por: Wang, Suyuchen, et al.
Publicado: (2025)
por: Wang, Suyuchen, et al.
Publicado: (2025)
AFlow: Automating Agentic Workflow Generation
por: Zhang, Jiayi, et al.
Publicado: (2024)
por: Zhang, Jiayi, et al.
Publicado: (2024)
Superlattice Engineering Regulation to Address Structural Defect Evolution in Ni‐Rich Layered Cathodes
por: Zhichen Hou, et al.
Publicado: (2026)
por: Zhichen Hou, et al.
Publicado: (2026)
TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
por: Wang, Ziyi, et al.
Publicado: (2026)
por: Wang, Ziyi, et al.
Publicado: (2026)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
por: Zhang, Yuxin, et al.
Publicado: (2025)
por: Zhang, Yuxin, et al.
Publicado: (2025)
JND-Guided Neural Watermarking with Spatial Transformer Decoding for Screen-Capture Robustness
por: Qin, Jiayi, et al.
Publicado: (2026)
por: Qin, Jiayi, et al.
Publicado: (2026)
What Deserve Studying the Most? A Q‐Methodology Approach to Explore Stakeholders' Perspectives on Research Priorities in GenAI ‐Supported Second Language Education
por: Hanwei Wu, et al.
Publicado: (2024)
por: Hanwei Wu, et al.
Publicado: (2024)
Self-Supervised Prompt Optimization
por: Xiang, Jinyu, et al.
Publicado: (2025)
por: Xiang, Jinyu, et al.
Publicado: (2025)
Latent Action Reparameterization for Efficient Agent Inference
por: Huang, Wenhao, et al.
Publicado: (2026)
por: Huang, Wenhao, et al.
Publicado: (2026)
Ejemplares similares
-
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
por: Kong, Fanqi, et al.
Publicado: (2026) -
Scalable Environments Drive Generalizable Agents
por: Zhang, Jiayi, et al.
Publicado: (2026) -
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
por: Zhang, Jiayi, et al.
Publicado: (2025) -
Enhancing Table Reasoning with Deterministic Table-State Rewards
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026) -
From Table to Cell: Attention for Better Reasoning with TABALIGN
por: Kwok, Tung Sum Thomas, et al.
Publicado: (2026)