InfoPO: Information-Driven Policy Optimization for User-Centric Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kong, Fanqi, Zhang, Jiayi, Deng, Mingyi, Wu, Chenglin, Luo, Yuyu, Liu, Bang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Environments Drive Generalizable Agents
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026)
Co-Evolution of Policy and Internal Reward for Language Agents
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
InfoPO: On Mutual Information Maximization for Large Language Model Alignment
von: Xiao, Teng, et al.
Veröffentlicht: (2025)
von: Xiao, Teng, et al.
Veröffentlicht: (2025)
ReCode: Unify Plan and Action for Universal Granularity Control
von: Yu, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoyang, et al.
Veröffentlicht: (2025)
AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
von: Ruan, Jianhao, et al.
Veröffentlicht: (2026)
von: Ruan, Jianhao, et al.
Veröffentlicht: (2026)
InfoAgent: Advancing Autonomous Information-Seeking Agents
von: Zhang, Gongrui, et al.
Veröffentlicht: (2025)
von: Zhang, Gongrui, et al.
Veröffentlicht: (2025)
InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
von: Luo, Kun, et al.
Veröffentlicht: (2025)
von: Luo, Kun, et al.
Veröffentlicht: (2025)
MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
von: Li, Ruoran, et al.
Veröffentlicht: (2026)
von: Li, Ruoran, et al.
Veröffentlicht: (2026)
Atom of Thoughts for Markov LLM Test-Time Scaling
von: Teng, Fengwei, et al.
Veröffentlicht: (2025)
von: Teng, Fengwei, et al.
Veröffentlicht: (2025)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
von: Xie, Yupeng, et al.
Veröffentlicht: (2025)
von: Xie, Yupeng, et al.
Veröffentlicht: (2025)
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
von: Liufu, Weijia, et al.
Veröffentlicht: (2026)
von: Liufu, Weijia, et al.
Veröffentlicht: (2026)
Self-Supervised Prompt Optimization
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
Phase Transition for Budgeted Multi-Agent Synergy
von: Liu, Bang, et al.
Veröffentlicht: (2026)
von: Liu, Bang, et al.
Veröffentlicht: (2026)
InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
von: Lee, Ka Yiu, et al.
Veröffentlicht: (2026)
von: Lee, Ka Yiu, et al.
Veröffentlicht: (2026)
InfoCon: Concept Discovery with Generative and Discriminative Informativeness
von: Liu, Ruizhe, et al.
Veröffentlicht: (2024)
von: Liu, Ruizhe, et al.
Veröffentlicht: (2024)
DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering
von: Lin, Teng, et al.
Veröffentlicht: (2026)
von: Lin, Teng, et al.
Veröffentlicht: (2026)
Harnessing Agentic Evolution
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026)
UserBench: An Interactive Gym Environment for User-Centric Agents
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Latent Action Reparameterization for Efficient Agent Inference
von: Huang, Wenhao, et al.
Veröffentlicht: (2026)
von: Huang, Wenhao, et al.
Veröffentlicht: (2026)
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
von: Zhang, Yanfei, et al.
Veröffentlicht: (2026)
von: Zhang, Yanfei, et al.
Veröffentlicht: (2026)
OPEx: A Component-Wise Analysis of LLM-Centric Agents in Embodied Instruction Following
von: Shi, Haochen, et al.
Veröffentlicht: (2024)
von: Shi, Haochen, et al.
Veröffentlicht: (2024)
ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
von: Zhang, Yiwen, et al.
Veröffentlicht: (2025)
von: Zhang, Yiwen, et al.
Veröffentlicht: (2025)
MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
von: Cheng, Weihua, et al.
Veröffentlicht: (2025)
von: Cheng, Weihua, et al.
Veröffentlicht: (2025)
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
von: Wu, Yifan, et al.
Veröffentlicht: (2026)
von: Wu, Yifan, et al.
Veröffentlicht: (2026)
Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
von: Wu, Yifan, et al.
Veröffentlicht: (2025)
von: Wu, Yifan, et al.
Veröffentlicht: (2025)
InteractComp: Evaluating Search Agents With Ambiguous Queries
von: Deng, Mingyi, et al.
Veröffentlicht: (2025)
von: Deng, Mingyi, et al.
Veröffentlicht: (2025)
InfoCom: Kilobyte-Scale Communication-Efficient Collaborative Perception with Information Bottleneck
von: Wei, Quanmin, et al.
Veröffentlicht: (2025)
von: Wei, Quanmin, et al.
Veröffentlicht: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
von: Yang, Chenglin
Veröffentlicht: (2026)
von: Yang, Chenglin
Veröffentlicht: (2026)
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
von: Li, Chenyi, et al.
Veröffentlicht: (2026)
von: Li, Chenyi, et al.
Veröffentlicht: (2026)
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
Mars-PO: Multi-Agent Reasoning System Preference Optimization
von: Lou, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Lou, Xiaoxuan, et al.
Veröffentlicht: (2024)
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
von: Du, Yaxin, et al.
Veröffentlicht: (2025)
von: Du, Yaxin, et al.
Veröffentlicht: (2025)
SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
von: Sun, Huashan, et al.
Veröffentlicht: (2025)
von: Sun, Huashan, et al.
Veröffentlicht: (2025)
RePO: Replay-Enhanced Policy Optimization
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
Foundation Protocol: A Coordination Layer for Agentic Society
von: Liu, Bang, et al.
Veröffentlicht: (2026)
von: Liu, Bang, et al.
Veröffentlicht: (2026)
Latent-Info and Low-Dimensional Learning for Human Mesh Recovery and Parallel Optimization
von: Zhang, Xiang, et al.
Veröffentlicht: (2025)
von: Zhang, Xiang, et al.
Veröffentlicht: (2025)
Agentic Enterprise: AI-Centric User to User-Centric AI
von: Narechania, Arpit, et al.
Veröffentlicht: (2025)
von: Narechania, Arpit, et al.
Veröffentlicht: (2025)
AFlow: Automating Agentic Workflow Generation
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scalable Environments Drive Generalizable Agents
von: Zhang, Jiayi, et al.
Veröffentlicht: (2026) -
Co-Evolution of Policy and Internal Reward for Language Agents
von: Wang, Xinyu, et al.
Veröffentlicht: (2026) -
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025) -
InfoPO: On Mutual Information Maximization for Large Language Model Alignment
von: Xiao, Teng, et al.
Veröffentlicht: (2025) -
ReCode: Unify Plan and Action for Universal Granularity Control
von: Yu, Zhaoyang, et al.
Veröffentlicht: (2025)