Adaptive Social Learning via Mode Policy Optimization for Language Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Minzheng, Li, Yongbin, Wang, Haobo, Zhang, Xinghua, Xu, Nan, Wu, Bingli, Huang, Fei, Yu, Haiyang, Mao, Wenji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
von: Zou, Tao, et al.
Veröffentlicht: (2025)
von: Zou, Tao, et al.
Veröffentlicht: (2025)
The Imperative of Conversation Analysis in the Era of LLMs: A Survey of Tasks, Techniques, and Trends
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
von: Wang, Minzheng, et al.
Veröffentlicht: (2026)
von: Wang, Minzheng, et al.
Veröffentlicht: (2026)
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
ExpSeek: Self-Triggered Experience Seeking for Web Agents
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2026)
EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
von: Zhang, Guibin, et al.
Veröffentlicht: (2026)
von: Zhang, Guibin, et al.
Veröffentlicht: (2026)
Extend Model Merging from Fine-Tuned to Pre-Trained Large Language Models via Weight Disentanglement
von: Yu, Le, et al.
Veröffentlicht: (2024)
von: Yu, Le, et al.
Veröffentlicht: (2024)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
Fine-Tuning Language Models with Reward Learning on Policy
von: Lang, Hao, et al.
Veröffentlicht: (2024)
von: Lang, Hao, et al.
Veröffentlicht: (2024)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction
von: Xiao, Xinglin, et al.
Veröffentlicht: (2023)
von: Xiao, Xinglin, et al.
Veröffentlicht: (2023)
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2025)
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
von: Yu, Le, et al.
Veröffentlicht: (2023)
von: Yu, Le, et al.
Veröffentlicht: (2023)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024)
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024)
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
von: Ye, Xinge, et al.
Veröffentlicht: (2025)
von: Ye, Xinge, et al.
Veröffentlicht: (2025)
Preference Ranking Optimization for Human Alignment
von: Song, Feifan, et al.
Veröffentlicht: (2023)
von: Song, Feifan, et al.
Veröffentlicht: (2023)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization
von: Cui, Chaoqun, et al.
Veröffentlicht: (2026)
von: Cui, Chaoqun, et al.
Veröffentlicht: (2026)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
von: Song, Feifan, et al.
Veröffentlicht: (2024)
von: Song, Feifan, et al.
Veröffentlicht: (2024)
SDPO: Segment-Level Direct Preference Optimization for Social Agents
von: Kong, Aobo, et al.
Veröffentlicht: (2025)
von: Kong, Aobo, et al.
Veröffentlicht: (2025)
Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
MOA: Multi-Objective Alignment for Role-Playing Agents
von: Liao, Chonghua, et al.
Veröffentlicht: (2025)
von: Liao, Chonghua, et al.
Veröffentlicht: (2025)
Perception-Aware Policy Optimization for Multimodal Reasoning
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
Improving Retrospective Language Agents via Joint Policy Gradient Optimization
von: Feng, Xueyang, et al.
Veröffentlicht: (2025)
von: Feng, Xueyang, et al.
Veröffentlicht: (2025)
A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
von: Zhao, Yingxiu, et al.
Veröffentlicht: (2023)
von: Zhao, Yingxiu, et al.
Veröffentlicht: (2023)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
von: Zhang, Wenqi, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2024)
Iterative Forward Tuning Boosts In-Context Learning in Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
von: Luo, Run, et al.
Veröffentlicht: (2024)
von: Luo, Run, et al.
Veröffentlicht: (2024)
Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration
von: Ma, Yingwei, et al.
Veröffentlicht: (2024)
von: Ma, Yingwei, et al.
Veröffentlicht: (2024)
Agentic Reinforcement Learning with Implicit Step Rewards
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2025)
DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
von: Li, Ruoran, et al.
Veröffentlicht: (2026)
von: Li, Ruoran, et al.
Veröffentlicht: (2026)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
Transferable Post-training via Inverse Value Learning
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
Debate Helps Weak-to-Strong Generalization
von: Lang, Hao, et al.
Veröffentlicht: (2025)
von: Lang, Hao, et al.
Veröffentlicht: (2025)
Selective Weak-to-Strong Generalization
von: Lang, Hao, et al.
Veröffentlicht: (2025)
von: Lang, Hao, et al.
Veröffentlicht: (2025)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
von: Wang, Minzheng, et al.
Veröffentlicht: (2024) -
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
von: Zou, Tao, et al.
Veröffentlicht: (2025) -
The Imperative of Conversation Analysis in the Era of LLMs: A Survey of Tasks, Techniques, and Trends
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024) -
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
von: Wang, Minzheng, et al.
Veröffentlicht: (2026) -
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)