Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Minzheng, Luo, Run, Wang, Yanbo, Liu, Zichen, Tan, Yuqiao, Tan, Tao, Nan, Xu, Zheng, Yinhe, Mao, Wenji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Social Learning via Mode Policy Optimization for Language Agents
di: Wang, Minzheng, et al.
Pubblicazione: (2025)
di: Wang, Minzheng, et al.
Pubblicazione: (2025)
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
di: Tan, Yuqiao, et al.
Pubblicazione: (2026)
di: Tan, Yuqiao, et al.
Pubblicazione: (2026)
YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction
di: Xiao, Xinglin, et al.
Pubblicazione: (2023)
di: Xiao, Xinglin, et al.
Pubblicazione: (2023)
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
di: Wang, Minzheng, et al.
Pubblicazione: (2024)
di: Wang, Minzheng, et al.
Pubblicazione: (2024)
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
Continuous QA Learning with Structured Prompts
di: Zheng, Yinhe
Pubblicazione: (2022)
di: Zheng, Yinhe
Pubblicazione: (2022)
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale
di: Zheng, Wenzhen, et al.
Pubblicazione: (2024)
di: Zheng, Wenzhen, et al.
Pubblicazione: (2024)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
di: Fan, Ziqing, et al.
Pubblicazione: (2025)
di: Fan, Ziqing, et al.
Pubblicazione: (2025)
Asymmetric Catalytic Aziridination to Synthesize Spiro‐aziridine Oxindoles
di: Yinhe Qu, et al.
Pubblicazione: (2025)
di: Yinhe Qu, et al.
Pubblicazione: (2025)
College Students' Behavioural Intentions of AI‐Assisted Language Learning: Based on the Technology Acceptance Model
di: Wenji Wang, et al.
Pubblicazione: (2025)
di: Wenji Wang, et al.
Pubblicazione: (2025)
Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery Policy
di: Cao, Chenyang, et al.
Pubblicazione: (2024)
di: Cao, Chenyang, et al.
Pubblicazione: (2024)
PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding
di: Che, Lirong, et al.
Pubblicazione: (2026)
di: Che, Lirong, et al.
Pubblicazione: (2026)
A Universal Vehicle-Trailer Navigation System with Neural Kinematics and Online Residual Learning
di: Chen, Yanbo, et al.
Pubblicazione: (2025)
di: Chen, Yanbo, et al.
Pubblicazione: (2025)
Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics
di: Yu, Zhaoxin, et al.
Pubblicazione: (2026)
di: Yu, Zhaoxin, et al.
Pubblicazione: (2026)
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
di: Zou, Tao, et al.
Pubblicazione: (2025)
di: Zou, Tao, et al.
Pubblicazione: (2025)
MuAP: Multi-step Adaptive Prompt Learning for Vision-Language Model with Missing Modality
di: Dai, Ruiting, et al.
Pubblicazione: (2024)
di: Dai, Ruiting, et al.
Pubblicazione: (2024)
Scaling Behaviors of Evolutionary Algorithms on GPUs: When Does Parallelism Pay Off?
di: Yu, Xinmeng, et al.
Pubblicazione: (2026)
di: Yu, Xinmeng, et al.
Pubblicazione: (2026)
Effects of different culture modes on the storage quality of tilapia (Oreochromis mossambicus) fillets at 4°C
di: Dongya Zhao, et al.
Pubblicazione: (2024)
di: Dongya Zhao, et al.
Pubblicazione: (2024)
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
di: Yuan, Siyu, et al.
Pubblicazione: (2024)
di: Yuan, Siyu, et al.
Pubblicazione: (2024)
Breaking Task Impasses Quickly: Adaptive Neuro-Symbolic Learning for Open-World Robotics
di: Lorang, Pierrick
Pubblicazione: (2026)
di: Lorang, Pierrick
Pubblicazione: (2026)
Large Language Model-Aided Evolutionary Search for Constrained Multiobjective Optimization
di: Wang, Zeyi, et al.
Pubblicazione: (2024)
di: Wang, Zeyi, et al.
Pubblicazione: (2024)
Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?
di: Zhao, Yibo, et al.
Pubblicazione: (2026)
di: Zhao, Yibo, et al.
Pubblicazione: (2026)
MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks
di: Hao, Run, et al.
Pubblicazione: (2026)
di: Hao, Run, et al.
Pubblicazione: (2026)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
di: Zhao, Shiwan, et al.
Pubblicazione: (2026)
di: Zhao, Shiwan, et al.
Pubblicazione: (2026)
Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Models
di: Liu, Yunqing, et al.
Pubblicazione: (2025)
di: Liu, Yunqing, et al.
Pubblicazione: (2025)
Cores and weights of multipartitions and blocks of Ariki-Koike algebras
di: Li, Yanbo, et al.
Pubblicazione: (2024)
di: Li, Yanbo, et al.
Pubblicazione: (2024)
Policy and World Modeling Co-Training for Language Agents
di: Lu, Ning, et al.
Pubblicazione: (2026)
di: Lu, Ning, et al.
Pubblicazione: (2026)
Re-Visioning Arts and Cultural Policy: Current Impasses and Future Directions
di: Craik, Jennifer
Pubblicazione: (2013)
di: Craik, Jennifer
Pubblicazione: (2013)
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
di: Shi, Yuchen, et al.
Pubblicazione: (2025)
di: Shi, Yuchen, et al.
Pubblicazione: (2025)
Minute-Long Videos with Dual Parallelisms
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
Fragments of Martin's axiom
di: Peng, Yinhe
Pubblicazione: (2025)
di: Peng, Yinhe
Pubblicazione: (2025)
A ccc indestructible construction with CH
di: Peng, Yinhe
Pubblicazione: (2025)
di: Peng, Yinhe
Pubblicazione: (2025)
Distinguishing Martin's axiom from its restrictions
di: Peng, Yinhe
Pubblicazione: (2024)
di: Peng, Yinhe
Pubblicazione: (2024)
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
Breaking the Charge Limitation Constrained by Triboelectrification through the Charge Multiplication Effect
di: Ai Chen, et al.
Pubblicazione: (2025)
di: Ai Chen, et al.
Pubblicazione: (2025)
Nonparametric Uniform Inference in Binary Classification and Policy Values
di: Liu, Nan, et al.
Pubblicazione: (2025)
di: Liu, Nan, et al.
Pubblicazione: (2025)
FOSP: Fine-tuning Offline Safe Policy through World Models
di: Cao, Chenyang, et al.
Pubblicazione: (2024)
di: Cao, Chenyang, et al.
Pubblicazione: (2024)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
di: Li, Jiaze, et al.
Pubblicazione: (2026)
di: Li, Jiaze, et al.
Pubblicazione: (2026)
A Knowledge Enhanced Learning and Semantic Composition Model for Multi-Claim Fact Checking
di: Wang, Shuai, et al.
Pubblicazione: (2021)
di: Wang, Shuai, et al.
Pubblicazione: (2021)
Documenti analoghi
-
Adaptive Social Learning via Mode Policy Optimization for Language Agents
di: Wang, Minzheng, et al.
Pubblicazione: (2025) -
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
di: Tan, Yuqiao, et al.
Pubblicazione: (2026) -
YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction
di: Xiao, Xinglin, et al.
Pubblicazione: (2023) -
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
di: Wang, Minzheng, et al.
Pubblicazione: (2024) -
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)