Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Lang, Lin, Jiahao, Xing, Dong, Zhang, Li, Ma, De, Pan, Gang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation Tree
by: Feng, Lang, et al.
Published: (2024)
by: Feng, Lang, et al.
Published: (2024)
Behavioral Embeddings of Programs: A Quasi-Dynamic Approach for Optimization Prediction
by: Pan, Haolin, et al.
Published: (2025)
by: Pan, Haolin, et al.
Published: (2025)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
BiBLDR: Bidirectional Behavior Learning for Drug Repositioning
by: Zhang, Renye, et al.
Published: (2025)
by: Zhang, Renye, et al.
Published: (2025)
Offline Behavior Distillation
by: Lei, Shiye, et al.
Published: (2024)
by: Lei, Shiye, et al.
Published: (2024)
The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning
by: Li, Yurui, et al.
Published: (2025)
by: Li, Yurui, et al.
Published: (2025)
Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation
by: Zhong, Yisheng, et al.
Published: (2026)
by: Zhong, Yisheng, et al.
Published: (2026)
Traversal-as-Policy: Log-Distilled Gated Behavior Trees as Externalized, Verifiable Policies for Safe, Robust, and Efficient Agents
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
Strategic Over-Parameterization for Generalizable Low-Rank Adaptation
by: Gao, Jing, et al.
Published: (2026)
by: Gao, Jing, et al.
Published: (2026)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
by: Chen, Zhaorun, et al.
Published: (2025)
by: Chen, Zhaorun, et al.
Published: (2025)
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
by: Chai, Jiajun, et al.
Published: (2025)
by: Chai, Jiajun, et al.
Published: (2025)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
by: Zhang, Shu-Hao, et al.
Published: (2026)
by: Zhang, Shu-Hao, et al.
Published: (2026)
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
by: Yang, Haochen, et al.
Published: (2026)
by: Yang, Haochen, et al.
Published: (2026)
Structured Agent Distillation for Large Language Model
by: Liu, Jun, et al.
Published: (2025)
by: Liu, Jun, et al.
Published: (2025)
LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
SmartPlay: A Benchmark for LLMs as Intelligent Agents
by: Wu, Yue, et al.
Published: (2023)
by: Wu, Yue, et al.
Published: (2023)
Group-in-Group Policy Optimization for LLM Agent Training
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
by: Caccia, Lucas, et al.
Published: (2025)
by: Caccia, Lucas, et al.
Published: (2025)
A Multi-Task Targeted Learning Framework for Lithium-Ion Battery State-of-Health and Remaining Useful Life
by: Wang, Chenhan, et al.
Published: (2026)
by: Wang, Chenhan, et al.
Published: (2026)
MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair
by: Li, Changqing, et al.
Published: (2025)
by: Li, Changqing, et al.
Published: (2025)
Role Play: Learning Adaptive Role-Specific Strategies in Multi-Agent Interactions
by: Long, Weifan, et al.
Published: (2024)
by: Long, Weifan, et al.
Published: (2024)
AgentOCR: Reimagining Agent History via Optical Self-Compression
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
EPR-GAIL: An EPR-Enhanced Hierarchical Imitation Learning Framework to Simulate Complex User Consumption Behaviors
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
EMOD: A Unified EEG Emotion Representation Framework Leveraging V-A Guided Contrastive Learning
by: Chen, Yuning, et al.
Published: (2025)
by: Chen, Yuning, et al.
Published: (2025)
Advancing Multi-Organ Disease Care: A Hierarchical Multi-Agent Reinforcement Learning Framework
by: Tan, Daniel J., et al.
Published: (2024)
by: Tan, Daniel J., et al.
Published: (2024)
Modeling Human Beliefs about AI Behavior for Scalable Oversight
by: Lang, Leon, et al.
Published: (2025)
by: Lang, Leon, et al.
Published: (2025)
MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
by: Duan, Wenchang, et al.
Published: (2026)
by: Duan, Wenchang, et al.
Published: (2026)
PlayGen-MoG: Framework for Diverse Multi-Agent Play Generation via Mixture-of-Gaussians Trajectory Prediction
by: Song, Kevin
Published: (2026)
by: Song, Kevin
Published: (2026)
Understanding Generalization in Role-Playing Models via Information Theory
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
Integrating Knowledge Distillation Methods: A Sequential Multi-Stage Framework
by: Tian, Yinxi, et al.
Published: (2026)
by: Tian, Yinxi, et al.
Published: (2026)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Learning Game-Playing Agents with Generative Code Optimization
by: Kuang, Zhiyi, et al.
Published: (2025)
by: Kuang, Zhiyi, et al.
Published: (2025)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
by: Yao, Bohan, et al.
Published: (2025)
by: Yao, Bohan, et al.
Published: (2025)
AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
by: Du, Jinye, et al.
Published: (2025)
by: Du, Jinye, et al.
Published: (2025)
Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent Cooperation
by: Liu, Sicong, et al.
Published: (2025)
by: Liu, Sicong, et al.
Published: (2025)
Similar Items
-
Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation Tree
by: Feng, Lang, et al.
Published: (2024) -
Behavioral Embeddings of Programs: A Quasi-Dynamic Approach for Optimization Prediction
by: Pan, Haolin, et al.
Published: (2025) -
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026) -
BiBLDR: Bidirectional Behavior Learning for Drug Repositioning
by: Zhang, Renye, et al.
Published: (2025) -
Offline Behavior Distillation
by: Lei, Shiye, et al.
Published: (2024)