Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Hanlin, Yao, Jian, Liu, Weiming, Wang, Qing, Qin, Hanmin, Kong, Hansheng, Tang, Kirk, Xiong, Jiechao, Yu, Chao, Li, Kai, Xing, Junliang, Chen, Hongwu, Zhuo, Juchao, Fu, Qiang, Wei, Yang, Fu, Haobo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
by: Xu, Hang, et al.
Published: (2024)
by: Xu, Hang, et al.
Published: (2024)
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
by: He, Jinmin, et al.
Published: (2025)
by: He, Jinmin, et al.
Published: (2025)
Deep (Predictive) Discounted Counterfactual Regret Minimization
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
by: He, Jinmin, et al.
Published: (2023)
by: He, Jinmin, et al.
Published: (2023)
Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
by: He, Jinmin, et al.
Published: (2025)
by: He, Jinmin, et al.
Published: (2025)
Divergence-Augmented Policy Optimization
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Diversity from Human Feedback
by: Wang, Ren-Jian, et al.
Published: (2023)
by: Wang, Ren-Jian, et al.
Published: (2023)
Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information
by: Yang, Jinghan, et al.
Published: (2025)
by: Yang, Jinghan, et al.
Published: (2025)
MARPO: A Reflective Policy Optimization for Multi Agent Reinforcement Learning
by: Wu, Cuiling, et al.
Published: (2025)
by: Wu, Cuiling, et al.
Published: (2025)
A wideband transmitarray antenna based on polarization conversion metasurface with 2‐bit phase compensation
by: Hanmin Zhang, et al.
Published: (2024)
by: Hanmin Zhang, et al.
Published: (2024)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Rethinking Adversarial Inverse Reinforcement Learning: Policy Imitation, Transferable Reward Recovery and Algebraic Equilibrium Proof
by: Zhang, Yangchun, et al.
Published: (2024)
by: Zhang, Yangchun, et al.
Published: (2024)
A note about why deep learning is deep: A discontinuous approximation perspective
by: Yongxin Li, et al.
Published: (2024)
by: Yongxin Li, et al.
Published: (2024)
Proper Dataset Valuation by Pointwise Mutual Information
by: Zheng, Shuran, et al.
Published: (2024)
by: Zheng, Shuran, et al.
Published: (2024)
On the Properties and Estimation of Pointwise Mutual Information Profiles
by: Czyż, Paweł, et al.
Published: (2023)
by: Czyż, Paweł, et al.
Published: (2023)
Social Imitation Dynamics of Vaccination Driven by Vaccine Effectiveness and Beliefs
by: Fu, Feng, et al.
Published: (2025)
by: Fu, Feng, et al.
Published: (2025)
pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
by: Chen, Hansheng, et al.
Published: (2025)
by: Chen, Hansheng, et al.
Published: (2025)
Toward Emergent Holism: A Mutually Constitutive Account for Systems Science and Holistic Philosophy
by: Qiang Fu, et al.
Published: (2026)
by: Qiang Fu, et al.
Published: (2026)
Enhance Reasoning for Large Language Models in the Game Werewolf
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
by: Wang, Zhengcheng, et al.
Published: (2025)
by: Wang, Zhengcheng, et al.
Published: (2025)
Maximum Entropy Heterogeneous-Agent Reinforcement Learning
by: Liu, Jiarong, et al.
Published: (2023)
by: Liu, Jiarong, et al.
Published: (2023)
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
Rethinking Mutual Information for Language Conditioned Skill Discovery on Imitation Learning
by: Ju, Zhaoxun, et al.
Published: (2024)
by: Ju, Zhaoxun, et al.
Published: (2024)
Circular economy meets building automation
by: Cai, Hanmin
Published: (2023)
by: Cai, Hanmin
Published: (2023)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
by: Li, Jiangnan, et al.
Published: (2025)
by: Li, Jiangnan, et al.
Published: (2025)
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)
by: Yu, Xingrui, et al.
Published: (2024)
FlowHFT: Imitation Learning via Flow Matching Policy for Optimal High-Frequency Trading under Diverse Market Conditions
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
MADECM: A Curiosity‐Augmented Evolutionary Algorithm for Multi‐Agent Policy Diversity Optimization
by: Jianyang Wu, et al.
Published: (2026)
by: Jianyang Wu, et al.
Published: (2026)
MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
FedMKT: Federated Mutual Knowledge Transfer for Large and Small Language Models
by: Fan, Tao, et al.
Published: (2024)
by: Fan, Tao, et al.
Published: (2024)
Wolff potential estimates for elliptic obstacle problems with generalized Orlicz growth
by: Xiong, Qi, et al.
Published: (2025)
by: Xiong, Qi, et al.
Published: (2025)
Heterogeneous Multi-agent Zero-Shot Coordination by Coevolution
by: Xue, Ke, et al.
Published: (2022)
by: Xue, Ke, et al.
Published: (2022)
Reflective Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
On approximating the $f$-divergence between two Ising models
by: Feng, Weiming, et al.
Published: (2025)
by: Feng, Weiming, et al.
Published: (2025)
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
by: Zhao, Guangyu, et al.
Published: (2024)
by: Zhao, Guangyu, et al.
Published: (2024)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
by: Jia, Nan, et al.
Published: (2026)
by: Jia, Nan, et al.
Published: (2026)
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
by: Xiong, Kai, et al.
Published: (2026)
by: Xiong, Kai, et al.
Published: (2026)
Comparison of Encryption Algorithms for Wearable Devices in IoT Systems
by: Yang, Haobo
Published: (2024)
by: Yang, Haobo
Published: (2024)
ACOCMPMI: An Ant Colony Optimization Algorithm Based on Composite Multiscale Part Mutual Information for Detecting Epistatic Interactions
by: Yan Sun, et al.
Published: (2025)
by: Yan Sun, et al.
Published: (2025)
Averages with the Gaussian divisor: Weighted Inequalities and the Pointwise Ergodic Theorem
by: Giannitsi, Christina, et al.
Published: (2024)
by: Giannitsi, Christina, et al.
Published: (2024)
Similar Items
-
Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
by: Xu, Hang, et al.
Published: (2024) -
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
by: He, Jinmin, et al.
Published: (2025) -
Deep (Predictive) Discounted Counterfactual Regret Minimization
by: Xu, Hang, et al.
Published: (2025) -
Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
by: He, Jinmin, et al.
Published: (2023) -
Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
by: He, Jinmin, et al.
Published: (2025)