Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Zhenglin, Yu, Xingrui, Bossens, David Mark, Lyu, Yueming, Guo, Qing, Fan, Flint Xiaofeng, Ong, Yew Soon, Tsang, Ivor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)
by: Yu, Xingrui, et al.
Published: (2024)
Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field
by: Tan, Kim Yong, et al.
Published: (2026)
by: Tan, Kim Yong, et al.
Published: (2026)
Fast Direct: Query-Efficient Online Black-box Guidance for Diffusion-model Target Generation
by: Tan, Kim Yong, et al.
Published: (2025)
by: Tan, Kim Yong, et al.
Published: (2025)
Distributional Multi-objective Black-box Optimization for Diffusion-model Inference-time Multi-Target Generation
by: Tan, Kim Yong, et al.
Published: (2025)
by: Tan, Kim Yong, et al.
Published: (2025)
The Digital Ecosystem of Beliefs: does evolution favour AI over humans?
by: Bossens, David M., et al.
Published: (2024)
by: Bossens, David M., et al.
Published: (2024)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol
by: Fan, Flint Xiaofeng, et al.
Published: (2026)
by: Fan, Flint Xiaofeng, et al.
Published: (2026)
Covariance-Adaptive Sequential Black-box Optimization for Diffusion Targeted Generation
by: Lyu, Yueming, et al.
Published: (2024)
by: Lyu, Yueming, et al.
Published: (2024)
Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs
by: Fan, Flint Xiaofeng, et al.
Published: (2025)
by: Fan, Flint Xiaofeng, et al.
Published: (2025)
FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
by: Fan, Flint Xiaofeng, et al.
Published: (2024)
by: Fan, Flint Xiaofeng, et al.
Published: (2024)
Policy Dispersion in Non-Markovian Environment
by: Qu, Bohao, et al.
Published: (2023)
by: Qu, Bohao, et al.
Published: (2023)
A Simple Yet Effective Approach for Diversified Session-Based Recommendation
by: Yin, Qing, et al.
Published: (2024)
by: Yin, Qing, et al.
Published: (2024)
Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
by: Wu, Jingxuan, et al.
Published: (2025)
by: Wu, Jingxuan, et al.
Published: (2025)
Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models
by: Wu, Jingxuan, et al.
Published: (2026)
by: Wu, Jingxuan, et al.
Published: (2026)
Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models
by: Sun, Ming, et al.
Published: (2026)
by: Sun, Ming, et al.
Published: (2026)
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
by: Wan, Zhenglin, et al.
Published: (2025)
by: Wan, Zhenglin, et al.
Published: (2025)
Lifting Traces to Logic: Programmatic Skill Induction with Neuro-Symbolic Learning for Long-Horizon Agentic Tasks
by: Shao, Jie-Jing, et al.
Published: (2026)
by: Shao, Jie-Jing, et al.
Published: (2026)
Lang-PINN: From Language to Physics-Informed Neural Networks via a Multi-Agent Framework
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions
by: He, Xin, et al.
Published: (2026)
by: He, Xin, et al.
Published: (2026)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
by: Pan, Xu, et al.
Published: (2026)
by: Pan, Xu, et al.
Published: (2026)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
Towards Harmless Rawlsian Fairness Regardless of Demographic Prior
by: Wang, Xuanqian, et al.
Published: (2024)
by: Wang, Xuanqian, et al.
Published: (2024)
Automated Large-scale CVRP Solver Design via LLM-assisted Flexible MCTS
by: Guo, Tong, et al.
Published: (2026)
by: Guo, Tong, et al.
Published: (2026)
Adversarial Dual On-Policy Distillation from Expressive Teacher
by: Wan, Zhenglin, et al.
Published: (2026)
by: Wan, Zhenglin, et al.
Published: (2026)
RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Where to Move Next: Zero-shot Generalization of LLMs for Next POI Recommendation
by: Feng, Shanshan, et al.
Published: (2024)
by: Feng, Shanshan, et al.
Published: (2024)
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
by: Bossens, David M.
Published: (2023)
by: Bossens, David M.
Published: (2023)
Learning with Foresight: Enhancing Neural Routing Policy via Multi-Node Lookahead Prediction
by: Jiang, Xia, et al.
Published: (2026)
by: Jiang, Xia, et al.
Published: (2026)
MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming
by: Zheng, Chengqi, et al.
Published: (2025)
by: Zheng, Chengqi, et al.
Published: (2025)
Reflection-Driven Control for Trustworthy Code Agents
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
by: Yan, Xin, et al.
Published: (2026)
by: Yan, Xin, et al.
Published: (2026)
Multi-Task Learning with Multi-Task Optimization
by: Bai, Lu, et al.
Published: (2024)
by: Bai, Lu, et al.
Published: (2024)
Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
by: Cao, Xiaofeng, et al.
Published: (2025)
by: Cao, Xiaofeng, et al.
Published: (2025)
IKNO: Infinite-order Kernel Neural Operators
by: Zhu, Pengyuan, et al.
Published: (2026)
by: Zhu, Pengyuan, et al.
Published: (2026)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025)
by: Zhang, Chubin, et al.
Published: (2025)
LLM-to-Phy3D: Physically Conform Online 3D Object Generation with LLMs
by: Wong, Melvin, et al.
Published: (2025)
by: Wong, Melvin, et al.
Published: (2025)
Mastering Continual Reinforcement Learning through Fine-Grained Sparse Network Allocation and Dormant Neuron Exploration
by: Zheng, Chengqi, et al.
Published: (2025)
by: Zheng, Chengqi, et al.
Published: (2025)
Transductive Reward Inference on Graph
by: Qu, Bohao, et al.
Published: (2024)
by: Qu, Bohao, et al.
Published: (2024)
A Continuous Encoding-Based Representation for Efficient Multi-Fidelity Multi-Objective Neural Architecture Search
by: Wei, Zhao, et al.
Published: (2025)
by: Wei, Zhao, et al.
Published: (2025)
Similar Items
-
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024) -
Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field
by: Tan, Kim Yong, et al.
Published: (2026) -
Fast Direct: Query-Efficient Online Black-box Guidance for Diffusion-model Target Generation
by: Tan, Kim Yong, et al.
Published: (2025) -
Distributional Multi-objective Black-box Optimization for Diffusion-model Inference-time Multi-Target Generation
by: Tan, Kim Yong, et al.
Published: (2025) -
The Digital Ecosystem of Beliefs: does evolution favour AI over humans?
by: Bossens, David M., et al.
Published: (2024)