Enter the Void - Planning to Seek Entropy When Reward is Scarce
Fuente:
arXiv
Saved in:
| Main Authors: | Sundar, Ashish, Luo, Chunbo, Wang, Xiaoyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
by: Lei, Jingdi, et al.
Published: (2025)
by: Lei, Jingdi, et al.
Published: (2025)
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
by: Wu, Boqian, et al.
Published: (2026)
by: Wu, Boqian, et al.
Published: (2026)
Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering
by: Ginting, Muhammad Fadhil, et al.
Published: (2025)
by: Ginting, Muhammad Fadhil, et al.
Published: (2025)
"When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
by: De, Somsubhra, et al.
Published: (2025)
by: De, Somsubhra, et al.
Published: (2025)
Scale Decoupled Distillation
by: Luo, Shicai Wei Chunbo Luo Yang
Published: (2024)
by: Luo, Shicai Wei Chunbo Luo Yang
Published: (2024)
CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
by: Liu, Zhipeng, et al.
Published: (2026)
by: Liu, Zhipeng, et al.
Published: (2026)
Robust Multimodal Learning via Representation Decoupling
by: Wei, Shicai, et al.
Published: (2024)
by: Wei, Shicai, et al.
Published: (2024)
Knowledge Localization: Mission Not Accomplished? Enter Query Localization!
by: Chen, Yuheng, et al.
Published: (2024)
by: Chen, Yuheng, et al.
Published: (2024)
Application of LLMs to Multi-Robot Path Planning and Task Allocation
by: Kumar, Ashish
Published: (2025)
by: Kumar, Ashish
Published: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Fairness Under Demographic Scarce Regime
by: Kenfack, Patrik Joslin, et al.
Published: (2023)
by: Kenfack, Patrik Joslin, et al.
Published: (2023)
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
by: Huang, Chen, et al.
Published: (2025)
by: Huang, Chen, et al.
Published: (2025)
Toto 2.0: Time Series Forecasting Enters the Scaling Era
by: Khwaja, Emaad, et al.
Published: (2026)
by: Khwaja, Emaad, et al.
Published: (2026)
PRInTS: Reward Modeling for Long-Horizon Information Seeking
by: Lee, Jaewoo, et al.
Published: (2025)
by: Lee, Jaewoo, et al.
Published: (2025)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
by: Xiong, Xuan, et al.
Published: (2026)
by: Xiong, Xuan, et al.
Published: (2026)
Causality Enhanced Origin-Destination Flow Prediction in Data-Scarce Cities
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
by: Li, Yanshi, et al.
Published: (2024)
by: Li, Yanshi, et al.
Published: (2024)
Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards
by: Lara, Luis, et al.
Published: (2026)
by: Lara, Luis, et al.
Published: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
by: Yang, Shidong, et al.
Published: (2026)
by: Yang, Shidong, et al.
Published: (2026)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
Improving Open-world Continual Learning under the Constraints of Scarce Labeled Data
by: Li, Yujie, et al.
Published: (2025)
by: Li, Yujie, et al.
Published: (2025)
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
Planning-Augmented Sampling with Early Guidance for High-Reward Discovery
by: Zhu, Rui, et al.
Published: (2025)
by: Zhu, Rui, et al.
Published: (2025)
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
by: Wang, Jiaxuan, et al.
Published: (2026)
by: Wang, Jiaxuan, et al.
Published: (2026)
Personalized Treatment Outcome Prediction from Scarce Data via Dual-Channel Knowledge Distillation and Adaptive Fusion
by: Chen, Wenjie, et al.
Published: (2025)
by: Chen, Wenjie, et al.
Published: (2025)
GOV-REK: Governed Reward Engineering Kernels for Designing Robust Multi-Agent Reinforcement Learning Systems
by: Rana, Ashish, et al.
Published: (2024)
by: Rana, Ashish, et al.
Published: (2024)
Scaling Autonomous Agents via Automatic Reward Modeling And Planning
by: Chen, Zhenfang, et al.
Published: (2025)
by: Chen, Zhenfang, et al.
Published: (2025)
Reward Bound for Behavioral Guarantee of Model-based Planning Agents
by: An, Zhiyu, et al.
Published: (2024)
by: An, Zhiyu, et al.
Published: (2024)
Achieving Equilibrium under Utility Heterogeneity: An Agent-Attention Framework for Multi-Agent Multi-Objective Reinforcement Learning
by: Li, Zhuhui, et al.
Published: (2025)
by: Li, Zhuhui, et al.
Published: (2025)
Logically Constrained Robotics Transformers for Enhanced Perception-Action Planning
by: Kapoor, Parv, et al.
Published: (2024)
by: Kapoor, Parv, et al.
Published: (2024)
Multi-View Subgraph Neural Networks: Self-Supervised Learning with Scarce Labeled Data
by: Wang, Zhenzhong, et al.
Published: (2024)
by: Wang, Zhenzhong, et al.
Published: (2024)
NanoNet: Parameter-Efficient Learning with Label-Scarce Supervision for Lightweight Text Mining Model
by: Mao, Qianren, et al.
Published: (2026)
by: Mao, Qianren, et al.
Published: (2026)
When Right Meets Wrong: Bilateral Context Conditioning with Reward-Confidence Correction for GRPO
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
by: Li, Tingyu, et al.
Published: (2025)
by: Li, Tingyu, et al.
Published: (2025)
Robust Molecular Property Prediction via Densifying Scarce Labeled Data
by: Kim, Jina, et al.
Published: (2025)
by: Kim, Jina, et al.
Published: (2025)
PMLBmini: A Tabular Classification Benchmark Suite for Data-Scarce Applications
by: Knauer, Ricardo, et al.
Published: (2024)
by: Knauer, Ricardo, et al.
Published: (2024)
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
by: Kong, Injin, et al.
Published: (2026)
by: Kong, Injin, et al.
Published: (2026)
When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models
by: Zheng, Haoran, et al.
Published: (2026)
by: Zheng, Haoran, et al.
Published: (2026)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
by: Zhao, Wenshuo, et al.
Published: (2026)
by: Zhao, Wenshuo, et al.
Published: (2026)
Similar Items
-
OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
by: Lei, Jingdi, et al.
Published: (2025) -
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
by: Wu, Boqian, et al.
Published: (2026) -
Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering
by: Ginting, Muhammad Fadhil, et al.
Published: (2025) -
"When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
by: De, Somsubhra, et al.
Published: (2025) -
Scale Decoupled Distillation
by: Luo, Shicai Wei Chunbo Luo Yang
Published: (2024)