HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhanyu, Hu, Qingguo, Wang, Ante, Liu, Chenqing, Xiang, Zhishang, Li, Hui, Qiu, Delai, Su, Jinsong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
von: Min, Zijun, et al.
Veröffentlicht: (2026)
von: Min, Zijun, et al.
Veröffentlicht: (2026)
A Multi-Agent Framework with Automated Decision Rule Optimization for Cross-Domain Misinformation Detection
von: Li, Hui, et al.
Veröffentlicht: (2025)
von: Li, Hui, et al.
Veröffentlicht: (2025)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
von: Xu, Huimin, et al.
Veröffentlicht: (2026)
von: Xu, Huimin, et al.
Veröffentlicht: (2026)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
von: Chen, Kun, et al.
Veröffentlicht: (2026)
von: Chen, Kun, et al.
Veröffentlicht: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation
von: Xiao, Meiman, et al.
Veröffentlicht: (2026)
von: Xiao, Meiman, et al.
Veröffentlicht: (2026)
From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
von: Xu, Donglai, et al.
Veröffentlicht: (2025)
von: Xu, Donglai, et al.
Veröffentlicht: (2025)
Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning
von: Wei, Shuyu, et al.
Veröffentlicht: (2026)
von: Wei, Shuyu, et al.
Veröffentlicht: (2026)
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation
von: Wu, Chuanjie, et al.
Veröffentlicht: (2026)
von: Wu, Chuanjie, et al.
Veröffentlicht: (2026)
Frequency Enhanced Pre-training for Cross-city Few-shot Traffic Forecasting
von: Liu, Zhanyu, et al.
Veröffentlicht: (2024)
von: Liu, Zhanyu, et al.
Veröffentlicht: (2024)
Entropy-Aware Structural Alignment for Zero-Shot Handwritten Chinese Character Recognition
von: Luo, Qiuming, et al.
Veröffentlicht: (2026)
von: Luo, Qiuming, et al.
Veröffentlicht: (2026)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
von: Qiu, Zhongxi, et al.
Veröffentlicht: (2025)
von: Qiu, Zhongxi, et al.
Veröffentlicht: (2025)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
von: Hu, Jiajun, et al.
Veröffentlicht: (2026)
von: Hu, Jiajun, et al.
Veröffentlicht: (2026)
Entropy-Tree: Tree-Based Decoding with Entropy-Guided Exploration
von: Wei, Longxuan, et al.
Veröffentlicht: (2026)
von: Wei, Longxuan, et al.
Veröffentlicht: (2026)
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Avoiding Premature Collapse: Adaptive Annealing for Entropy-Regularized Structural Inference
von: Liu, Yizhi
Veröffentlicht: (2026)
von: Liu, Yizhi
Veröffentlicht: (2026)
Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
von: Jiang, Tianjiao, et al.
Veröffentlicht: (2025)
von: Jiang, Tianjiao, et al.
Veröffentlicht: (2025)
Response Enhanced Semi-supervised Dialogue Query Generation
von: Huang, Jianheng, et al.
Veröffentlicht: (2023)
von: Huang, Jianheng, et al.
Veröffentlicht: (2023)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
von: Li, Qingbin, et al.
Veröffentlicht: (2025)
von: Li, Qingbin, et al.
Veröffentlicht: (2025)
TTCS: Test-Time Curriculum Synthesis for Self-Evolving
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
von: Yang, Chengyi, et al.
Veröffentlicht: (2026)
Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration
von: Wang, Ante, et al.
Veröffentlicht: (2025)
von: Wang, Ante, et al.
Veröffentlicht: (2025)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
Entropy Myth: Collapse of the Disorder Dogma
von: Shaub, Jarid Shaub
Veröffentlicht: (2025)
von: Shaub, Jarid Shaub
Veröffentlicht: (2025)
FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation
von: Zhang, Qinggang, et al.
Veröffentlicht: (2025)
von: Zhang, Qinggang, et al.
Veröffentlicht: (2025)
When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation
von: Xiang, Zhishang, et al.
Veröffentlicht: (2025)
von: Xiang, Zhishang, et al.
Veröffentlicht: (2025)
Countering the Over-Reliance Trap: Mitigating Object Hallucination for LVLMs via a Self-Validation Framework
von: Liu, Shiyu, et al.
Veröffentlicht: (2026)
von: Liu, Shiyu, et al.
Veröffentlicht: (2026)
A Dual-Perspective Metaphor Detection Framework Using Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2024)
von: Lin, Yujie, et al.
Veröffentlicht: (2024)
Enhancing Information Maximization with Distance-Aware Contrastive Learning for Source-Free Cross-Domain Few-Shot Learning
von: Xu, Huali, et al.
Veröffentlicht: (2024)
von: Xu, Huali, et al.
Veröffentlicht: (2024)
GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
von: Zhao, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2025)
Domain-Rectifying Adapter for Cross-Domain Few-Shot Segmentation
von: Su, Jiapeng, et al.
Veröffentlicht: (2024)
von: Su, Jiapeng, et al.
Veröffentlicht: (2024)
PromptAL: Sample-Aware Dynamic Soft Prompts for Few-Shot Active Learning
von: Xiang, Hui, et al.
Veröffentlicht: (2025)
von: Xiang, Hui, et al.
Veröffentlicht: (2025)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
von: Ren, Yanwei, et al.
Veröffentlicht: (2026)
von: Ren, Yanwei, et al.
Veröffentlicht: (2026)
Interpretable Cross-Domain Few-Shot Learning with Rectified Target-Domain Local Alignment
von: Zhao, Yaze, et al.
Veröffentlicht: (2026)
von: Zhao, Yaze, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
von: Hu, Qingguo, et al.
Veröffentlicht: (2025) -
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026) -
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
von: Min, Zijun, et al.
Veröffentlicht: (2026) -
A Multi-Agent Framework with Automated Decision Rule Optimization for Cross-Domain Misinformation Detection
von: Li, Hui, et al.
Veröffentlicht: (2025) -
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
von: Xu, Huimin, et al.
Veröffentlicht: (2026)