ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Jia, He, ChangYi, Lin, YingQiao, Yang, MingMin, Shen, FeiYang, Liu, ShaoGuo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
por: Yang, Pinci, et al.
Publicado: (2025)
por: Yang, Pinci, et al.
Publicado: (2025)
Mechanical Properties and Deformation Behavior of a Novel 3D Printed Tubular TPMS Structure
por: ShaoGuo Zhang, et al.
Publicado: (2026)
por: ShaoGuo Zhang, et al.
Publicado: (2026)
Chain of Time: In-Context Physical Simulation with Image Generation Models
por: Wang, YingQiao, et al.
Publicado: (2025)
por: Wang, YingQiao, et al.
Publicado: (2025)
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
por: Xu, Fangzhi, et al.
Publicado: (2025)
por: Xu, Fangzhi, et al.
Publicado: (2025)
The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective
por: Yan, Renye, et al.
Publicado: (2024)
por: Yan, Renye, et al.
Publicado: (2024)
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
por: Nguyen, Ha-Thanh, et al.
Publicado: (2024)
por: Nguyen, Ha-Thanh, et al.
Publicado: (2024)
Query Decomposition for RAG: Balancing Exploration-Exploitation
por: Petcu, Roxana, et al.
Publicado: (2025)
por: Petcu, Roxana, et al.
Publicado: (2025)
Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
por: Niu, Zenghao, et al.
Publicado: (2025)
por: Niu, Zenghao, et al.
Publicado: (2025)
In-context Exploration-Exploitation for Reinforcement Learning
por: Dai, Zhenwen, et al.
Publicado: (2024)
por: Dai, Zhenwen, et al.
Publicado: (2024)
Exploration, Exploitation, and Organizational Coordination Mechanisms
por: Silvio Popadiuk
Publicado: (2016)
por: Silvio Popadiuk
Publicado: (2016)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
por: Huang, Xu, et al.
Publicado: (2024)
por: Huang, Xu, et al.
Publicado: (2024)
Scale-Adaptive Balancing of Exploration and Exploitation in Classical Planning
por: Wissow, Stephen, et al.
Publicado: (2023)
por: Wissow, Stephen, et al.
Publicado: (2023)
SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation
por: Chen, Yixiang, et al.
Publicado: (2025)
por: Chen, Yixiang, et al.
Publicado: (2025)
Landmark Guided Active Exploration with State-specific Balance Coefficient
por: Cui, Fei, et al.
Publicado: (2023)
por: Cui, Fei, et al.
Publicado: (2023)
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
por: Yang, Lu, et al.
Publicado: (2026)
por: Yang, Lu, et al.
Publicado: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
por: Huang, Fanding, et al.
Publicado: (2025)
por: Huang, Fanding, et al.
Publicado: (2025)
ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection
por: Gao, Changjiang, et al.
Publicado: (2026)
por: Gao, Changjiang, et al.
Publicado: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
por: Liang, Kun, et al.
Publicado: (2026)
por: Liang, Kun, et al.
Publicado: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
por: Chen, Peter, et al.
Publicado: (2025)
por: Chen, Peter, et al.
Publicado: (2025)
A Goal-Oriented Approach for Active Object Detection with Exploration-Exploitation Balance
por: Yu, Yalei, et al.
Publicado: (2025)
por: Yu, Yalei, et al.
Publicado: (2025)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
por: Chen, Zhipeng, et al.
Publicado: (2025)
por: Chen, Zhipeng, et al.
Publicado: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
por: Zeng, Weihao, et al.
Publicado: (2024)
por: Zeng, Weihao, et al.
Publicado: (2024)
A Balanced Approach of Rapid Genetic Exploration and Surrogate Exploitation for Hyperparameter Optimization
por: Kim, Chul, et al.
Publicado: (2025)
por: Kim, Chul, et al.
Publicado: (2025)
Energy Exploration & Exploitation
Publicado: (2020)
Publicado: (2020)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
por: Su, Xuerui, et al.
Publicado: (2025)
por: Su, Xuerui, et al.
Publicado: (2025)
From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
por: Xu, Donglai, et al.
Publicado: (2025)
por: Xu, Donglai, et al.
Publicado: (2025)
Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action
por: Gao, Gong, et al.
Publicado: (2026)
por: Gao, Gong, et al.
Publicado: (2026)
Engineered Protein Fibers with Reinforced Mechanical Properties Via β‐Sheet High‐Order Assembly
por: Ming Li, et al.
Publicado: (2024)
por: Ming Li, et al.
Publicado: (2024)
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
por: Liu, Yujian, et al.
Publicado: (2025)
por: Liu, Yujian, et al.
Publicado: (2025)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
por: Zhao, Chu, et al.
Publicado: (2026)
por: Zhao, Chu, et al.
Publicado: (2026)
MRSO: Balancing Exploration and Exploitation through Modified Rat Swarm Optimization for Global Optimization
por: Abdulla, Hemin Sardar, et al.
Publicado: (2024)
por: Abdulla, Hemin Sardar, et al.
Publicado: (2024)
EEA: Exploration-Exploitation Agent for Long Video Understanding
por: Yang, Te, et al.
Publicado: (2025)
por: Yang, Te, et al.
Publicado: (2025)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
por: Yao, Jiashu, et al.
Publicado: (2026)
por: Yao, Jiashu, et al.
Publicado: (2026)
Real-Time Auto-Optimization in Unknown Environments via Structure-Exploiting Dual Control for Exploration and Exploitation
por: Dong, Shiying, et al.
Publicado: (2026)
por: Dong, Shiying, et al.
Publicado: (2026)
Disentangling Exploration from Exploitation
por: Lizzeri, Alessandro, et al.
Publicado: (2024)
por: Lizzeri, Alessandro, et al.
Publicado: (2024)
Marine Exploration and Exploitation of Hydrocarbons
por: Radovich, Violeta S.
Publicado: (2025)
por: Radovich, Violeta S.
Publicado: (2025)
Organizational Factors for Exploration and Exploitation
por: Sharadindu Pandey
Publicado: (2009)
por: Sharadindu Pandey
Publicado: (2009)
LLM-Empowered State Representation for Reinforcement Learning
por: Wang, Boyuan, et al.
Publicado: (2024)
por: Wang, Boyuan, et al.
Publicado: (2024)
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
por: Wang, Chen, et al.
Publicado: (2026)
por: Wang, Chen, et al.
Publicado: (2026)
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
por: Zhai, Naixin, et al.
Publicado: (2026)
por: Zhai, Naixin, et al.
Publicado: (2026)
Ejemplares similares
-
Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
por: Yang, Pinci, et al.
Publicado: (2025) -
Mechanical Properties and Deformation Behavior of a Novel 3D Printed Tubular TPMS Structure
por: ShaoGuo Zhang, et al.
Publicado: (2026) -
Chain of Time: In-Context Physical Simulation with Image Generation Models
por: Wang, YingQiao, et al.
Publicado: (2025) -
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
por: Xu, Fangzhi, et al.
Publicado: (2025) -
The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective
por: Yan, Renye, et al.
Publicado: (2024)