Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Wenshuo, Zhu, Qi, Zeng, Xingshan, Mi, Fei, Shang, Lifeng, R., Yi, Fung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
SELF: Self-Evolution with Language Feedback
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025)
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
von: Karlekar, Sweta, et al.
Veröffentlicht: (2026)
von: Karlekar, Sweta, et al.
Veröffentlicht: (2026)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
von: Cui, Yingqian, et al.
Veröffentlicht: (2025)
von: Cui, Yingqian, et al.
Veröffentlicht: (2025)
Inference-Time Scaling for Generalist Reward Modeling
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
Reinforcement Learning Teachers of Test Time Scaling
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2025)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
YODA: Teacher-Student Progressive Learning for Language Models
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection
von: Yu, Xiao, et al.
Veröffentlicht: (2023)
von: Yu, Xiao, et al.
Veröffentlicht: (2023)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
von: Chen, Peter Baile, et al.
Veröffentlicht: (2025)
von: Chen, Peter Baile, et al.
Veröffentlicht: (2025)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
It's Not That Simple. An Analysis of Simple Test-Time Scaling
von: Wu, Guojun
Veröffentlicht: (2025)
von: Wu, Guojun
Veröffentlicht: (2025)
On the Role of Temperature Sampling in Test-Time Scaling
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Crosslingual Reasoning through Test-Time Scaling
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
von: Tejaswi, Atula, et al.
Veröffentlicht: (2026)
von: Tejaswi, Atula, et al.
Veröffentlicht: (2026)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
von: Wang, Zige, et al.
Veröffentlicht: (2025)
von: Wang, Zige, et al.
Veröffentlicht: (2025)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
von: Guo, Jizhou, et al.
Veröffentlicht: (2025)
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
von: Monjur, Ocean, et al.
Veröffentlicht: (2026)
Mode-Conditioning Unlocks Superior Test-Time Scaling
von: Wu, Chen Henry, et al.
Veröffentlicht: (2025)
von: Wu, Chen Henry, et al.
Veröffentlicht: (2025)
Parallel Test-Time Scaling for Latent Reasoning Models
von: You, Runyang, et al.
Veröffentlicht: (2025)
von: You, Runyang, et al.
Veröffentlicht: (2025)
Atom of Thoughts for Markov LLM Test-Time Scaling
von: Teng, Fengwei, et al.
Veröffentlicht: (2025)
von: Teng, Fengwei, et al.
Veröffentlicht: (2025)
Iterative Deepening Sampling as Efficient Test-Time Scaling
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
Efficient Test-Time Scaling via Self-Calibration
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
von: Cao, Lang, et al.
Veröffentlicht: (2025)
von: Cao, Lang, et al.
Veröffentlicht: (2025)
What Scales in Cross-Entropy Scaling Law?
von: Yan, Junxi, et al.
Veröffentlicht: (2025)
von: Yan, Junxi, et al.
Veröffentlicht: (2025)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
von: Liu, Shang, et al.
Veröffentlicht: (2024)
von: Liu, Shang, et al.
Veröffentlicht: (2024)
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation
von: Bin, Yi, et al.
Veröffentlicht: (2025)
von: Bin, Yi, et al.
Veröffentlicht: (2025)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
von: Wang, Youkang, et al.
Veröffentlicht: (2026)
von: Wang, Youkang, et al.
Veröffentlicht: (2026)
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
von: Chen, Yanxi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
von: Zeng, Xingshan, et al.
Veröffentlicht: (2025) -
SELF: Self-Evolution with Language Feedback
von: Lu, Jianqiao, et al.
Veröffentlicht: (2023) -
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
von: Huang, Yuzhen, et al.
Veröffentlicht: (2025) -
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025) -
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)