PEAR: Phase Entropy Aware Reward for Efficient Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Chen, Lu, Wei, Zhang, Wenxuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
di: Yang, Shidong, et al.
Pubblicazione: (2026)
di: Yang, Shidong, et al.
Pubblicazione: (2026)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
di: Xiong, Xuan, et al.
Pubblicazione: (2026)
di: Xiong, Xuan, et al.
Pubblicazione: (2026)
InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning
di: Wei, Chengwei, et al.
Pubblicazione: (2026)
di: Wei, Chengwei, et al.
Pubblicazione: (2026)
Efficient Reasoning via Reward Model
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
Promoting Efficient Reasoning with Verifiable Stepwise Reward
di: Yue, Chuhuai, et al.
Pubblicazione: (2025)
di: Yue, Chuhuai, et al.
Pubblicazione: (2025)
Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
di: Li, Yanhao, et al.
Pubblicazione: (2025)
di: Li, Yanhao, et al.
Pubblicazione: (2025)
PEAR: Pixel-aligned Expressive humAn mesh Recovery
di: Wu, Jiahao, et al.
Pubblicazione: (2026)
di: Wu, Jiahao, et al.
Pubblicazione: (2026)
Strikingness-Aware Evaluation for Temporal Knowledge Graph Reasoning
di: Huang, Rikui, et al.
Pubblicazione: (2026)
di: Huang, Rikui, et al.
Pubblicazione: (2026)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
di: Zhang, Yao, et al.
Pubblicazione: (2025)
di: Zhang, Yao, et al.
Pubblicazione: (2025)
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
di: Yang, Yankai, et al.
Pubblicazione: (2026)
di: Yang, Yankai, et al.
Pubblicazione: (2026)
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead
di: Tan, Tao, et al.
Pubblicazione: (2024)
di: Tan, Tao, et al.
Pubblicazione: (2024)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
di: Su, Tiancheng, et al.
Pubblicazione: (2025)
di: Su, Tiancheng, et al.
Pubblicazione: (2025)
Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
di: Cao, Hongye, et al.
Pubblicazione: (2025)
di: Cao, Hongye, et al.
Pubblicazione: (2025)
RHyVE: Competence-Aware Verification and Phase-Aware Deployment for LLM-Generated Reward Hypotheses
di: Wu, Feiyu, et al.
Pubblicazione: (2026)
di: Wu, Feiyu, et al.
Pubblicazione: (2026)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
di: Li, Gang, et al.
Pubblicazione: (2025)
di: Li, Gang, et al.
Pubblicazione: (2025)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
di: Sun, Wei, et al.
Pubblicazione: (2025)
di: Sun, Wei, et al.
Pubblicazione: (2025)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
The Art of Efficient Reasoning: Data, Reward, and Optimization
di: Wu, Taiqiang, et al.
Pubblicazione: (2026)
di: Wu, Taiqiang, et al.
Pubblicazione: (2026)
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
di: Zabounidis, Renos, et al.
Pubblicazione: (2025)
di: Zabounidis, Renos, et al.
Pubblicazione: (2025)
Process Rewards with Learned Reliability
di: Li, Jinyuan, et al.
Pubblicazione: (2026)
di: Li, Jinyuan, et al.
Pubblicazione: (2026)
Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language Models
di: Liu, Yan, et al.
Pubblicazione: (2026)
di: Liu, Yan, et al.
Pubblicazione: (2026)
Hear the Heartbeat in Phases: Physiologically Grounded Phase-Aware ECG Biometrics
di: Huang, Jintao, et al.
Pubblicazione: (2026)
di: Huang, Jintao, et al.
Pubblicazione: (2026)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
di: Chen, Xinzhu, et al.
Pubblicazione: (2025)
di: Chen, Xinzhu, et al.
Pubblicazione: (2025)
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
di: Kim, Yoonjeon, et al.
Pubblicazione: (2025)
di: Kim, Yoonjeon, et al.
Pubblicazione: (2025)
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
di: Zhao, Qiannian, et al.
Pubblicazione: (2026)
di: Zhao, Qiannian, et al.
Pubblicazione: (2026)
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
di: Tian, Wei, et al.
Pubblicazione: (2026)
di: Tian, Wei, et al.
Pubblicazione: (2026)
iPEAR: Iterative Pyramid Estimation with Attention and Residuals for Deformable Medical Image Registration
di: Wu, Heming, et al.
Pubblicazione: (2025)
di: Wu, Heming, et al.
Pubblicazione: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
Verifiable Process Rewards for Agentic Reasoning
di: Yuan, Huining, et al.
Pubblicazione: (2026)
di: Yuan, Huining, et al.
Pubblicazione: (2026)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
di: Chen, Sirui, et al.
Pubblicazione: (2026)
di: Chen, Sirui, et al.
Pubblicazione: (2026)
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward
di: Hussain, Mustafa Anis, et al.
Pubblicazione: (2026)
di: Hussain, Mustafa Anis, et al.
Pubblicazione: (2026)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
di: Yu, Zishun, et al.
Pubblicazione: (2025)
di: Yu, Zishun, et al.
Pubblicazione: (2025)
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
di: Xie, Shaoan, et al.
Pubblicazione: (2025)
di: Xie, Shaoan, et al.
Pubblicazione: (2025)
ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference
di: Wang, Junda, et al.
Pubblicazione: (2026)
di: Wang, Junda, et al.
Pubblicazione: (2026)
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
di: Zhang, Xiangxiang, et al.
Pubblicazione: (2025)
di: Zhang, Xiangxiang, et al.
Pubblicazione: (2025)
Entropy-Gated Branching for Efficient Test-Time Reasoning
di: Li, Xianzhi, et al.
Pubblicazione: (2025)
di: Li, Xianzhi, et al.
Pubblicazione: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
di: Yang, Shidong, et al.
Pubblicazione: (2026) -
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
di: Xiong, Xuan, et al.
Pubblicazione: (2026) -
InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning
di: Wei, Chengwei, et al.
Pubblicazione: (2026) -
Efficient Reasoning via Reward Model
di: Wang, Yuhao, et al.
Pubblicazione: (2025) -
Promoting Efficient Reasoning with Verifiable Stepwise Reward
di: Yue, Chuhuai, et al.
Pubblicazione: (2025)