The Art of Efficient Reasoning: Data, Reward, and Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Taiqiang, Xu, Zenan, Zhou, Bo, Wong, Ngai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revisiting Model Interpolation for Efficient Reasoning
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
Timber: Training-free Instruct Model Refining with Base via Effective Rank
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024)
di: Yang, Runming, et al.
Pubblicazione: (2024)
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
di: Hu, Minda, et al.
Pubblicazione: (2026)
di: Hu, Minda, et al.
Pubblicazione: (2026)
Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
di: Guan, Xin, et al.
Pubblicazione: (2026)
di: Guan, Xin, et al.
Pubblicazione: (2026)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
di: Li, Zhen, et al.
Pubblicazione: (2025)
di: Li, Zhen, et al.
Pubblicazione: (2025)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
di: Sun, Wei, et al.
Pubblicazione: (2025)
di: Sun, Wei, et al.
Pubblicazione: (2025)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
di: Fan, Zhiting, et al.
Pubblicazione: (2026)
di: Fan, Zhiting, et al.
Pubblicazione: (2026)
A Survey on the Honesty of Large Language Models
di: Li, Siheng, et al.
Pubblicazione: (2024)
di: Li, Siheng, et al.
Pubblicazione: (2024)
InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning
di: Wei, Chengwei, et al.
Pubblicazione: (2026)
di: Wei, Chengwei, et al.
Pubblicazione: (2026)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
di: Xiong, Xuan, et al.
Pubblicazione: (2026)
di: Xiong, Xuan, et al.
Pubblicazione: (2026)
Prototypical Reward Network for Data-Efficient RLHF
di: Zhang, Jinghan, et al.
Pubblicazione: (2024)
di: Zhang, Jinghan, et al.
Pubblicazione: (2024)
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
di: Yang, Cehao, et al.
Pubblicazione: (2025)
di: Yang, Cehao, et al.
Pubblicazione: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
di: Sanders, Kate, et al.
Pubblicazione: (2026)
di: Sanders, Kate, et al.
Pubblicazione: (2026)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
di: Wang, Rui, et al.
Pubblicazione: (2025)
di: Wang, Rui, et al.
Pubblicazione: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
di: Jiang, Fengqing, et al.
Pubblicazione: (2024)
di: Jiang, Fengqing, et al.
Pubblicazione: (2024)
ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
di: Zhang, Qing, et al.
Pubblicazione: (2025)
di: Zhang, Qing, et al.
Pubblicazione: (2025)
Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards
di: Han, Tianyang, et al.
Pubblicazione: (2026)
di: Han, Tianyang, et al.
Pubblicazione: (2026)
Exploring Reasoning Reward Model for Agents
di: Fan, Kaixuan, et al.
Pubblicazione: (2026)
di: Fan, Kaixuan, et al.
Pubblicazione: (2026)
Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration
di: He, Bowei, et al.
Pubblicazione: (2026)
di: He, Bowei, et al.
Pubblicazione: (2026)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
di: Wen, Xumeng, et al.
Pubblicazione: (2025)
di: Wen, Xumeng, et al.
Pubblicazione: (2025)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
di: Yuan, Wenhao, et al.
Pubblicazione: (2026)
di: Yuan, Wenhao, et al.
Pubblicazione: (2026)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
di: Liu, Tao, et al.
Pubblicazione: (2026)
di: Liu, Tao, et al.
Pubblicazione: (2026)
Mixture-of-Subspaces in Low-Rank Adaptation
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning
di: Han, Jiuzhou, et al.
Pubblicazione: (2025)
di: Han, Jiuzhou, et al.
Pubblicazione: (2025)
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
di: Wu, Feijie, et al.
Pubblicazione: (2025)
di: Wu, Feijie, et al.
Pubblicazione: (2025)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
di: Zhang, Zhenru, et al.
Pubblicazione: (2025)
di: Zhang, Zhenru, et al.
Pubblicazione: (2025)
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
di: Yang, Wen, et al.
Pubblicazione: (2025)
di: Yang, Wen, et al.
Pubblicazione: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
di: Chen, Sirui, et al.
Pubblicazione: (2026)
di: Chen, Sirui, et al.
Pubblicazione: (2026)
UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward
di: Liu, Yile, et al.
Pubblicazione: (2026)
di: Liu, Yile, et al.
Pubblicazione: (2026)
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
di: Tian, Wei, et al.
Pubblicazione: (2026)
di: Tian, Wei, et al.
Pubblicazione: (2026)
ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
di: Chen, Bin, et al.
Pubblicazione: (2025)
di: Chen, Bin, et al.
Pubblicazione: (2025)
Quantifying Compositionality of Classic and State-of-the-Art Embeddings
di: Guo, Zhijin, et al.
Pubblicazione: (2025)
di: Guo, Zhijin, et al.
Pubblicazione: (2025)
Enhancing LLM Reasoning with Reward-guided Tree Search
di: Jiang, Jinhao, et al.
Pubblicazione: (2024)
di: Jiang, Jinhao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Revisiting Model Interpolation for Efficient Reasoning
di: Wu, Taiqiang, et al.
Pubblicazione: (2025) -
Timber: Training-free Instruct Model Refining with Base via Effective Rank
di: Wu, Taiqiang, et al.
Pubblicazione: (2025) -
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
di: Wu, Taiqiang, et al.
Pubblicazione: (2024) -
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024) -
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
di: Hu, Minda, et al.
Pubblicazione: (2026)