DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Tian, Jiao, Wenxiang, He, Zhiwei, Xu, Jiahao, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025)
by: Zhang, Ziyin, et al.
Published: (2025)
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
by: He, Zhiwei, et al.
Published: (2025)
by: He, Zhiwei, et al.
Published: (2025)
CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
by: Fan, Zhenxuan, et al.
Published: (2026)
by: Fan, Zhenxuan, et al.
Published: (2026)
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
by: He, Zhiwei, et al.
Published: (2024)
by: He, Zhiwei, et al.
Published: (2024)
Free(): Learning to Forget in Malloc-Only Reasoning Models
by: Zheng, Yilun, et al.
Published: (2026)
by: Zheng, Yilun, et al.
Published: (2026)
HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows
by: Yao, Wenlin, et al.
Published: (2024)
by: Yao, Wenlin, et al.
Published: (2024)
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
by: Hu, Minda, et al.
Published: (2026)
by: Hu, Minda, et al.
Published: (2026)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
by: Yue, Murong, et al.
Published: (2024)
by: Yue, Murong, et al.
Published: (2024)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
by: Panaganti, Kishan, et al.
Published: (2026)
by: Panaganti, Kishan, et al.
Published: (2026)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
by: Li, Mukai, et al.
Published: (2025)
by: Li, Mukai, et al.
Published: (2025)
Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
by: Li, Chengzhengxu, et al.
Published: (2025)
by: Li, Chengzhengxu, et al.
Published: (2025)
CoT-Valve: Length-Compressible Chain-of-Thought Tuning
by: Ma, Xinyin, et al.
Published: (2025)
by: Ma, Xinyin, et al.
Published: (2025)
Self-Consistency Boosts Calibration for Math Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework
by: Huang, Kerui, et al.
Published: (2025)
by: Huang, Kerui, et al.
Published: (2025)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
by: Xiong, Xuan, et al.
Published: (2026)
by: Xiong, Xuan, et al.
Published: (2026)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Exploring Reasoning Reward Model for Agents
by: Fan, Kaixuan, et al.
Published: (2026)
by: Fan, Kaixuan, et al.
Published: (2026)
SmartThinker: Learning to Compress and Preserve Reasoning by Step-Level Length Control
by: He, Xingyang, et al.
Published: (2025)
by: He, Xingyang, et al.
Published: (2025)
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
by: Tian, Wei, et al.
Published: (2026)
by: Tian, Wei, et al.
Published: (2026)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
Optimizing Length Compression in Large Reasoning Models
by: Cheng, Zhengxiang, et al.
Published: (2025)
by: Cheng, Zhengxiang, et al.
Published: (2025)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
by: Wang, Weilan, et al.
Published: (2025)
by: Wang, Weilan, et al.
Published: (2025)
Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decoding
by: Liu, Tianqiao, et al.
Published: (2024)
by: Liu, Tianqiao, et al.
Published: (2024)
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
by: Shan, Hongyu, et al.
Published: (2025)
by: Shan, Hongyu, et al.
Published: (2025)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
by: Turpin, Miles, et al.
Published: (2025)
by: Turpin, Miles, et al.
Published: (2025)
Reinforcement Learning for Chain of Thought Compression with One-Domain-to-All Generalization
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
DRT: Deep Reasoning Translation via Long Chain-of-Thought
by: Wang, Jiaan, et al.
Published: (2024)
by: Wang, Jiaan, et al.
Published: (2024)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
by: Cai, Zefan, et al.
Published: (2025)
by: Cai, Zefan, et al.
Published: (2025)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
by: You, Wangjie, et al.
Published: (2025)
by: You, Wangjie, et al.
Published: (2025)
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
by: Chen, Zhuoen, et al.
Published: (2026)
by: Chen, Zhuoen, et al.
Published: (2026)
Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning
by: Chen, Yiqun, et al.
Published: (2026)
by: Chen, Yiqun, et al.
Published: (2026)
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression
by: Choi, Eunseong, et al.
Published: (2024)
by: Choi, Eunseong, et al.
Published: (2024)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
by: Hu, Jinwu, et al.
Published: (2025)
by: Hu, Jinwu, et al.
Published: (2025)
Similar Items
-
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025) -
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025) -
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
by: He, Zhiwei, et al.
Published: (2025) -
CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
by: Fan, Zhenxuan, et al.
Published: (2026) -
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
by: He, Zhiwei, et al.
Published: (2024)