Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zabounidis, Renos, Golatkar, Aditya, Kleinman, Michael, Achille, Alessandro, Xia, Wei, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
e1: Learning Adaptive Control of Reasoning Effort
von: Kleinman, Michael, et al.
Veröffentlicht: (2025)
von: Kleinman, Michael, et al.
Veröffentlicht: (2025)
Critical Learning Periods Emerge Even in Deep Linear Networks
von: Kleinman, Michael, et al.
Veröffentlicht: (2023)
von: Kleinman, Michael, et al.
Veröffentlicht: (2023)
Training Data Protection with Compositional Diffusion Models
von: Golatkar, Aditya, et al.
Veröffentlicht: (2023)
von: Golatkar, Aditya, et al.
Veröffentlicht: (2023)
AI Agents as Universal Task Solvers
von: Achille, Alessandro, et al.
Veröffentlicht: (2025)
von: Achille, Alessandro, et al.
Veröffentlicht: (2025)
PICASO: Permutation-Invariant Context Composition with State Space Models
von: Liu, Tian Yu, et al.
Veröffentlicht: (2025)
von: Liu, Tian Yu, et al.
Veröffentlicht: (2025)
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
von: Stein, Adam, et al.
Veröffentlicht: (2025)
von: Stein, Adam, et al.
Veröffentlicht: (2025)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
von: Biggs, Benjamin, et al.
Veröffentlicht: (2024)
von: Biggs, Benjamin, et al.
Veröffentlicht: (2024)
LATTS: Locally Adaptive Test-Time Scaling
von: Uscidda, Theo, et al.
Veröffentlicht: (2025)
von: Uscidda, Theo, et al.
Veröffentlicht: (2025)
Tangent Transformers for Composition, Privacy and Removal
von: Liu, Tian Yu, et al.
Veröffentlicht: (2023)
von: Liu, Tian Yu, et al.
Veröffentlicht: (2023)
Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion Predictions
von: Zhao, Albert, et al.
Veröffentlicht: (2025)
von: Zhao, Albert, et al.
Veröffentlicht: (2025)
CPR: Retrieval Augmented Generation for Copyright Protection
von: Golatkar, Aditya, et al.
Veröffentlicht: (2024)
von: Golatkar, Aditya, et al.
Veröffentlicht: (2024)
Dynamic Chain-of-Thought: Towards Adaptive Deep Reasoning
von: Wang, Libo
Veröffentlicht: (2025)
von: Wang, Libo
Veröffentlicht: (2025)
CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction
von: Park, Jueon, et al.
Veröffentlicht: (2025)
von: Park, Jueon, et al.
Veröffentlicht: (2025)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
von: Trager, Matthew, et al.
Veröffentlicht: (2023)
von: Trager, Matthew, et al.
Veröffentlicht: (2023)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
von: Li, Xintong, et al.
Veröffentlicht: (2026)
von: Li, Xintong, et al.
Veröffentlicht: (2026)
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
von: Nunez, Elvis, et al.
Veröffentlicht: (2024)
von: Nunez, Elvis, et al.
Veröffentlicht: (2024)
Unveiling Confirmation Bias in Chain-of-Thought Reasoning
von: Wan, Yue, et al.
Veröffentlicht: (2025)
von: Wan, Yue, et al.
Veröffentlicht: (2025)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
Chain-of-Thought Predictive Control
von: Jia, Zhiwei, et al.
Veröffentlicht: (2023)
von: Jia, Zhiwei, et al.
Veröffentlicht: (2023)
Fractured Chain-of-Thought Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Priming: Hybrid State Space Models From Pre-trained Transformers
von: Chattopadhyay, Aditya, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Aditya, et al.
Veröffentlicht: (2026)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
Scaling Graph Chain-of-Thought Reasoning: A Multi-Agent Framework with Efficient LLM Serving
von: Huan, Chengying, et al.
Veröffentlicht: (2025)
von: Huan, Chengying, et al.
Veröffentlicht: (2025)
Efficient Embedding-based Synthetic Data Generation for Complex Reasoning Tasks
von: Jayaraman, Srideepika, et al.
Veröffentlicht: (2026)
von: Jayaraman, Srideepika, et al.
Veröffentlicht: (2026)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025)
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025)
Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
von: Chen, Zhikang, et al.
Veröffentlicht: (2025)
von: Chen, Zhikang, et al.
Veröffentlicht: (2025)
AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
von: Lou, Chenwei, et al.
Veröffentlicht: (2025)
von: Lou, Chenwei, et al.
Veröffentlicht: (2025)
Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
Improving Chain-of-Thought for Logical Reasoning via Attention-Aware Intervention
von: Phuong, Nguyen Minh, et al.
Veröffentlicht: (2026)
von: Phuong, Nguyen Minh, et al.
Veröffentlicht: (2026)
Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2025)
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2025)
STree: Speculative Tree Decoding for Hybrid State-Space Models
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
Scalable Chain of Thoughts via Elastic Reasoning
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
e1: Learning Adaptive Control of Reasoning Effort
von: Kleinman, Michael, et al.
Veröffentlicht: (2025) -
Critical Learning Periods Emerge Even in Deep Linear Networks
von: Kleinman, Michael, et al.
Veröffentlicht: (2023) -
Training Data Protection with Compositional Diffusion Models
von: Golatkar, Aditya, et al.
Veröffentlicht: (2023) -
AI Agents as Universal Task Solvers
von: Achille, Alessandro, et al.
Veröffentlicht: (2025) -
PICASO: Permutation-Invariant Context Composition with State Space Models
von: Liu, Tian Yu, et al.
Veröffentlicht: (2025)