Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Jinyan, Cardie, Claire |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions
by: Su, Jinyan, et al.
Published: (2026)
by: Su, Jinyan, et al.
Published: (2026)
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023)
by: Su, Jinyan, et al.
Published: (2023)
CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization
by: Su, Jinyan, et al.
Published: (2026)
by: Su, Jinyan, et al.
Published: (2026)
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Reasoning Court: Combining Reasoning, Action, and Judgment for Multi-Hop Reasoning
by: Wu, Jingtian, et al.
Published: (2025)
by: Wu, Jingtian, et al.
Published: (2025)
SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
by: Huang, Weiyang, et al.
Published: (2026)
by: Huang, Weiyang, et al.
Published: (2026)
Multi-Hop Question Answering: When Can Humans Help, and Where do They Struggle?
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
by: Huang, Chengyu, et al.
Published: (2025)
by: Huang, Chengyu, et al.
Published: (2025)
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
by: Jin, Zhensheng, et al.
Published: (2025)
by: Jin, Zhensheng, et al.
Published: (2025)
Better LLM Reasoning via Dual-Play
by: Zhang, Zhengxin, et al.
Published: (2025)
by: Zhang, Zhengxin, et al.
Published: (2025)
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
by: Li, Ruosen, et al.
Published: (2025)
by: Li, Ruosen, et al.
Published: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
MASH: Modeling Abstention via Selective Help-Seeking
by: Gul, Mustafa Omer, et al.
Published: (2025)
by: Gul, Mustafa Omer, et al.
Published: (2025)
Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
Think Twice: Branch-and-Rethink Reasoning Reward Model
by: Jiao, Yizhu, et al.
Published: (2025)
by: Jiao, Yizhu, et al.
Published: (2025)
Corpus Poisoning via Approximate Greedy Gradient Descent
by: Su, Jinyan, et al.
Published: (2024)
by: Su, Jinyan, et al.
Published: (2024)
Token-weighted Direct Preference Optimization with Attention
by: Huang, Chengyu, et al.
Published: (2026)
by: Huang, Chengyu, et al.
Published: (2026)
Efficient Reasoning with Balanced Thinking
by: Li, Yulin, et al.
Published: (2026)
by: Li, Yulin, et al.
Published: (2026)
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
by: Wang, Binghai, et al.
Published: (2026)
by: Wang, Binghai, et al.
Published: (2026)
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
by: Rui, Shaohao, et al.
Published: (2025)
by: Rui, Shaohao, et al.
Published: (2025)
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
by: Tan, Wenhui, et al.
Published: (2025)
by: Tan, Wenhui, et al.
Published: (2025)
Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA
by: Su, Jinyan, et al.
Published: (2026)
by: Su, Jinyan, et al.
Published: (2026)
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
by: Huang, Chengyu, et al.
Published: (2026)
by: Huang, Chengyu, et al.
Published: (2026)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
by: Qi, Jirui, et al.
Published: (2025)
by: Qi, Jirui, et al.
Published: (2025)
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
by: Wu, Jiayun, et al.
Published: (2026)
by: Wu, Jiayun, et al.
Published: (2026)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
by: Zeng, Qingcheng, et al.
Published: (2025)
by: Zeng, Qingcheng, et al.
Published: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
WildChat: 1M ChatGPT Interaction Logs in the Wild
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Are Triggers Needed for Document-Level Event Extraction?
by: Shaar, Shaden, et al.
Published: (2024)
by: Shaar, Shaden, et al.
Published: (2024)
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
by: Zhang, Junyu, et al.
Published: (2025)
by: Zhang, Junyu, et al.
Published: (2025)
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
by: Yuan, Danlong, et al.
Published: (2025)
by: Yuan, Danlong, et al.
Published: (2025)
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
by: Ling, Zehui, et al.
Published: (2025)
by: Ling, Zehui, et al.
Published: (2025)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
by: Singh, Joykirat, et al.
Published: (2025)
by: Singh, Joykirat, et al.
Published: (2025)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
by: He, Yanjie
Published: (2026)
by: He, Yanjie
Published: (2026)
Adaptive Deep Reasoning: Triggering Deep Thinking When Needed
by: Wang, Yunhao, et al.
Published: (2025)
by: Wang, Yunhao, et al.
Published: (2025)
Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains
by: Xie, Hui, et al.
Published: (2026)
by: Xie, Hui, et al.
Published: (2026)
ThinkSwitcher: When to Think Hard, When to Think Fast
by: Liang, Guosheng, et al.
Published: (2025)
by: Liang, Guosheng, et al.
Published: (2025)
Similar Items
-
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025) -
Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions
by: Su, Jinyan, et al.
Published: (2026) -
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023) -
CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization
by: Su, Jinyan, et al.
Published: (2026) -
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control
by: Su, Jinyan, et al.
Published: (2025)