Does Your Reasoning Model Implicitly Know When to Stop Thinking?
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Zixuan, Xia, Xin, Ren, Yuxi, Zheng, Jianbin, Wang, Xuanda, Zhang, Zhixia, Xie, Hongyan, Liang, Songshi, Chen, Zehao, Xiao, Xuefeng, Zhuang, Fuzhen, Li, Jianxin, Wang, Deqing, Ban, Yikun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Real-Time Aligned Reward Model beyond Semantics
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
UniFAR: A Unified Facet-Aware Retrieval Framework for Scientific Documents
di: Dou, Zheng, et al.
Pubblicazione: (2026)
di: Dou, Zheng, et al.
Pubblicazione: (2026)
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
di: Chen, Zehao, et al.
Pubblicazione: (2026)
di: Chen, Zehao, et al.
Pubblicazione: (2026)
Heterogeneous Agent Collaborative Reinforcement Learning
di: Zhang, Zhixia, et al.
Pubblicazione: (2026)
di: Zhang, Zhixia, et al.
Pubblicazione: (2026)
Your Group-Relative Advantage Is Biased
di: Yang, Fengkai, et al.
Pubblicazione: (2026)
di: Yang, Fengkai, et al.
Pubblicazione: (2026)
UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment
di: Xie, Hongyan, et al.
Pubblicazione: (2026)
di: Xie, Hongyan, et al.
Pubblicazione: (2026)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
Policy Improvement Reinforcement Learning
di: Wang, Huaiyang, et al.
Pubblicazione: (2026)
di: Wang, Huaiyang, et al.
Pubblicazione: (2026)
LLMBoost: Make Large Language Models Stronger with Boosting
di: Chen, Zehao, et al.
Pubblicazione: (2025)
di: Chen, Zehao, et al.
Pubblicazione: (2025)
Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
di: Xie, Hongyan, et al.
Pubblicazione: (2025)
di: Xie, Hongyan, et al.
Pubblicazione: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
di: Lu, Xiaodong, et al.
Pubblicazione: (2026)
di: Lu, Xiaodong, et al.
Pubblicazione: (2026)
Bridging Social Psychology and LLM Reasoning: Conflict-Aware Meta-Review Generation via Cognitive Alignment
di: Chen, Wei, et al.
Pubblicazione: (2025)
di: Chen, Wei, et al.
Pubblicazione: (2025)
Federated Reasoning Distillation Framework with Model Learnability-Aware Data Allocation
di: Guo, Wei, et al.
Pubblicazione: (2026)
di: Guo, Wei, et al.
Pubblicazione: (2026)
One for Dozens: Adaptive REcommendation for All Domains with Counterfactual Augmentation
di: Luo, Huishi, et al.
Pubblicazione: (2024)
di: Luo, Huishi, et al.
Pubblicazione: (2024)
CDC: Causal Domain Clustering for Multi-Domain Recommendation
di: Luo, Huishi, et al.
Pubblicazione: (2025)
di: Luo, Huishi, et al.
Pubblicazione: (2025)
FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
di: Dou, Zheng, et al.
Pubblicazione: (2025)
di: Dou, Zheng, et al.
Pubblicazione: (2025)
When does word order matter and when doesn't it?
di: Chen, Xuanda, et al.
Pubblicazione: (2024)
di: Chen, Xuanda, et al.
Pubblicazione: (2024)
LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
di: Chen, Yiwen, et al.
Pubblicazione: (2026)
di: Chen, Yiwen, et al.
Pubblicazione: (2026)
Counterfactual Credit Policy Optimization for Multi-Agent Collaboration
di: Li, Zhongyi, et al.
Pubblicazione: (2026)
di: Li, Zhongyi, et al.
Pubblicazione: (2026)
Base Models Know How to Reason, Thinking Models Learn When
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
di: Sun, Renliang, et al.
Pubblicazione: (2025)
di: Sun, Renliang, et al.
Pubblicazione: (2025)
Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals
di: Xie, Roy, et al.
Pubblicazione: (2025)
di: Xie, Roy, et al.
Pubblicazione: (2025)
Implicit Reasoning in Transformers is Reasoning through Shortcuts
di: Lin, Tianhe, et al.
Pubblicazione: (2025)
di: Lin, Tianhe, et al.
Pubblicazione: (2025)
Hyperbolic Diffusion Recommender Model
di: Yuan, Meng, et al.
Pubblicazione: (2025)
di: Yuan, Meng, et al.
Pubblicazione: (2025)
TextBridgeGNN: Pre-training Graph Neural Network for Cross-Domain Recommendation via Text-Guided Transfer
di: Chen, Yiwen, et al.
Pubblicazione: (2025)
di: Chen, Yiwen, et al.
Pubblicazione: (2025)
Know When To Stop: A Study of Semantic Drift in Text Generation
di: Spataru, Ava, et al.
Pubblicazione: (2024)
di: Spataru, Ava, et al.
Pubblicazione: (2024)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
di: Zeng, Qingcheng, et al.
Pubblicazione: (2025)
di: Zeng, Qingcheng, et al.
Pubblicazione: (2025)
A Survey on Causal Inference for Recommendation
di: Luo, Huishi, et al.
Pubblicazione: (2023)
di: Luo, Huishi, et al.
Pubblicazione: (2023)
Knowing When to Stop Matters: A Unified Algorithm for Online Conversion under Horizon Uncertainty
di: Wang, Yanzhao, et al.
Pubblicazione: (2025)
di: Wang, Yanzhao, et al.
Pubblicazione: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
Thinking Out of Order: When Output Order Stops Reflecting Reasoning Order in Diffusion Language Models
di: Yu, Longxuan, et al.
Pubblicazione: (2026)
di: Yu, Longxuan, et al.
Pubblicazione: (2026)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
FairDgcl: Fairness-aware Recommendation with Dynamic Graph Contrastive Learning
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
di: Han, Qijun, et al.
Pubblicazione: (2026)
di: Han, Qijun, et al.
Pubblicazione: (2026)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
When Your Model Stops Working: Anytime-Valid Calibration Monitoring
di: Farran, Tristan
Pubblicazione: (2026)
di: Farran, Tristan
Pubblicazione: (2026)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
di: Qi, Jirui, et al.
Pubblicazione: (2025)
di: Qi, Jirui, et al.
Pubblicazione: (2025)
When Does Your Brain Know You? Segment Length and Its Impact on EEG-based Biometric Authentication Accuracy
di: Alzahab, Nibras Abo, et al.
Pubblicazione: (2024)
di: Alzahab, Nibras Abo, et al.
Pubblicazione: (2024)
Boosting Accuracy & Efficiency: Teaching LLMs to "Know When to Stop" with Absolute Quality Assessment
di: Etsuro, Minami, et al.
Pubblicazione: (2025)
di: Etsuro, Minami, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Real-Time Aligned Reward Model beyond Semantics
di: Huang, Zixuan, et al.
Pubblicazione: (2026) -
UniFAR: A Unified Facet-Aware Retrieval Framework for Scientific Documents
di: Dou, Zheng, et al.
Pubblicazione: (2026) -
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
di: Chen, Zehao, et al.
Pubblicazione: (2026) -
Heterogeneous Agent Collaborative Reinforcement Learning
di: Zhang, Zhixia, et al.
Pubblicazione: (2026) -
Your Group-Relative Advantage Is Biased
di: Yang, Fengkai, et al.
Pubblicazione: (2026)