ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junda, Yang, Zhichao, Zhang, Dongxu, Batra, Sanjit Singh, Tillman, Robert E. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast and Effective On-policy Distillation from Reasoning Prefixes
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
von: Kwok, Tung Sum Thomas, et al.
Veröffentlicht: (2026)
von: Kwok, Tung Sum Thomas, et al.
Veröffentlicht: (2026)
Imputation of Unknown Missingness in Sparse Electronic Health Records
von: Han, Jun, et al.
Veröffentlicht: (2026)
von: Han, Jun, et al.
Veröffentlicht: (2026)
POET: Protocol Optimization via Eligibility Tuning
von: Das, Trisha, et al.
Veröffentlicht: (2026)
von: Das, Trisha, et al.
Veröffentlicht: (2026)
MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
Statistical Early Stopping for Reasoning Models
von: Xie, Yangxinyu, et al.
Veröffentlicht: (2026)
von: Xie, Yangxinyu, et al.
Veröffentlicht: (2026)
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2025)
von: Kambhampati, Subbarao, et al.
Veröffentlicht: (2025)
TARSE: Test-Time Adaptation via Retrieval of Skills and Experience for Reasoning Agents
von: Wang, Junda, et al.
Veröffentlicht: (2026)
von: Wang, Junda, et al.
Veröffentlicht: (2026)
Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning
von: Zhou, Zhaomeng, et al.
Veröffentlicht: (2026)
von: Zhou, Zhaomeng, et al.
Veröffentlicht: (2026)
PersonalQ: Select, Quantize, and Serve Personalized Diffusion Models for Efficient Inference
von: Wang, Qirui, et al.
Veröffentlicht: (2026)
von: Wang, Qirui, et al.
Veröffentlicht: (2026)
Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
von: Cao, Hongye, et al.
Veröffentlicht: (2025)
von: Cao, Hongye, et al.
Veröffentlicht: (2025)
Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
von: Lin, Fulin, et al.
Veröffentlicht: (2025)
von: Lin, Fulin, et al.
Veröffentlicht: (2025)
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
von: Huang, Jingkai, et al.
Veröffentlicht: (2026)
von: Huang, Jingkai, et al.
Veröffentlicht: (2026)
Early Stopping for Large Reasoning Models via Confidence Dynamics
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
ChatThero: An LLM-Supported Chatbot for Behavior Change and Therapeutic Support in Addiction Recovery
von: Wang, Junda, et al.
Veröffentlicht: (2025)
von: Wang, Junda, et al.
Veröffentlicht: (2025)
A*-Decoding: Token-Efficient Inference Scaling
von: Chatziveroglou, Giannis
Veröffentlicht: (2025)
von: Chatziveroglou, Giannis
Veröffentlicht: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
von: Chen, Benteng, et al.
Veröffentlicht: (2026)
von: Chen, Benteng, et al.
Veröffentlicht: (2026)
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference
von: Xie, Wenxuan, et al.
Veröffentlicht: (2026)
von: Xie, Wenxuan, et al.
Veröffentlicht: (2026)
ESPO: Early-Stopping Proximal Policy Optimization
von: Li, Zihang, et al.
Veröffentlicht: (2026)
von: Li, Zihang, et al.
Veröffentlicht: (2026)
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
von: Zeng, Runjia, et al.
Veröffentlicht: (2026)
von: Zeng, Runjia, et al.
Veröffentlicht: (2026)
GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts
von: Zeng, Wenhao, et al.
Veröffentlicht: (2026)
von: Zeng, Wenhao, et al.
Veröffentlicht: (2026)
Rethinking Early Stopping: Refine, Then Calibrate
von: Berta, Eugène, et al.
Veröffentlicht: (2025)
von: Berta, Eugène, et al.
Veröffentlicht: (2025)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
von: Huang, Zixuan, et al.
Veröffentlicht: (2026)
von: Huang, Zixuan, et al.
Veröffentlicht: (2026)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
von: Nagle, Alliot, et al.
Veröffentlicht: (2026)
Token-Efficient RL for LLM Reasoning
von: Lee, Alan, et al.
Veröffentlicht: (2025)
von: Lee, Alan, et al.
Veröffentlicht: (2025)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
S2O: Early Stopping for Sparse Attention via Online Permutation
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
von: Yu, Zewei, et al.
Veröffentlicht: (2026)
von: Yu, Zewei, et al.
Veröffentlicht: (2026)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
von: Yu, Zishun, et al.
Veröffentlicht: (2025)
von: Yu, Zishun, et al.
Veröffentlicht: (2025)
ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning
von: Yi, Jingyang, et al.
Veröffentlicht: (2025)
von: Yi, Jingyang, et al.
Veröffentlicht: (2025)
Towards Competitive Search Relevance For Inference-Free Learned Sparse Retrievers
von: Geng, Zhichao, et al.
Veröffentlicht: (2024)
von: Geng, Zhichao, et al.
Veröffentlicht: (2024)
Exploring $\ell_0$ Sparsification for Inference-free Sparse Retrievers
von: Shen, Xinjie, et al.
Veröffentlicht: (2025)
von: Shen, Xinjie, et al.
Veröffentlicht: (2025)
Learning to Stop Cut Generation for Efficient Mixed-Integer Linear Programming
von: Ling, Haotian, et al.
Veröffentlicht: (2024)
von: Ling, Haotian, et al.
Veröffentlicht: (2024)
EMS: Multi-Agent Voting via Efficient Majority-then-Stopping
von: Liu, Yiqing, et al.
Veröffentlicht: (2026)
von: Liu, Yiqing, et al.
Veröffentlicht: (2026)
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast and Effective On-policy Distillation from Reasoning Prefixes
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026) -
Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
von: Yang, Zhichao, et al.
Veröffentlicht: (2026) -
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
von: Kwok, Tung Sum Thomas, et al.
Veröffentlicht: (2026) -
Imputation of Unknown Missingness in Sparse Electronic Health Records
von: Han, Jun, et al.
Veröffentlicht: (2026) -
POET: Protocol Optimization via Eligibility Tuning
von: Das, Trisha, et al.
Veröffentlicht: (2026)