SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zheng, Dong, Qingxiu, Ma, Jingyuan, Zhang, Di, Jia, Kai, Sui, Zhifang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Boosting Large Language Models with Synthetic Preference Data
von: Dong, Qingxiu, et al.
Veröffentlicht: (2024)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2024)
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
von: Yang, Yixin, et al.
Veröffentlicht: (2026)
von: Yang, Yixin, et al.
Veröffentlicht: (2026)
Chain-of-Thought Tokens are Computer Program Variables
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)
Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge
von: Tang, Yao, et al.
Veröffentlicht: (2026)
von: Tang, Yao, et al.
Veröffentlicht: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
A Survey on In-context Learning
von: Dong, Qingxiu, et al.
Veröffentlicht: (2022)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2022)
Hierarchical Budget Policy Optimization for Adaptive Reasoning
von: Lyu, Shangke, et al.
Veröffentlicht: (2025)
von: Lyu, Shangke, et al.
Veröffentlicht: (2025)
LLM-REVal: Can We Trust LLM Reviewers Yet?
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
von: Mishra, Kshitij, et al.
Veröffentlicht: (2026)
von: Mishra, Kshitij, et al.
Veröffentlicht: (2026)
LLM-Oriented Token-Adaptive Knowledge Distillation
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
ARS: Adaptive Reasoning Suppression for Efficient Large Reasoning Language Models
von: Zheng, Dongqi
Veröffentlicht: (2025)
von: Zheng, Dongqi
Veröffentlicht: (2025)
Extending Token Computation for LLM Reasoning
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
Steering LLM Thinking with Budget Guidance
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows
von: Wang, Xinglin, et al.
Veröffentlicht: (2026)
von: Wang, Xinglin, et al.
Veröffentlicht: (2026)
Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning
von: Wang, Xinglin, et al.
Veröffentlicht: (2024)
von: Wang, Xinglin, et al.
Veröffentlicht: (2024)
A Scaling Law for Token Efficiency in LLM Fine-Tuning Under Fixed Compute Budgets
von: Lagasse, Ryan, et al.
Veröffentlicht: (2025)
von: Lagasse, Ryan, et al.
Veröffentlicht: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
von: Zhou, Zhi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhi, et al.
Veröffentlicht: (2025)
CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation
von: Petullo, James, et al.
Veröffentlicht: (2026)
von: Petullo, James, et al.
Veröffentlicht: (2026)
AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
von: Huang, Shijue, et al.
Veröffentlicht: (2025)
von: Huang, Shijue, et al.
Veröffentlicht: (2025)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
von: Sabry, Mohammed, et al.
Veröffentlicht: (2026)
von: Sabry, Mohammed, et al.
Veröffentlicht: (2026)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
von: Yang, Wang, et al.
Veröffentlicht: (2026)
von: Yang, Wang, et al.
Veröffentlicht: (2026)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025)
von: Su, DiJia, et al.
Veröffentlicht: (2025)
HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic Reasoning
von: Yang, Jinning, et al.
Veröffentlicht: (2025)
von: Yang, Jinning, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
MPO: Boosting LLM Agents with Meta Plan Optimization
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
von: Chen, Xuhang, et al.
Veröffentlicht: (2025)
von: Chen, Xuhang, et al.
Veröffentlicht: (2025)
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
von: Zhang, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyun, et al.
Veröffentlicht: (2025)
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
FLM-101B: An Open LLM and How to Train It with $100K Budget
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
von: Qian, Chen, et al.
Veröffentlicht: (2025)
von: Qian, Chen, et al.
Veröffentlicht: (2025)
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
von: Wang, Jingyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Boosting Large Language Models with Synthetic Preference Data
von: Dong, Qingxiu, et al.
Veröffentlicht: (2024) -
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024) -
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025) -
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
von: Yang, Yixin, et al.
Veröffentlicht: (2026) -
Chain-of-Thought Tokens are Computer Program Variables
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)