Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Wenya, Shaochen, Zhong, Le, Hoang Anh Duy, Xu, Zhaozhuo, Xie, Jianwen, Liu, Zirui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
Do LLMs Know to Respect Copyright Notice?
von: Xu, Jialiang, et al.
Veröffentlicht: (2024)
von: Xu, Jialiang, et al.
Veröffentlicht: (2024)
AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
von: Luo, Feng, et al.
Veröffentlicht: (2025)
von: Luo, Feng, et al.
Veröffentlicht: (2025)
Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
von: Liu, Zirui, et al.
Veröffentlicht: (2023)
von: Liu, Zirui, et al.
Veröffentlicht: (2023)
A Neural Model for Word Repetition
von: Dager, Daniel, et al.
Veröffentlicht: (2025)
von: Dager, Daniel, et al.
Veröffentlicht: (2025)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
von: Le, Hoang Anh Duy, et al.
Veröffentlicht: (2026)
von: Le, Hoang Anh Duy, et al.
Veröffentlicht: (2026)
On Repetitive Finite Automata with Translucent Words
von: Mráz, František, et al.
Veröffentlicht: (2025)
von: Mráz, František, et al.
Veröffentlicht: (2025)
Large Language Models Know What Makes Exemplary Contexts
von: Long, Quanyu, et al.
Veröffentlicht: (2024)
von: Long, Quanyu, et al.
Veröffentlicht: (2024)
Sensitivity Meets Sparsity: The Impact of Extremely Sparse Parameter Patterns on Theory-of-Mind of Large Language Models
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT
von: Nguyen, Duy Anh
Veröffentlicht: (2026)
von: Nguyen, Duy Anh
Veröffentlicht: (2026)
What Do Self-Supervised Speech Models Know About Words?
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
Schema Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding
von: Le, Yifan
Veröffentlicht: (2026)
von: Le, Yifan
Veröffentlicht: (2026)
XAI-enhanced Comparative Opinion Mining via Aspect-based Scoring and Semantic Reasoning
von: Le, Ngoc-Quang, et al.
Veröffentlicht: (2026)
von: Le, Ngoc-Quang, et al.
Veröffentlicht: (2026)
The Diminishing Returns of Early-Exit Decoding in Modern LLMs
von: Wei, Rui, et al.
Veröffentlicht: (2026)
von: Wei, Rui, et al.
Veröffentlicht: (2026)
Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning
von: Nguyen, Manh, et al.
Veröffentlicht: (2026)
von: Nguyen, Manh, et al.
Veröffentlicht: (2026)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
Deductive Beam Search: Decoding Deducible Rationale for Chain-of-Thought Reasoning
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
Self-consistent Reasoning For Solving Math Word Problems
von: Xiong, Jing, et al.
Veröffentlicht: (2022)
von: Xiong, Jing, et al.
Veröffentlicht: (2022)
Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers
von: Chen, Xin, et al.
Veröffentlicht: (2026)
von: Chen, Xin, et al.
Veröffentlicht: (2026)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
von: Long, Quanyu, et al.
Veröffentlicht: (2026)
von: Long, Quanyu, et al.
Veröffentlicht: (2026)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
von: Bai, Yuyang, et al.
Veröffentlicht: (2026)
Large Language Model-Enhanced Symbolic Reasoning for Knowledge Base Completion
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
von: He, Qiyuan, et al.
Veröffentlicht: (2025)
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2023)
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2023)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
von: Long, Jikai, et al.
Veröffentlicht: (2025)
von: Long, Jikai, et al.
Veröffentlicht: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers
von: Nguyen, Viet-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Viet-Anh, et al.
Veröffentlicht: (2025)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
Prompt Repetition Improves Non-Reasoning LLMs
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2025)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2025)
Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation
von: Li, Bo, et al.
Veröffentlicht: (2025)
von: Li, Bo, et al.
Veröffentlicht: (2025)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
von: Liu, Hui, et al.
Veröffentlicht: (2023)
von: Liu, Hui, et al.
Veröffentlicht: (2023)
OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice
von: Wang, Rongyang, et al.
Veröffentlicht: (2026)
von: Wang, Rongyang, et al.
Veröffentlicht: (2026)
Sequence Repetition Enhances Token Embeddings and Improves Sequence Labeling with Decoder-only Language Models
von: Kukić, Matija Luka, et al.
Veröffentlicht: (2026)
von: Kukić, Matija Luka, et al.
Veröffentlicht: (2026)
DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
von: Xu, Zicheng, et al.
Veröffentlicht: (2025)
von: Xu, Zicheng, et al.
Veröffentlicht: (2025)
Adaptive Two-Phase Finetuning LLMs for Japanese Legal Text Retrieval
von: Trung, Quang Hoang, et al.
Veröffentlicht: (2024)
von: Trung, Quang Hoang, et al.
Veröffentlicht: (2024)
Repetition Neurons: How Do Language Models Produce Repetitions?
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
von: Wu, Yuheng, et al.
Veröffentlicht: (2025) -
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024) -
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024) -
Do LLMs Know to Respect Copyright Notice?
von: Xu, Jialiang, et al.
Veröffentlicht: (2024) -
AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
von: Luo, Feng, et al.
Veröffentlicht: (2025)