When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Keyu, Lyu, Tian, Su, Guinan, Geiping, Jonas, Yin, Lu, Canini, Marco, Liu, Shiwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
by: Su, Guinan, et al.
Published: (2025)
by: Su, Guinan, et al.
Published: (2025)
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
by: Li, Xueyan, et al.
Published: (2025)
by: Li, Xueyan, et al.
Published: (2025)
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
by: Su, Guinan, et al.
Published: (2026)
by: Su, Guinan, et al.
Published: (2026)
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
Efficient Test-Time Inference via Deterministic Exploration of Truncated Decoding Trees
by: Li, Xueyan, et al.
Published: (2026)
by: Li, Xueyan, et al.
Published: (2026)
Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
by: Song, Xinyuan, et al.
Published: (2025)
by: Song, Xinyuan, et al.
Published: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
by: Xin, Jihao, et al.
Published: (2026)
by: Xin, Jihao, et al.
Published: (2026)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation
by: Yang, Yanwu, et al.
Published: (2025)
by: Yang, Yanwu, et al.
Published: (2025)
Why Smaller Is Slower? Dimensional Misalignment in Compressed LLMs
by: Xin, Jihao, et al.
Published: (2026)
by: Xin, Jihao, et al.
Published: (2026)
Fewer is More: Boosting LLM Reasoning with Reinforced Context Pruning
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
by: Li, Pengxiang, et al.
Published: (2024)
by: Li, Pengxiang, et al.
Published: (2024)
Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs
by: Xu, Longhuan, et al.
Published: (2026)
by: Xu, Longhuan, et al.
Published: (2026)
When, Where and Why to Average Weights?
by: Ajroldi, Niccolò, et al.
Published: (2025)
by: Ajroldi, Niccolò, et al.
Published: (2025)
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
by: Egashira, Kazuki, et al.
Published: (2025)
by: Egashira, Kazuki, et al.
Published: (2025)
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
by: Lu, Haiquan, et al.
Published: (2024)
by: Lu, Haiquan, et al.
Published: (2024)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
by: He, Di, et al.
Published: (2026)
by: He, Di, et al.
Published: (2026)
Reassessing Layer Pruning in LLMs: New Insights and Methods
by: Lu, Yao, et al.
Published: (2024)
by: Lu, Yao, et al.
Published: (2024)
When the Chain Breaks: Interactive Diagnosis of LLM Chain-of-Thought Reasoning Errors
by: Chen, Shiwei, et al.
Published: (2026)
by: Chen, Shiwei, et al.
Published: (2026)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
More Data, Fewer Diacritics: Scaling Arabic TTS
by: Musleh, Ahmed, et al.
Published: (2026)
by: Musleh, Ahmed, et al.
Published: (2026)
Straightforward Layer-wise Pruning for More Efficient Visual Adaptation
by: Han, Ruizi, et al.
Published: (2024)
by: Han, Ruizi, et al.
Published: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
by: Chen, Xinrui, et al.
Published: (2025)
by: Chen, Xinrui, et al.
Published: (2025)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
by: Kharrat, Salma, et al.
Published: (2024)
by: Kharrat, Salma, et al.
Published: (2024)
Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer Proxies
by: Zhu, Yuehua, et al.
Published: (2020)
by: Zhu, Yuehua, et al.
Published: (2020)
High-Layer Attention Pruning with Rescaling
by: Liu, Songtao, et al.
Published: (2025)
by: Liu, Songtao, et al.
Published: (2025)
A Layer Selection Approach to Test Time Adaptation
by: Sahoo, Sabyasachi, et al.
Published: (2024)
by: Sahoo, Sabyasachi, et al.
Published: (2024)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
by: Zhang, Zhenyu, et al.
Published: (2024)
by: Zhang, Zhenyu, et al.
Published: (2024)
NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist
by: Bertram, Johannes, et al.
Published: (2026)
by: Bertram, Johannes, et al.
Published: (2026)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
by: Zhou, Shu, et al.
Published: (2026)
by: Zhou, Shu, et al.
Published: (2026)
Layer Collapse in Diffusion Language Models
by: Conzelmann, Alexander, et al.
Published: (2026)
by: Conzelmann, Alexander, et al.
Published: (2026)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Adam-mini: Use Fewer Learning Rates To Gain More
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
Similar Items
-
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
by: Su, Guinan, et al.
Published: (2025) -
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
by: Su, Guinan, et al.
Published: (2025) -
Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
by: Su, Guinan, et al.
Published: (2025) -
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
by: Li, Xueyan, et al.
Published: (2025) -
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
by: Su, Guinan, et al.
Published: (2026)