Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Minkyu, Yun, Vincent-Daniel, Kim, Youngrae, Cho, Suin, Lim, Woosang, Lee, Sunwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Locality-Aware Redundancy Pruning for LLM Depth Compression
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025)
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2025)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
von: Kim, Taesu, et al.
Veröffentlicht: (2024)
von: Kim, Taesu, et al.
Veröffentlicht: (2024)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
Test-time Alignment of Diffusion Models without Reward Over-optimization
von: Kim, Sunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Sunwoo, et al.
Veröffentlicht: (2025)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
LLMs Position Themselves as More Rational Than Humans: Emergence of AI Self-Awareness Measured Through Game Theory
von: Kim, Kyung-Hoon
Veröffentlicht: (2025)
von: Kim, Kyung-Hoon
Veröffentlicht: (2025)
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
von: Men, Xin, et al.
Veröffentlicht: (2024)
von: Men, Xin, et al.
Veröffentlicht: (2024)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
Knowledge Synthesis of Photosynthesis Research Using a Large Language Model
von: Yoon, Seungri, et al.
Veröffentlicht: (2025)
von: Yoon, Seungri, et al.
Veröffentlicht: (2025)
Measuring the Depth of LLM Unlearning via Activation Patching
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
Fewer is More: Boosting LLM Reasoning with Reinforced Context Pruning
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
von: Huang, Xijie, et al.
Veröffentlicht: (2023)
The Algebra of Meaning: Why Machines Need Montague More Than Moore's Law
von: Jeong, Cheonkam, et al.
Veröffentlicht: (2025)
von: Jeong, Cheonkam, et al.
Veröffentlicht: (2025)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
Text Change Detection in Multilingual Documents Using Image Comparison
von: Park, Doyoung, et al.
Veröffentlicht: (2024)
von: Park, Doyoung, et al.
Veröffentlicht: (2024)
Think Clearly: Improving Reasoning via Redundant Token Pruning
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Prompt-based Depth Pruning of Large Language Models
von: Wee, Juyun, et al.
Veröffentlicht: (2025)
von: Wee, Juyun, et al.
Veröffentlicht: (2025)
DeepPrune: Parallel Scaling without Inter-trace Redundancy
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
von: Kurz, Simon, et al.
Veröffentlicht: (2024)
von: Kurz, Simon, et al.
Veröffentlicht: (2024)
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
von: Bhattacharyya, Chaitali, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Chaitali, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
LLM-Enhanced Black-Litterman Portfolio Optimization
von: Lee, Youngbin, et al.
Veröffentlicht: (2025)
von: Lee, Youngbin, et al.
Veröffentlicht: (2025)
Trivial Vocabulary Bans Improve LLM Reasoning More Than Deep Linguistic Constraints
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
von: Shin, Minkyu, et al.
Veröffentlicht: (2023)
von: Shin, Minkyu, et al.
Veröffentlicht: (2023)
RoToR: Towards More Reliable Responses for Order-Invariant Inputs
von: Yoon, Soyoung, et al.
Veröffentlicht: (2025)
von: Yoon, Soyoung, et al.
Veröffentlicht: (2025)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
von: Qin, Jiayu, et al.
Veröffentlicht: (2025)
von: Qin, Jiayu, et al.
Veröffentlicht: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
Development of Flexure‐Based Supernumerary Robotic Finger for Hand Function Augmentation
von: Junmo Yang, et al.
Veröffentlicht: (2024)
von: Junmo Yang, et al.
Veröffentlicht: (2024)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
von: Lee, Gisang, et al.
Veröffentlicht: (2024)
von: Lee, Gisang, et al.
Veröffentlicht: (2024)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
von: Moumen, Adel, et al.
Veröffentlicht: (2026)
von: Moumen, Adel, et al.
Veröffentlicht: (2026)
Raon-Speech Technical Report
von: Kim, Beomsoo, et al.
Veröffentlicht: (2026)
von: Kim, Beomsoo, et al.
Veröffentlicht: (2026)
Discrete Audio Tokens: More Than a Survey!
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
von: Lim, Woosang, et al.
Veröffentlicht: (2025)
von: Lim, Woosang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Locality-Aware Redundancy Pruning for LLM Depth Compression
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026) -
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025) -
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026) -
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025) -
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
von: Park, Dongmin, et al.
Veröffentlicht: (2024)