Locality-Aware Redundancy Pruning for LLM Depth Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Yun, Vincent-Daniel, Kim, Youngrae, Lim, Woosang, Heo, YoungJin, Kim, Minkyu, Lee, Sunwoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
by: Yun, Vincent-Daniel, et al.
Published: (2025)
by: Yun, Vincent-Daniel, et al.
Published: (2025)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
Test-time Alignment of Diffusion Models without Reward Over-optimization
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
Deep Learning and Matrix Completion-aided IoT Network Localization in the Outlier Scenarios
by: Kim, Sunwoo
Published: (2025)
by: Kim, Sunwoo
Published: (2025)
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
by: Lim, Woosang, et al.
Published: (2025)
by: Lim, Woosang, et al.
Published: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Explicit Feature Interaction-aware Graph Neural Networks
by: Kim, Minkyu, et al.
Published: (2022)
by: Kim, Minkyu, et al.
Published: (2022)
Interaction-Aware Influence Functions for Group Attribution
by: Heo, Jaeseung, et al.
Published: (2026)
by: Heo, Jaeseung, et al.
Published: (2026)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
FlowerFormer: Empowering Neural Architecture Encoding using a Flow-aware Graph Transformer
by: Hwang, Dongyeong, et al.
Published: (2024)
by: Hwang, Dongyeong, et al.
Published: (2024)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Measuring the Depth of LLM Unlearning via Activation Patching
by: Lee, Jaeung, et al.
Published: (2026)
by: Lee, Jaeung, et al.
Published: (2026)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
CASK: Core-Aware Selective KV Compression for Reasoning Traces
by: Kim, Buseong, et al.
Published: (2026)
by: Kim, Buseong, et al.
Published: (2026)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
by: Wagner, Moritz, et al.
Published: (2025)
by: Wagner, Moritz, et al.
Published: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
Feature-Centric Unsupervised Node Representation Learning Without Homophily Assumption
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
Multi-View Node Pruning for Accurate Graph Representation
by: Kim, Hanjin, et al.
Published: (2025)
by: Kim, Hanjin, et al.
Published: (2025)
Efficient and Scalable Estimation of Tool Representations in Vector Space
by: Moon, Suhong, et al.
Published: (2024)
by: Moon, Suhong, et al.
Published: (2024)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Subgraph Federated Learning for Local Generalization
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
EPIC: Graph Augmentation with Edit Path Interpolation via Learnable Cost
by: Heo, Jaeseung, et al.
Published: (2023)
by: Heo, Jaeseung, et al.
Published: (2023)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning
by: Song, Sanghyeob, et al.
Published: (2026)
by: Song, Sanghyeob, et al.
Published: (2026)
GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models
by: Liu, Xiaoyun, et al.
Published: (2026)
by: Liu, Xiaoyun, et al.
Published: (2026)
RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction
by: Ko, Hanbum, et al.
Published: (2026)
by: Ko, Hanbum, et al.
Published: (2026)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
by: Choi, Minsik, et al.
Published: (2025)
by: Choi, Minsik, et al.
Published: (2025)
Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
by: Kim, Sung-Hyun, et al.
Published: (2025)
by: Kim, Sung-Hyun, et al.
Published: (2025)
PaCA: Partial Connection Adaptation for Efficient Fine-Tuning
by: Woo, Sunghyeon, et al.
Published: (2025)
by: Woo, Sunghyeon, et al.
Published: (2025)
Severing Spurious Correlations with Data Pruning
by: Mulchandani, Varun, et al.
Published: (2025)
by: Mulchandani, Varun, et al.
Published: (2025)
Learning-augmented robotic automation for real-world manufacturing
by: Kim, Yunho, et al.
Published: (2026)
by: Kim, Yunho, et al.
Published: (2026)
SAMOSA: Sharpness Aware Minimization for Open Set Active learning
by: Kim, Young In, et al.
Published: (2025)
by: Kim, Young In, et al.
Published: (2025)
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
by: Kim, Hee-Sung, et al.
Published: (2026)
by: Kim, Hee-Sung, et al.
Published: (2026)
Posterior Label Smoothing for Node Classification
by: Heo, Jaeseung, et al.
Published: (2024)
by: Heo, Jaeseung, et al.
Published: (2024)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Similar Items
-
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026) -
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
by: Yun, Vincent-Daniel, et al.
Published: (2025) -
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026) -
Test-time Alignment of Diffusion Models without Reward Over-optimization
by: Kim, Sunwoo, et al.
Published: (2025) -
Deep Learning and Matrix Completion-aided IoT Network Localization in the Outlier Scenarios
by: Kim, Sunwoo
Published: (2025)