SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Jiwon, Oh, Kyungseok, Kim, Taesu, Kim, Hyungjun, Kim, Yulhwa, Kim, Jae-Joon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025)
by: Jo, Dongwon, et al.
Published: (2025)
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
by: Jo, Dongwon, et al.
Published: (2024)
by: Jo, Dongwon, et al.
Published: (2024)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024)
by: Jeon, Hyesung, et al.
Published: (2024)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)
by: Song, Jiwon, et al.
Published: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
RelayGen: Intra-Generation Model Switching for Efficient Reasoning
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026)
by: Kwon, Omin, et al.
Published: (2026)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
by: Song, Jiwon, et al.
Published: (2026)
by: Song, Jiwon, et al.
Published: (2026)
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
by: Lee, Changhun, et al.
Published: (2023)
by: Lee, Changhun, et al.
Published: (2023)
MedRep: Medical Concept Representation for General Electronic Health Record Foundation Models
by: Kim, Junmo, et al.
Published: (2025)
by: Kim, Junmo, et al.
Published: (2025)
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2025)
by: Jeon, Hyesung, et al.
Published: (2025)
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
by: Jung, Yeonjoon, et al.
Published: (2025)
by: Jung, Yeonjoon, et al.
Published: (2025)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
by: Kim, Jongsuk, et al.
Published: (2024)
by: Kim, Jongsuk, et al.
Published: (2024)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
by: Oh, Hyungjun, et al.
Published: (2024)
by: Oh, Hyungjun, et al.
Published: (2024)
ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
by: Lee, Hyunseok, et al.
Published: (2025)
by: Lee, Hyunseok, et al.
Published: (2025)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025)
by: Kang, Beomseok, et al.
Published: (2025)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
by: Yang, Jaewoo, et al.
Published: (2024)
by: Yang, Jaewoo, et al.
Published: (2024)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
by: Lim, Junghwan, et al.
Published: (2025)
by: Lim, Junghwan, et al.
Published: (2025)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
by: Lee, Yuna, et al.
Published: (2026)
by: Lee, Yuna, et al.
Published: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
by: Noh, Kanghyun, et al.
Published: (2026)
by: Noh, Kanghyun, et al.
Published: (2026)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
by: Ho, Namgyu, et al.
Published: (2024)
by: Ho, Namgyu, et al.
Published: (2024)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
by: Kim, Seonwu, et al.
Published: (2025)
by: Kim, Seonwu, et al.
Published: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
by: Kim, Doyoung, et al.
Published: (2025)
by: Kim, Doyoung, et al.
Published: (2025)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
by: Kim, Seongho, et al.
Published: (2024)
by: Kim, Seongho, et al.
Published: (2024)
Exploring Changes in Nation Perception with Nationality-Assigned Personas in LLMs
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
Enhancing Clinical Efficiency through LLM: Discharge Note Generation for Cardiac Patients
by: Jung, HyoJe, et al.
Published: (2024)
by: Jung, HyoJe, et al.
Published: (2024)
Compressed Context Memory For Online Language Model Interaction
by: Kim, Jang-Hyun, et al.
Published: (2023)
by: Kim, Jang-Hyun, et al.
Published: (2023)
Hyperloop Transformers
by: Zeitoun, Abbas, et al.
Published: (2026)
by: Zeitoun, Abbas, et al.
Published: (2026)
Few-shot Personalization of LLMs with Mis-aligned Responses
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Surrogate modeling for interpreting black-box LLMs in medical predictions
by: Han, Changho, et al.
Published: (2026)
by: Han, Changho, et al.
Published: (2026)
Task Schema and Binding: A Double Dissociation Study of In-Context Learning
by: Kim, Chaeha
Published: (2025)
by: Kim, Chaeha
Published: (2025)
A Baseline for Self-state Identification and Classification in Mental Health Data: CLPsych 2025 Task
by: Kim, Laerdon
Published: (2025)
by: Kim, Laerdon
Published: (2025)
Similar Items
-
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
by: Jo, Dongwon, et al.
Published: (2025) -
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
by: Jo, Dongwon, et al.
Published: (2024) -
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024) -
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025) -
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)