DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
Fuente:
arXiv
Saved in:
| Main Authors: | Woo, Sunghyeon, Park, Baeseong, Kim, Byeongwook, Jo, Minjung, Kwon, Se Jung, Jeon, Dongsuk, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
by: Lee, Jeongtae, et al.
Published: (2026)
by: Lee, Jeongtae, et al.
Published: (2026)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022)
by: Park, Gunho, et al.
Published: (2022)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
PaCA: Partial Connection Adaptation for Efficient Fine-Tuning
by: Woo, Sunghyeon, et al.
Published: (2025)
by: Woo, Sunghyeon, et al.
Published: (2025)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
by: Bae, Jeongin, et al.
Published: (2026)
by: Bae, Jeongin, et al.
Published: (2026)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
by: Yoon, Kanghoon, et al.
Published: (2025)
by: Yoon, Kanghoon, et al.
Published: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Salt Water Drops Slide Faster: Ionic Modulation of Drop Friction
by: Dongho Shin, et al.
Published: (2026)
by: Dongho Shin, et al.
Published: (2026)
I Dropped a Neural Net
by: Park, Hyunwoo
Published: (2026)
by: Park, Hyunwoo
Published: (2026)
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025)
by: Lee, Joonhyung, et al.
Published: (2025)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026)
by: Rajabzadeh, Hossein, et al.
Published: (2026)
On the Threshold of Drop Fragmentation under Impulsive Acceleration
by: Parik, Aditya, et al.
Published: (2022)
by: Parik, Aditya, et al.
Published: (2022)
DropLoRA: Sparse Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
by: Zhang, Haojie
Published: (2025)
by: Zhang, Haojie
Published: (2025)
Resonance and Damping in Drop-Cantilever Interactions
by: Fowler, Crystal, et al.
Published: (2024)
by: Fowler, Crystal, et al.
Published: (2024)
DropGaussian: Structural Regularization for Sparse-view Gaussian Splatting
by: Park, Hyunwoo, et al.
Published: (2025)
by: Park, Hyunwoo, et al.
Published: (2025)
OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance
by: Park, Yeo Jeong, et al.
Published: (2026)
by: Park, Yeo Jeong, et al.
Published: (2026)
WACA-UNet: Weakness-Aware Channel Attention for Static IR Drop Prediction in Integrated Circuit Design
by: Seo, Youngmin, et al.
Published: (2025)
by: Seo, Youngmin, et al.
Published: (2025)
39‐1: Invited Paper: A Multi‐Drop High‐Speed Link with Foveated Up‐Scaler to Reduce Wires and Data Bandwidth in LED‐on‐Silicon‐Backplane for AR Glasses
by: Hyun-Wook Lim, et al.
Published: (2025)
by: Hyun-Wook Lim, et al.
Published: (2025)
CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection based View Transformation
by: Lee, In-Jae, et al.
Published: (2025)
by: Lee, In-Jae, et al.
Published: (2025)
Bursting Drops
by: Kulkarni, Varun, et al.
Published: (2022)
by: Kulkarni, Varun, et al.
Published: (2022)
Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model
by: Park, Sangjune, et al.
Published: (2026)
by: Park, Sangjune, et al.
Published: (2026)
Learn&Drop: Fast Learning of CNNs based on Layer Dropping
by: Cruciata, Giorgio, et al.
Published: (2026)
by: Cruciata, Giorgio, et al.
Published: (2026)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
by: Kim, Han-Byul, et al.
Published: (2025)
by: Kim, Han-Byul, et al.
Published: (2025)
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
by: Scribano, Carmelo, et al.
Published: (2026)
by: Scribano, Carmelo, et al.
Published: (2026)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
by: Xing, Long, et al.
Published: (2024)
by: Xing, Long, et al.
Published: (2024)
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
by: Wang, Guoxin, et al.
Published: (2025)
by: Wang, Guoxin, et al.
Published: (2025)
ADEdgeDrop: Adversarial Edge Dropping for Robust Graph Neural Networks
by: Chen, Zhaoliang, et al.
Published: (2024)
by: Chen, Zhaoliang, et al.
Published: (2024)
Drop to Gate Nasal Drops Attenuates Sepsis‐Induced Cognitive Dysfunction
by: Yaping Zhuang, et al.
Published: (2024)
by: Yaping Zhuang, et al.
Published: (2024)
Spreading of Dynamically Crosslinked Polydimethylsiloxane Drops
by: Kyujin Ko, et al.
Published: (2024)
by: Kyujin Ko, et al.
Published: (2024)
Similar Items
-
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026) -
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
by: Lee, Jeongtae, et al.
Published: (2026) -
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022) -
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
by: Woo, Sunghyeon, et al.
Published: (2026) -
PaCA: Partial Connection Adaptation for Efficient Fine-Tuning
by: Woo, Sunghyeon, et al.
Published: (2025)