Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Hayun, Shin, Dongkun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
von: Yu, Seungmin, et al.
Veröffentlicht: (2024)
von: Yu, Seungmin, et al.
Veröffentlicht: (2024)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025)
Cognitive Chunking for Soft Prompts: Accelerating Compressor Learning via Block-wise Causal Masking
von: Liu, Guojie, et al.
Veröffentlicht: (2026)
von: Liu, Guojie, et al.
Veröffentlicht: (2026)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
von: Zhang, Kunlong, et al.
Veröffentlicht: (2025)
von: Zhang, Kunlong, et al.
Veröffentlicht: (2025)
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
von: Shing, Makoto, et al.
Veröffentlicht: (2025)
von: Shing, Makoto, et al.
Veröffentlicht: (2025)
An Efficient Training Algorithm for Models with Block-wise Sparsity
von: Zhu, Ding, et al.
Veröffentlicht: (2025)
von: Zhu, Ding, et al.
Veröffentlicht: (2025)
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
von: Russinovich, Mark, et al.
Veröffentlicht: (2026)
von: Russinovich, Mark, et al.
Veröffentlicht: (2026)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
von: Qin, Jiayu, et al.
Veröffentlicht: (2025)
von: Qin, Jiayu, et al.
Veröffentlicht: (2025)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
von: Razin, Noam, et al.
Veröffentlicht: (2024)
von: Razin, Noam, et al.
Veröffentlicht: (2024)
Understanding the Effects of Safety Unalignment on Large Language Models
von: Halloran, John T.
Veröffentlicht: (2026)
von: Halloran, John T.
Veröffentlicht: (2026)
Content-Style Learning from Unaligned Domains: Identifiability under Unknown Latent Dimensions
von: Shrestha, Sagar, et al.
Veröffentlicht: (2024)
von: Shrestha, Sagar, et al.
Veröffentlicht: (2024)
GED-Consistent Disentanglement of Aligned and Unaligned Substructures for Graph Similarity Learning
von: Zhan, Zhentao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhentao, et al.
Veröffentlicht: (2025)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
von: Demirkiran, Cansu, et al.
Veröffentlicht: (2023)
von: Demirkiran, Cansu, et al.
Veröffentlicht: (2023)
Proto-EVFL: Enhanced Vertical Federated Learning via Dual Prototype with Extremely Unaligned Data
von: Guo, Wei, et al.
Veröffentlicht: (2025)
von: Guo, Wei, et al.
Veröffentlicht: (2025)
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
von: Jung, Jaeheun, et al.
Veröffentlicht: (2025)
von: Jung, Jaeheun, et al.
Veröffentlicht: (2025)
DNNShifter: An Efficient DNN Pruning System for Edge Computing
von: Eccles, Bailey J., et al.
Veröffentlicht: (2023)
von: Eccles, Bailey J., et al.
Veröffentlicht: (2023)
Differentially Private Block-wise Gradient Shuffle for Deep Learning
von: Zagardo, David
Veröffentlicht: (2024)
von: Zagardo, David
Veröffentlicht: (2024)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
von: Ilin, Ivan, et al.
Veröffentlicht: (2025)
von: Ilin, Ivan, et al.
Veröffentlicht: (2025)
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
von: Xu, Zihuai, et al.
Veröffentlicht: (2025)
von: Xu, Zihuai, et al.
Veröffentlicht: (2025)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
von: Li, Meng, et al.
Veröffentlicht: (2026)
von: Li, Meng, et al.
Veröffentlicht: (2026)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
von: Jo, Hyun-rae, et al.
Veröffentlicht: (2024)
von: Jo, Hyun-rae, et al.
Veröffentlicht: (2024)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
von: Abillama, Pierre, et al.
Veröffentlicht: (2025)
von: Abillama, Pierre, et al.
Veröffentlicht: (2025)
Skip2-LoRA: A Lightweight On-device DNN Fine-tuning Method for Low-cost Edge Devices
von: Matsutani, Hiroki, et al.
Veröffentlicht: (2024)
von: Matsutani, Hiroki, et al.
Veröffentlicht: (2024)
DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with Pruning
von: Han, Lixiang, et al.
Veröffentlicht: (2024)
von: Han, Lixiang, et al.
Veröffentlicht: (2024)
MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
von: Zou, Xingze, et al.
Veröffentlicht: (2026)
von: Zou, Xingze, et al.
Veröffentlicht: (2026)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
von: Le, Qi, et al.
Veröffentlicht: (2025)
von: Le, Qi, et al.
Veröffentlicht: (2025)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
von: Lee, Dongyeop, et al.
Veröffentlicht: (2025)
von: Lee, Dongyeop, et al.
Veröffentlicht: (2025)
AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
SAFFIRA: a Framework for Assessing the Reliability of Systolic-Array-Based DNN Accelerators
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
von: Taheri, Mahdi, et al.
Veröffentlicht: (2024)
DapperFL: Domain Adaptive Federated Learning with Model Fusion Pruning for Edge Devices
von: Jia, Yongzhe, et al.
Veröffentlicht: (2024)
von: Jia, Yongzhe, et al.
Veröffentlicht: (2024)
Benchmarking Mobile Device Control Agents across Diverse Configurations
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Multi-Dimensional Pruning: Joint Channel, Layer and Block Pruning with Latency Constraint
von: Sun, Xinglong, et al.
Veröffentlicht: (2024)
von: Sun, Xinglong, et al.
Veröffentlicht: (2024)
'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators
von: Han, Ruichi, et al.
Veröffentlicht: (2026)
von: Han, Ruichi, et al.
Veröffentlicht: (2026)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
von: Qin, Yifan, et al.
Veröffentlicht: (2023)
von: Qin, Yifan, et al.
Veröffentlicht: (2023)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
von: Cho, Yeseul, et al.
Veröffentlicht: (2025)
von: Cho, Yeseul, et al.
Veröffentlicht: (2025)
HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
von: Zhao, Huaqin, et al.
Veröffentlicht: (2024)
von: Zhao, Huaqin, et al.
Veröffentlicht: (2024)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
von: Choi, Moonseok, et al.
Veröffentlicht: (2023)
von: Choi, Moonseok, et al.
Veröffentlicht: (2023)
Harnessing Neuron Stability to Improve DNN Verification
von: Duong, Hai, et al.
Veröffentlicht: (2024)
von: Duong, Hai, et al.
Veröffentlicht: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
von: Yu, Seungmin, et al.
Veröffentlicht: (2024) -
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2025) -
Cognitive Chunking for Soft Prompts: Accelerating Compressor Learning via Block-wise Causal Masking
von: Liu, Guojie, et al.
Veröffentlicht: (2026) -
Hardware-Aware DNN Compression for Homogeneous Edge Devices
von: Zhang, Kunlong, et al.
Veröffentlicht: (2025) -
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
von: Shing, Makoto, et al.
Veröffentlicht: (2025)