Hardware-Aware DNN Compression for Homogeneous Edge Devices
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Kunlong, Li, Guiying, Lu, Ning, Yang, Peng, Tang, Ke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hardware-Aware DNN Compression for Homogeneous Edge Devices
por: Zhang, Kunlong, et al.
Publicado: (2025)
por: Zhang, Kunlong, et al.
Publicado: (2025)
SCAN-Edge: Finding MobileNet-speed Hybrid Networks for Diverse Edge Devices via Hardware-Aware Evolutionary Search
por: Chiang, Hung-Yueh, et al.
Publicado: (2024)
por: Chiang, Hung-Yueh, et al.
Publicado: (2024)
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
por: Wu, Jiahao, et al.
Publicado: (2026)
por: Wu, Jiahao, et al.
Publicado: (2026)
Skip2-LoRA: A Lightweight On-device DNN Fine-tuning Method for Low-cost Edge Devices
por: Matsutani, Hiroki, et al.
Publicado: (2024)
por: Matsutani, Hiroki, et al.
Publicado: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
por: Wang, Kun, et al.
Publicado: (2024)
por: Wang, Kun, et al.
Publicado: (2024)
Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
por: Cheng, Zhipeng, et al.
Publicado: (2025)
por: Cheng, Zhipeng, et al.
Publicado: (2025)
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
por: Wang, Ziyi, et al.
Publicado: (2025)
por: Wang, Ziyi, et al.
Publicado: (2025)
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
por: Lee, Hayun, et al.
Publicado: (2024)
por: Lee, Hayun, et al.
Publicado: (2024)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
por: Xie, Jianhang, et al.
Publicado: (2025)
por: Xie, Jianhang, et al.
Publicado: (2025)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
por: Zhang, Tuo, et al.
Publicado: (2025)
por: Zhang, Tuo, et al.
Publicado: (2025)
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
por: Jiang, Xinting, et al.
Publicado: (2026)
por: Jiang, Xinting, et al.
Publicado: (2026)
Recall: Empowering Multimodal Embedding for Edge Devices
por: Cai, Dongqi, et al.
Publicado: (2024)
por: Cai, Dongqi, et al.
Publicado: (2024)
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
por: Xiao, Jinqi, et al.
Publicado: (2023)
por: Xiao, Jinqi, et al.
Publicado: (2023)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
por: Ramachandran, Akshat, et al.
Publicado: (2024)
por: Ramachandran, Akshat, et al.
Publicado: (2024)
From LLMs to Edge: Parameter-Efficient Fine-Tuning on Edge Devices
por: Slamanig, Georg, et al.
Publicado: (2025)
por: Slamanig, Georg, et al.
Publicado: (2025)
Multi-Objective Hardware Aware Neural Architecture Search using Hardware Cost Diversity
por: Sinha, Nilotpal, et al.
Publicado: (2024)
por: Sinha, Nilotpal, et al.
Publicado: (2024)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
por: Yu, Xiaoming, et al.
Publicado: (2026)
por: Yu, Xiaoming, et al.
Publicado: (2026)
Compressing Neural Networks Using Tensor Networks with Exponentially Fewer Variational Parameters
por: Qing, Yong, et al.
Publicado: (2023)
por: Qing, Yong, et al.
Publicado: (2023)
Collaborative Compression for Large-Scale MoE Deployment on Edge
por: Chen, Yixiao, et al.
Publicado: (2025)
por: Chen, Yixiao, et al.
Publicado: (2025)
AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN Accelerators
por: Taheri, Mahdi, et al.
Publicado: (2024)
por: Taheri, Mahdi, et al.
Publicado: (2024)
A&B BNN: Add&Bit-Operation-Only Hardware-Friendly Binary Neural Network
por: Ma, Ruichen, et al.
Publicado: (2024)
por: Ma, Ruichen, et al.
Publicado: (2024)
HASTE: Hardware-Aware Dynamic Sparse Training for Large Output Spaces
por: Ullah, Nasib, et al.
Publicado: (2026)
por: Ullah, Nasib, et al.
Publicado: (2026)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference Systems
por: Zhou, Ao, et al.
Publicado: (2024)
por: Zhou, Ao, et al.
Publicado: (2024)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
por: Tan, Qitao, et al.
Publicado: (2025)
por: Tan, Qitao, et al.
Publicado: (2025)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
por: Gong, Ping, et al.
Publicado: (2025)
por: Gong, Ping, et al.
Publicado: (2025)
Robust Federated Learning on Edge Devices with Domain Heterogeneity
por: Le, Huy Q., et al.
Publicado: (2025)
por: Le, Huy Q., et al.
Publicado: (2025)
CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation
por: Yang, Ning, et al.
Publicado: (2026)
por: Yang, Ning, et al.
Publicado: (2026)
MetaCLBench: Meta Continual Learning Benchmark on Resource-Constrained Edge Devices
por: Li, Sijia, et al.
Publicado: (2025)
por: Li, Sijia, et al.
Publicado: (2025)
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
por: Bao, Rui, et al.
Publicado: (2025)
por: Bao, Rui, et al.
Publicado: (2025)
Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge Deployment
por: Ji, Yuhao, et al.
Publicado: (2024)
por: Ji, Yuhao, et al.
Publicado: (2024)
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
por: Qin, Ruiyang, et al.
Publicado: (2024)
por: Qin, Ruiyang, et al.
Publicado: (2024)
Hardware Aware Ensemble Selection for Balancing Predictive Accuracy and Cost
por: Maier, Jannis, et al.
Publicado: (2024)
por: Maier, Jannis, et al.
Publicado: (2024)
DNNShifter: An Efficient DNN Pruning System for Edge Computing
por: Eccles, Bailey J., et al.
Publicado: (2023)
por: Eccles, Bailey J., et al.
Publicado: (2023)
MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
por: Li, Guangyan, et al.
Publicado: (2025)
por: Li, Guangyan, et al.
Publicado: (2025)
LoCA: Location-Aware Cosine Adaptation for Parameter-Efficient Fine-Tuning
por: Du, Zhekai, et al.
Publicado: (2025)
por: Du, Zhekai, et al.
Publicado: (2025)
GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-Search
por: Zhou, Ao, et al.
Publicado: (2025)
por: Zhou, Ao, et al.
Publicado: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
por: Huang, Wei, et al.
Publicado: (2023)
por: Huang, Wei, et al.
Publicado: (2023)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
por: Zhang, Jun, et al.
Publicado: (2025)
por: Zhang, Jun, et al.
Publicado: (2025)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
por: Ran, Dezhi, et al.
Publicado: (2025)
por: Ran, Dezhi, et al.
Publicado: (2025)
Ejemplares similares
-
Hardware-Aware DNN Compression for Homogeneous Edge Devices
por: Zhang, Kunlong, et al.
Publicado: (2025) -
SCAN-Edge: Finding MobileNet-speed Hybrid Networks for Diverse Edge Devices via Hardware-Aware Evolutionary Search
por: Chiang, Hung-Yueh, et al.
Publicado: (2024) -
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
por: Wu, Jiahao, et al.
Publicado: (2026) -
Skip2-LoRA: A Lightweight On-device DNN Fine-tuning Method for Low-cost Edge Devices
por: Matsutani, Hiroki, et al.
Publicado: (2024) -
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
por: Wang, Kun, et al.
Publicado: (2024)