From a Lossless (~1.5:1) Compression Algorithm for Llama2 7B Weights to Variable Precision, Variable Range, Compressed Numeric Data Types for CNNs and LLMs
Fuente:
arXiv
Saved in:
| Main Author: | Liguori, Vincenzo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Procrastination Is All You Need: Exponent Indexed Accumulators for Floating Point, Posits and Logarithmic Numbers
by: Liguori, Vincenzo
Published: (2024)
by: Liguori, Vincenzo
Published: (2024)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Fair and Square: Replacing One Real Multiplication with a Single Square and One Complex Multiplication with Three Squares When Performing Matrix Multiplication and Convolutions
by: Liguori, Vincenzo
Published: (2026)
by: Liguori, Vincenzo
Published: (2026)
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
by: Li, Shuaiting, et al.
Published: (2024)
by: Li, Shuaiting, et al.
Published: (2024)
Efficient FIR filtering with Bit Layer Multiply Accumulator
by: Liguori, Vincenzo
Published: (2024)
by: Liguori, Vincenzo
Published: (2024)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
RAS: A Bit-Exact rANS Accelerator For High-Performance Neural Lossless Compression
by: Qin, Yuchao, et al.
Published: (2025)
by: Qin, Yuchao, et al.
Published: (2025)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
by: Yubeaton, Patrick, et al.
Published: (2025)
by: Yubeaton, Patrick, et al.
Published: (2025)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
by: Toupas, Petros, et al.
Published: (2024)
by: Toupas, Petros, et al.
Published: (2024)
GLANCE: Gaze-Led Attention Network for Compressed Edge-inference
by: Solanki, Neeraj, et al.
Published: (2026)
by: Solanki, Neeraj, et al.
Published: (2026)
Dedicated Inference Engine and Binary-Weight Neural Networks for Lightweight Instance Segmentation
by: Chen, Tse-Wei, et al.
Published: (2025)
by: Chen, Tse-Wei, et al.
Published: (2025)
ASC: Adaptive Scale Feature Map Compression for Deep Neural Network
by: Yao, Yuan, et al.
Published: (2023)
by: Yao, Yuan, et al.
Published: (2023)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024)
by: Xu, Han, et al.
Published: (2024)
LEXI: Lossless Exponent Coding for Efficient Inter-Chiplet Communication in Hybrid LLMs
by: Sun, Miao, et al.
Published: (2026)
by: Sun, Miao, et al.
Published: (2026)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
by: Fu, Yonggan, et al.
Published: (2023)
by: Fu, Yonggan, et al.
Published: (2023)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
by: Huang, Sixiao, et al.
Published: (2025)
by: Huang, Sixiao, et al.
Published: (2025)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
by: Ferry, Corentin, et al.
Published: (2024)
by: Ferry, Corentin, et al.
Published: (2024)
Precomputed 1D-CNNs for Atrial Fibrillation Detection on Tiny Smart Sensor Systems
by: Einhaus, Lukas, et al.
Published: (2026)
by: Einhaus, Lukas, et al.
Published: (2026)
M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
by: Hu, Weiming, et al.
Published: (2025)
by: Hu, Weiming, et al.
Published: (2025)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
by: Zhang, Kunlong, et al.
Published: (2025)
by: Zhang, Kunlong, et al.
Published: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026)
by: Dong, Pingcheng, et al.
Published: (2026)
Real-Time Spacecraft Pose Estimation Using Mixed-Precision Quantized Neural Network on COTS Reconfigurable MPSoC
by: Posso, Julien, et al.
Published: (2024)
by: Posso, Julien, et al.
Published: (2024)
Taming Performance Variability caused by Client-Side Hardware Configuration
by: Antoniou, Georgia, et al.
Published: (2024)
by: Antoniou, Georgia, et al.
Published: (2024)
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
by: Oh, Hyunwoo, et al.
Published: (2025)
by: Oh, Hyunwoo, et al.
Published: (2025)
Implementation and Evaluation of GBDI Memory Compression Algorithm Using C/C++ on a Broader Range of Workloads
by: Aina, Adeyemi
Published: (2025)
by: Aina, Adeyemi
Published: (2025)
A High-Throughput Hardware Accelerator for Lempel-Ziv 4 Compression Algorithm
by: Chen, Tao, et al.
Published: (2024)
by: Chen, Tao, et al.
Published: (2024)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
by: Zhao, Zhiyuan, et al.
Published: (2024)
by: Zhao, Zhiyuan, et al.
Published: (2024)
AppSign: Multi-level Approximate Computing for Real-Time Traffic Sign Recognition in Autonomous Vehicles
by: Omidian, Fatemeh, et al.
Published: (2024)
by: Omidian, Fatemeh, et al.
Published: (2024)
Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
by: Carrigg, Kieran, et al.
Published: (2026)
by: Carrigg, Kieran, et al.
Published: (2026)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
by: Zhang, Baoheng, et al.
Published: (2024)
by: Zhang, Baoheng, et al.
Published: (2024)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
Primitive-Driven Acceleration of Hyperdimensional Computing for Real-Time Image Classification
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
by: Miao, Wenxuan, et al.
Published: (2025)
by: Miao, Wenxuan, et al.
Published: (2025)
Vision Transformer Computation and Resilience for Dynamic Inference
by: Sreedhar, Kavya, et al.
Published: (2022)
by: Sreedhar, Kavya, et al.
Published: (2022)
Identifying Unnecessary 3D Gaussians using Clustering for Fast Rendering of 3D Gaussian Splatting
by: Jo, Joongho, et al.
Published: (2024)
by: Jo, Joongho, et al.
Published: (2024)
Real-Time Object Detection and Classification using YOLO for Edge FPGAs
by: Amin, Rashed Al, et al.
Published: (2025)
by: Amin, Rashed Al, et al.
Published: (2025)
SteROI-D: System Design and Mapping for Stereo Depth Inference on Regions of Interest
by: Erhardt, Jack, et al.
Published: (2025)
by: Erhardt, Jack, et al.
Published: (2025)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
by: Durvasula, Sankeerth, et al.
Published: (2025)
by: Durvasula, Sankeerth, et al.
Published: (2025)
Similar Items
-
Procrastination Is All You Need: Exponent Indexed Accumulators for Floating Point, Posits and Logarithmic Numbers
by: Liguori, Vincenzo
Published: (2024) -
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025) -
Fair and Square: Replacing One Real Multiplication with a Single Square and One Complex Multiplication with Three Squares When Performing Matrix Multiplication and Convolutions
by: Liguori, Vincenzo
Published: (2026) -
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
by: Li, Shuaiting, et al.
Published: (2024) -
Efficient FIR filtering with Bit Layer Multiply Accumulator
by: Liguori, Vincenzo
Published: (2024)