Effective Interplay between Sparsity and Quantization: From Theory to Practice
Fuente:
arXiv
Guardado en:
| Autores principales: | Harma, Simla Burcu, Chakraborty, Ayan, Kostenok, Elizaveta, Mishin, Danila, Ha, Dongho, Falsafi, Babak, Jaggi, Martin, Liu, Ming, Oh, Yunho, Subramanian, Suvinay, Yazdanbakhsh, Amir |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training
por: Harma, Simla Burcu, et al.
Publicado: (2022)
por: Harma, Simla Burcu, et al.
Publicado: (2022)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2024)
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2024)
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
por: Jiang, Wenqi, et al.
Publicado: (2025)
por: Jiang, Wenqi, et al.
Publicado: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
por: Mozaffari, Mohammad, et al.
Publicado: (2024)
por: Mozaffari, Mohammad, et al.
Publicado: (2024)
Uncertainty Estimation of Transformers' Predictions via Topological Analysis of the Attention Matrices
por: Kostenok, Elizaveta, et al.
Publicado: (2023)
por: Kostenok, Elizaveta, et al.
Publicado: (2023)
Calibration-Reasoning Framework for Descriptive Speech Quality Assessment
por: Kostenok, Elizaveta, et al.
Publicado: (2026)
por: Kostenok, Elizaveta, et al.
Publicado: (2026)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
por: Yazdanbakhsh, Amir
Publicado: (2025)
por: Yazdanbakhsh, Amir
Publicado: (2025)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
por: Jin, Tian, et al.
Publicado: (2025)
por: Jin, Tian, et al.
Publicado: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
por: Vishwanathan, Manoj, et al.
Publicado: (2026)
por: Vishwanathan, Manoj, et al.
Publicado: (2026)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
por: Durvasula, Sankeerth, et al.
Publicado: (2025)
por: Durvasula, Sankeerth, et al.
Publicado: (2025)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
por: Jin, Tian, et al.
Publicado: (2025)
por: Jin, Tian, et al.
Publicado: (2025)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
por: Choi, Sangun, et al.
Publicado: (2025)
por: Choi, Sangun, et al.
Publicado: (2025)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2025)
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2025)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
por: Ding, Shaojin, et al.
Publicado: (2023)
por: Ding, Shaojin, et al.
Publicado: (2023)
Exploring the Sparsity-Quantization Interplay on a Novel Hybrid SNN Event-Driven Architecture
por: Aliyev, Ilkin, et al.
Publicado: (2024)
por: Aliyev, Ilkin, et al.
Publicado: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
por: Wu, Hanjiang, et al.
Publicado: (2026)
por: Wu, Hanjiang, et al.
Publicado: (2026)
The use of general anesthesia for dental treatment of children with special healthcare needs in Alberta, Canada
por: Elnaz Yazdanbakhsh, et al.
Publicado: (2024)
por: Elnaz Yazdanbakhsh, et al.
Publicado: (2024)
Tao: Re-Thinking DL-based Microarchitecture Simulation
por: Pandey, Santosh, et al.
Publicado: (2024)
por: Pandey, Santosh, et al.
Publicado: (2024)
On a relationship between grain boundary free energy, grain boundary segregation, and grain boundary diffusion
por: Mishin, Yuri
Publicado: (2026)
por: Mishin, Yuri
Publicado: (2026)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
por: Sengupta, Ayan, et al.
Publicado: (2025)
por: Sengupta, Ayan, et al.
Publicado: (2025)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
por: Khaki, Samir, et al.
Publicado: (2025)
por: Khaki, Samir, et al.
Publicado: (2025)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
por: You, Haoran, et al.
Publicado: (2024)
por: You, Haoran, et al.
Publicado: (2024)
Capacity Matters: a Proof-of-Concept for Transformer Memorization on Real-World Data
por: Changalidis, Anton, et al.
Publicado: (2025)
por: Changalidis, Anton, et al.
Publicado: (2025)
Triangular Decomposition of the Crystal Lattice of Quantized Function Algebras: Revisited
por: Dey, Ayan
Publicado: (2026)
por: Dey, Ayan
Publicado: (2026)
Spark Transformer: Reactivating Sparsity in FFN and Attention
por: You, Chong, et al.
Publicado: (2025)
por: You, Chong, et al.
Publicado: (2025)
Bycatch of marine mammals in the Northwest Atlantic during commercial fishery (based on literature materials and observations by the Polar branch of VNIRO in 2013–2020)
por: Mishin, Т. V.
Publicado: (2022)
por: Mishin, Т. V.
Publicado: (2022)
Green Dentistry Education in the Post‐Minamata Era: Knowledge, Attitudes, and Practices in Türkiye
por: Mehmet Buldur, et al.
Publicado: (2026)
por: Mehmet Buldur, et al.
Publicado: (2026)
On the Interplay Between Sparsity and Training in Deep Reinforcement Learning
por: Davelouis, Fatima, et al.
Publicado: (2025)
por: Davelouis, Fatima, et al.
Publicado: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
por: Panferov, Andrei, et al.
Publicado: (2026)
por: Panferov, Andrei, et al.
Publicado: (2026)
On the Interplay of Privacy, Persuasion and Quantization
por: Anand, Anju, et al.
Publicado: (2025)
por: Anand, Anju, et al.
Publicado: (2025)
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning
por: Oh, Minjae, et al.
Publicado: (2025)
por: Oh, Minjae, et al.
Publicado: (2025)
Illposedness via degenerate dispersion for generalized surface quasi-geostrophic equations with singular velocities
por: Chae, Dongho, et al.
Publicado: (2023)
por: Chae, Dongho, et al.
Publicado: (2023)
Enabling Dynamic Sparsity in Quantized LLM Inference
por: Wang, Rongxiang, et al.
Publicado: (2025)
por: Wang, Rongxiang, et al.
Publicado: (2025)
Compression Scaling Laws:Unifying Sparsity and Quantization
por: Frantar, Elias, et al.
Publicado: (2025)
por: Frantar, Elias, et al.
Publicado: (2025)
Hodge Splittings and Einstein 4-manifolds
por: Aazami, Amir Babak
Publicado: (2025)
por: Aazami, Amir Babak
Publicado: (2025)
The Eisenhart Lift and Hamiltonian Systems
por: Aazami, Amir Babak
Publicado: (2024)
por: Aazami, Amir Babak
Publicado: (2024)
On the Petrov Type of a 4-manifold
por: Aazami, Amir Babak
Publicado: (2023)
por: Aazami, Amir Babak
Publicado: (2023)
Petrov Types for the Weyl Tensor via the Riemannian-to-Lorentzian Bridge
por: Aazami, Amir Babak
Publicado: (2024)
por: Aazami, Amir Babak
Publicado: (2024)
Geometry via Plane wave limits
por: Aazami, Amir Babak
Publicado: (2024)
por: Aazami, Amir Babak
Publicado: (2024)
On the Curvature and Topology of Compact Stationary Spacetimes
por: Aazami, Amir Babak
Publicado: (2025)
por: Aazami, Amir Babak
Publicado: (2025)
Ejemplares similares
-
Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training
por: Harma, Simla Burcu, et al.
Publicado: (2022) -
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2024) -
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
por: Jiang, Wenqi, et al.
Publicado: (2025) -
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
por: Mozaffari, Mohammad, et al.
Publicado: (2024) -
Uncertainty Estimation of Transformers' Predictions via Topological Analysis of the Attention Matrices
por: Kostenok, Elizaveta, et al.
Publicado: (2023)