PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hourri, Younes, Mozaffari, Mohammad, Dehnavi, Maryam Mehri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
by: Mozaffari, Mohammad, et al.
Published: (2025)
by: Mozaffari, Mohammad, et al.
Published: (2025)
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
by: Mozaffari, Mohammad, et al.
Published: (2026)
by: Mozaffari, Mohammad, et al.
Published: (2026)
SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates
by: Mozaffari, Mohammad, et al.
Published: (2023)
by: Mozaffari, Mohammad, et al.
Published: (2023)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
LLMs for Analog Circuit Design Continuum (ACDC)
by: Esfandiari, Yasaman, et al.
Published: (2025)
by: Esfandiari, Yasaman, et al.
Published: (2025)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
MoEITS: A Green AI approach for simplifying MoE-LLMs
by: Balderas, Luis, et al.
Published: (2026)
by: Balderas, Luis, et al.
Published: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
by: Yao, Feiyu, et al.
Published: (2026)
by: Yao, Feiyu, et al.
Published: (2026)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Time-Efficient Hybrid Hyperparameter Tuning Approach for Cardiovascular Disease Classification
by: Pathak, Abhay Kumar, et al.
Published: (2024)
by: Pathak, Abhay Kumar, et al.
Published: (2024)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
by: Tran, Nguyen Phuc, et al.
Published: (2025)
by: Tran, Nguyen Phuc, et al.
Published: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
by: Bi, Zhen, et al.
Published: (2026)
by: Bi, Zhen, et al.
Published: (2026)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
by: Hangun, Batuhan, et al.
Published: (2025)
by: Hangun, Batuhan, et al.
Published: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
by: Hossain, Md Arafat, et al.
Published: (2025)
by: Hossain, Md Arafat, et al.
Published: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
by: Yin, Yishu, et al.
Published: (2025)
by: Yin, Yishu, et al.
Published: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
by: Zhao, Yushang, et al.
Published: (2025)
by: Zhao, Yushang, et al.
Published: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
Exchangeability in Neural Network and its Application to Dynamic Pruning
by: Pu, et al.
Published: (2025)
by: Pu, et al.
Published: (2025)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
by: Almurshed, Osama, et al.
Published: (2025)
by: Almurshed, Osama, et al.
Published: (2025)
The Race to Efficiency: A New Perspective on AI Scaling Laws
by: Lu, Chien-Ping
Published: (2025)
by: Lu, Chien-Ping
Published: (2025)
Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
by: Avinash, MSR
Published: (2025)
by: Avinash, MSR
Published: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
AdaGradSelect: An adaptive gradient-guided layer selection method for efficient fine-tuning of SLMs
by: Kumar, Anshul, et al.
Published: (2025)
by: Kumar, Anshul, et al.
Published: (2025)
On the Sustainability of AI Inferences in the Edge
by: Sobhani, Ghazal, et al.
Published: (2025)
by: Sobhani, Ghazal, et al.
Published: (2025)
Knowledge Distillation for Reservoir-based Classifier: Human Activity Recognition
by: Kagiyama, Masaharu, et al.
Published: (2025)
by: Kagiyama, Masaharu, et al.
Published: (2025)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
by: Estevez, Melissa, et al.
Published: (2025)
by: Estevez, Melissa, et al.
Published: (2025)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025)
by: Yi, Qingao, et al.
Published: (2025)
Similar Items
-
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024) -
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
by: Mozaffari, Mohammad, et al.
Published: (2025) -
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
by: Mozaffari, Mohammad, et al.
Published: (2026) -
SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs
by: Mozaffari, Mohammad, et al.
Published: (2024) -
MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates
by: Mozaffari, Mohammad, et al.
Published: (2023)