TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
Fuente:
arXiv
Salvato in:
| Autori principali: | Gu, Hongyaoxing, Chen, Xinzhe, Hu, Lijuan, Liu, Fangfang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026)
LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
A method of using RSVD in residual calculation of LowBit GEMM
di: Gu, Hongyaoxing
Pubblicazione: (2024)
di: Gu, Hongyaoxing
Pubblicazione: (2024)
Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
TileLang: A Composable Tiled Programming Model for AI Systems
di: Wang, Lei, et al.
Pubblicazione: (2025)
di: Wang, Lei, et al.
Pubblicazione: (2025)
The Tile: A 2D Map of Ranking Scores for Two-Class Classification
di: Piérard, Sébastien, et al.
Pubblicazione: (2024)
di: Piérard, Sébastien, et al.
Pubblicazione: (2024)
TiledAttention: a CUDA Tile SDPA Kernel for PyTorch
di: Khan, Taimur
Pubblicazione: (2026)
di: Khan, Taimur
Pubblicazione: (2026)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
di: Ding, Yifu, et al.
Pubblicazione: (2026)
di: Ding, Yifu, et al.
Pubblicazione: (2026)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
di: Chowdhury, Mohammed Nowaz Rabbani, et al.
Pubblicazione: (2026)
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
di: Huang, Beichen, et al.
Pubblicazione: (2025)
di: Huang, Beichen, et al.
Pubblicazione: (2025)
Rhomboid Tiling for Geometric Graph Deep Learning
di: Zhang, Yipeng, et al.
Pubblicazione: (2025)
di: Zhang, Yipeng, et al.
Pubblicazione: (2025)
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels
di: Zhao, Yifan, et al.
Pubblicazione: (2026)
di: Zhao, Yifan, et al.
Pubblicazione: (2026)
An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
di: Li, Tseng-Jen, et al.
Pubblicazione: (2025)
di: Li, Tseng-Jen, et al.
Pubblicazione: (2025)
Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation
di: Ding, Yaoyao, et al.
Pubblicazione: (2025)
di: Ding, Yaoyao, et al.
Pubblicazione: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
di: Wang, Wenfeng, et al.
Pubblicazione: (2025)
di: Wang, Wenfeng, et al.
Pubblicazione: (2025)
HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
di: Wei, Jia, et al.
Pubblicazione: (2026)
di: Wei, Jia, et al.
Pubblicazione: (2026)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts
di: Yang, Minghao, et al.
Pubblicazione: (2026)
di: Yang, Minghao, et al.
Pubblicazione: (2026)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
di: Bouquet, Yann, et al.
Pubblicazione: (2026)
di: Bouquet, Yann, et al.
Pubblicazione: (2026)
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
di: Yin, Xiangyang, et al.
Pubblicazione: (2026)
di: Yin, Xiangyang, et al.
Pubblicazione: (2026)
Neural-Schwarz Tiling for Geometry-Universal PDE Solving at Scale
di: Secchi, Paolo, et al.
Pubblicazione: (2026)
di: Secchi, Paolo, et al.
Pubblicazione: (2026)
Machine Learning Toric Duality in Brane Tilings
di: Capuozzo, Pietro, et al.
Pubblicazione: (2024)
di: Capuozzo, Pietro, et al.
Pubblicazione: (2024)
SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling
di: Ye, Fanjiang, et al.
Pubblicazione: (2025)
di: Ye, Fanjiang, et al.
Pubblicazione: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
di: Beck, Maximilian, et al.
Pubblicazione: (2025)
di: Beck, Maximilian, et al.
Pubblicazione: (2025)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
di: Truong, Tuan, et al.
Pubblicazione: (2025)
di: Truong, Tuan, et al.
Pubblicazione: (2025)
MoRE: A Mixture of Low-Rank Experts for Adaptive Multi-Task Learning
di: Zhang, Dacao, et al.
Pubblicazione: (2025)
di: Zhang, Dacao, et al.
Pubblicazione: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
di: He, Yifei, et al.
Pubblicazione: (2025)
di: He, Yifei, et al.
Pubblicazione: (2025)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
di: Hourri, Younes, et al.
Pubblicazione: (2025)
di: Hourri, Younes, et al.
Pubblicazione: (2025)
MoPEQ: Mixture of Mixed Precision Quantized Experts
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
di: Yang, Cheng, et al.
Pubblicazione: (2024)
di: Yang, Cheng, et al.
Pubblicazione: (2024)
A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models
di: Sun, Mengyang, et al.
Pubblicazione: (2025)
di: Sun, Mengyang, et al.
Pubblicazione: (2025)
Design-Specification Tiling for ICL-based CAD Code Generation
di: Du, Yali, et al.
Pubblicazione: (2026)
di: Du, Yali, et al.
Pubblicazione: (2026)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2026)
Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism
di: Sun, Mengyang, et al.
Pubblicazione: (2026)
di: Sun, Mengyang, et al.
Pubblicazione: (2026)
Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling
di: Kwon, Young D., et al.
Pubblicazione: (2025)
di: Kwon, Young D., et al.
Pubblicazione: (2025)
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
di: Liu, Zijie, et al.
Pubblicazione: (2026)
di: Liu, Zijie, et al.
Pubblicazione: (2026)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
di: Hu, Xing, et al.
Pubblicazione: (2025)
di: Hu, Xing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026) -
LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026) -
A method of using RSVD in residual calculation of LowBit GEMM
di: Gu, Hongyaoxing
Pubblicazione: (2024) -
Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
di: Liu, Zhenyu, et al.
Pubblicazione: (2025) -
TileLang: A Composable Tiled Programming Model for AI Systems
di: Wang, Lei, et al.
Pubblicazione: (2025)