FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
Fuente:
arXiv
Salvato in:
| Autori principali: | Baek, Daehyeon, Choi, Jieun, Son, Jimyoung, Bin, Kyungmin, Choi, Seungbeom, Moon, Kihyo, Jang, Minsung, Lee, Hyojung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
LaDiMo: Layer-wise Distillation Inspired MoEfier
di: Kim, Sungyoon, et al.
Pubblicazione: (2024)
di: Kim, Sungyoon, et al.
Pubblicazione: (2024)
RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attribution
di: Picov, Isaac, et al.
Pubblicazione: (2026)
di: Picov, Isaac, et al.
Pubblicazione: (2026)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
di: Qiao, Ye, et al.
Pubblicazione: (2025)
di: Qiao, Ye, et al.
Pubblicazione: (2025)
ReRoPE: Repurposing RoPE for Relative Camera Control
di: Li, Chunyang, et al.
Pubblicazione: (2026)
di: Li, Chunyang, et al.
Pubblicazione: (2026)
Periodic RoPE for Infinite Context LLMs
di: Huo, Simin
Pubblicazione: (2026)
di: Huo, Simin
Pubblicazione: (2026)
Base of RoPE Bounds Context Length
di: Men, Xin, et al.
Pubblicazione: (2024)
di: Men, Xin, et al.
Pubblicazione: (2024)
Scaling Laws of RoPE-based Extrapolation
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
di: Choi, Seungbeom, et al.
Pubblicazione: (2025)
di: Choi, Seungbeom, et al.
Pubblicazione: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
di: Cheng, Yuan, et al.
Pubblicazione: (2026)
di: Cheng, Yuan, et al.
Pubblicazione: (2026)
TurboQuant on Quantized Models: Solving Compound Quantization Error with Pre-RoPE Compression and KV Compaction
di: Gökyıldız, Onur
Pubblicazione: (2026)
di: Gökyıldız, Onur
Pubblicazione: (2026)
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Frayed RoPE and Long Inputs: A Geometric Perspective
di: Wertheimer, Davis, et al.
Pubblicazione: (2026)
di: Wertheimer, Davis, et al.
Pubblicazione: (2026)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
di: Qiao, Ye, et al.
Pubblicazione: (2025)
di: Qiao, Ye, et al.
Pubblicazione: (2025)
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
di: Pan, Dayan, et al.
Pubblicazione: (2025)
di: Pan, Dayan, et al.
Pubblicazione: (2025)
Limits of Rank Recovery in Bilinear Observation Problems
di: Choi, Seungbeom
Pubblicazione: (2026)
di: Choi, Seungbeom
Pubblicazione: (2026)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
di: Xin, Jihao, et al.
Pubblicazione: (2026)
di: Xin, Jihao, et al.
Pubblicazione: (2026)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
di: Mikaeili, Aryan, et al.
Pubblicazione: (2026)
di: Mikaeili, Aryan, et al.
Pubblicazione: (2026)
A Circular Argument : Does RoPE need to be Equivariant for Vision?
di: van de Geijn, Chase, et al.
Pubblicazione: (2025)
di: van de Geijn, Chase, et al.
Pubblicazione: (2025)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
di: Zhong, Meizhi, et al.
Pubblicazione: (2024)
di: Zhong, Meizhi, et al.
Pubblicazione: (2024)
On the token distance modeling ability of higher RoPE attention dimension
di: Hong, Xiangyu, et al.
Pubblicazione: (2024)
di: Hong, Xiangyu, et al.
Pubblicazione: (2024)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
di: Yang, Ning, et al.
Pubblicazione: (2026)
di: Yang, Ning, et al.
Pubblicazione: (2026)
Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
di: Wei, Xiaokai, et al.
Pubblicazione: (2025)
di: Wei, Xiaokai, et al.
Pubblicazione: (2025)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
di: Li, Haoran, et al.
Pubblicazione: (2026)
di: Li, Haoran, et al.
Pubblicazione: (2026)
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
di: Gokmen, Ahmet Berke, et al.
Pubblicazione: (2025)
di: Gokmen, Ahmet Berke, et al.
Pubblicazione: (2025)
Rethinking RoPE: A Mathematical Blueprint for N-dimensional Positional Embedding
di: Liu, Haiping, et al.
Pubblicazione: (2025)
di: Liu, Haiping, et al.
Pubblicazione: (2025)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
di: Wang, Suyuchen, et al.
Pubblicazione: (2024)
di: Wang, Suyuchen, et al.
Pubblicazione: (2024)
Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
di: Alman, Josh, et al.
Pubblicazione: (2025)
di: Alman, Josh, et al.
Pubblicazione: (2025)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
di: Li, Zichong, et al.
Pubblicazione: (2026)
di: Li, Zichong, et al.
Pubblicazione: (2026)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2026)
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2026)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
di: Du, Yufeng, et al.
Pubblicazione: (2026)
di: Du, Yufeng, et al.
Pubblicazione: (2026)
Modeling the Spatiotemporal Spread and Control of African Swine Fever in the Republic of Korea Using a Patch‐Based Stochastic Framework
di: Changdae Son, et al.
Pubblicazione: (2026)
di: Changdae Son, et al.
Pubblicazione: (2026)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
di: Liu, Yuxi, et al.
Pubblicazione: (2026)
di: Liu, Yuxi, et al.
Pubblicazione: (2026)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration
di: Tan, Shihong, et al.
Pubblicazione: (2026)
di: Tan, Shihong, et al.
Pubblicazione: (2026)
Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
di: Kung, Gustavo Chau Loo, et al.
Pubblicazione: (2026)
di: Kung, Gustavo Chau Loo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
di: Bin, Kyungmin, et al.
Pubblicazione: (2025) -
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
di: Yang, Mingyu, et al.
Pubblicazione: (2025) -
LaDiMo: Layer-wise Distillation Inspired MoEfier
di: Kim, Sungyoon, et al.
Pubblicazione: (2024) -
RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attribution
di: Picov, Isaac, et al.
Pubblicazione: (2026) -
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
di: Qiao, Ye, et al.
Pubblicazione: (2025)