LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Xi, Haocheng, Singh, Harman, Hu, Yuezhou, Hooper, Coleman, Tiwari, Rishabh, Tomar, Aditya, Lee, Minjae, Kang, Wonjun, Mahoney, Michael, Xu, Chenfeng, Keutzer, Kurt, Gholami, Amir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
Residual Context Diffusion Language Models
di: Hu, Yuezhou, et al.
Pubblicazione: (2026)
di: Hu, Yuezhou, et al.
Pubblicazione: (2026)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)
CDLM: Consistency Diffusion Language Models For Faster Sampling
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
SciML Agents: Write the Solver, Not the Solution
di: Gaonkar, Saarth, et al.
Pubblicazione: (2025)
di: Gaonkar, Saarth, et al.
Pubblicazione: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
AI and Memory Wall
di: Gholami, Amir, et al.
Pubblicazione: (2024)
di: Gholami, Amir, et al.
Pubblicazione: (2024)
Multipole Attention for Efficient Long Context Reasoning
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
SPEED: Speculative Pipelined Execution for Efficient Decoding
di: Hooper, Coleman, et al.
Pubblicazione: (2023)
di: Hooper, Coleman, et al.
Pubblicazione: (2023)
ETS: Efficient Tree Search for Inference-Time Scaling
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
di: Hooper, Coleman, et al.
Pubblicazione: (2026)
di: Hooper, Coleman, et al.
Pubblicazione: (2026)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
di: Yang, Shuo, et al.
Pubblicazione: (2025)
di: Yang, Shuo, et al.
Pubblicazione: (2025)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2026)
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2026)
Learned Best-Effort LLM Serving
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
di: Khaki, Samir, et al.
Pubblicazione: (2025)
di: Khaki, Samir, et al.
Pubblicazione: (2025)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
di: Subramanian, Shashank, et al.
Pubblicazione: (2023)
di: Subramanian, Shashank, et al.
Pubblicazione: (2023)
An LLM Compiler for Parallel Function Calling
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
TinyAgent: Function Calling at the Edge
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024)
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
di: Zhang, Jintao, et al.
Pubblicazione: (2026)
di: Zhang, Jintao, et al.
Pubblicazione: (2026)
Characterizing Prompt Compression Methods for Long Context Inference
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
Agentic Test-Time Scaling for WebAgents
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
di: Kang, Wonjun, et al.
Pubblicazione: (2025)
di: Kang, Wonjun, et al.
Pubblicazione: (2025)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
di: Sheng, Ying, et al.
Pubblicazione: (2023)
di: Sheng, Ying, et al.
Pubblicazione: (2023)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
di: Liu, Zhijian, et al.
Pubblicazione: (2024)
di: Liu, Zhijian, et al.
Pubblicazione: (2024)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
di: Lee, Nicholas, et al.
Pubblicazione: (2024)
di: Lee, Nicholas, et al.
Pubblicazione: (2024)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
di: Li, Xingyang, et al.
Pubblicazione: (2025)
di: Li, Xingyang, et al.
Pubblicazione: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
di: Xi, Haocheng, et al.
Pubblicazione: (2025)
di: Xi, Haocheng, et al.
Pubblicazione: (2025)
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
di: Peng, Chensheng, et al.
Pubblicazione: (2024)
di: Peng, Chensheng, et al.
Pubblicazione: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
di: Moon, Suhong, et al.
Pubblicazione: (2024)
di: Moon, Suhong, et al.
Pubblicazione: (2024)
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
di: Li, Yiheng, et al.
Pubblicazione: (2024)
di: Li, Yiheng, et al.
Pubblicazione: (2024)
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
di: Li, Yiheng, et al.
Pubblicazione: (2025)
di: Li, Yiheng, et al.
Pubblicazione: (2025)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
di: Liang, Feng, et al.
Pubblicazione: (2024)
di: Liang, Feng, et al.
Pubblicazione: (2024)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
di: Wang, Qinsi, et al.
Pubblicazione: (2025)
di: Wang, Qinsi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
di: Tomar, Aditya, et al.
Pubblicazione: (2025) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025) -
Residual Context Diffusion Language Models
di: Hu, Yuezhou, et al.
Pubblicazione: (2026) -
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
di: Kim, Minseo, et al.
Pubblicazione: (2025) -
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)