vAttention: Verified Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Desai, Aditya, Agrawal, Kumar Krishna, Yang, Shuo, Cuadron, Alejandro, Schroeder, Luis Gaspar, Zaharia, Matei, Gonzalez, Joseph E., Stoica, Ion |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
vCache: Verified Semantic Prompt Caching
by: Schroeder, Luis Gaspar, et al.
Published: (2025)
by: Schroeder, Luis Gaspar, et al.
Published: (2025)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Inference Time Context Sparsity: Illusion or Opportunity?
by: Joshi, Sahil, et al.
Published: (2026)
by: Joshi, Sahil, et al.
Published: (2026)
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
by: Cuadron, Alejandro, et al.
Published: (2025)
by: Cuadron, Alejandro, et al.
Published: (2025)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design
by: Davis, Jared Quincy, et al.
Published: (2024)
by: Davis, Jared Quincy, et al.
Published: (2024)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025)
by: Zhu, Alan, et al.
Published: (2025)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Barbarians at the Gate: How AI is Upending Systems Research
by: Cheng, Audrey, et al.
Published: (2025)
by: Cheng, Audrey, et al.
Published: (2025)
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
by: Agarwal, Shubham, et al.
Published: (2026)
by: Agarwal, Shubham, et al.
Published: (2026)
AI-Driven Research for Databases
by: Cheng, Audrey, et al.
Published: (2026)
by: Cheng, Audrey, et al.
Published: (2026)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
by: Chen, Lingjiao, et al.
Published: (2024)
by: Chen, Lingjiao, et al.
Published: (2024)
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025)
by: Chen, Lingjiao, et al.
Published: (2025)
Optimizing LLM Queries in Relational Data Analytics Workloads
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
by: Li, Dacheng, et al.
Published: (2023)
by: Li, Dacheng, et al.
Published: (2023)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
by: Cao, Shiyi, et al.
Published: (2026)
by: Cao, Shiyi, et al.
Published: (2026)
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
The Time is Here for Just-in-Time Systems: Challenges and Opportunities
by: Liu, Shu, et al.
Published: (2026)
by: Liu, Shu, et al.
Published: (2026)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
RedunCut: Measurement-Driven Sampling and Accuracy Performance Modeling for Low-Cost Live Video Analytics
by: Sela, Gur-Eyal, et al.
Published: (2025)
by: Sela, Gur-Eyal, et al.
Published: (2025)
Let the Barbarians In: How AI Can Accelerate Systems Performance Research
by: Cheng, Audrey, et al.
Published: (2025)
by: Cheng, Audrey, et al.
Published: (2025)
AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
by: Cemri, Mert, et al.
Published: (2026)
by: Cemri, Mert, et al.
Published: (2026)
Specifications: The missing link to making the development of LLM systems an engineering discipline
by: Stoica, Ion, et al.
Published: (2024)
by: Stoica, Ion, et al.
Published: (2024)
World Model on Million-Length Video And Language With Blockwise RingAttention
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026)
by: Mishra, Mayank, et al.
Published: (2026)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Similar Items
-
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024) -
vCache: Verified Semantic Prompt Caching
by: Schroeder, Luis Gaspar, et al.
Published: (2025) -
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024) -
Inference Time Context Sparsity: Illusion or Opportunity?
by: Joshi, Sahil, et al.
Published: (2026) -
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)