Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Le, Hoang Anh Duy, Joshi, Sahil, Yang, Zeyu, Xu, Zhaozhuo, Shrivastava, Anshumali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
di: Joshi, Sahil, et al.
Pubblicazione: (2026)
di: Joshi, Sahil, et al.
Pubblicazione: (2026)
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
di: Yang, Zeyu, et al.
Pubblicazione: (2026)
di: Yang, Zeyu, et al.
Pubblicazione: (2026)
Inference Time Context Sparsity: Illusion or Opportunity?
di: Joshi, Sahil, et al.
Pubblicazione: (2026)
di: Joshi, Sahil, et al.
Pubblicazione: (2026)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
di: Monteiro, João, et al.
Pubblicazione: (2024)
di: Monteiro, João, et al.
Pubblicazione: (2024)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
SnapKV: LLM Knows What You are Looking for Before Generation
di: Li, Yuhong, et al.
Pubblicazione: (2024)
di: Li, Yuhong, et al.
Pubblicazione: (2024)
Look Before You Leap: Autonomous Exploration for LLM Agents
di: Ye, Ziang, et al.
Pubblicazione: (2026)
di: Ye, Ziang, et al.
Pubblicazione: (2026)
LLM Multi-Agent Systems: Challenges and Open Problems
di: Han, Shanshan, et al.
Pubblicazione: (2024)
di: Han, Shanshan, et al.
Pubblicazione: (2024)
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
Plan Before You Trade: Inference-Time Optimization for RL Trading Agents
di: Go, Eun, et al.
Pubblicazione: (2026)
di: Go, Eun, et al.
Pubblicazione: (2026)
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
di: Stripelis, Dimitris, et al.
Pubblicazione: (2024)
di: Stripelis, Dimitris, et al.
Pubblicazione: (2024)
Summarize Before You Speak with ARACH: A Training-Free Inference-Time Plug-In for Enhancing LLMs via Global Attention Reallocation
di: Wang, Jingtao, et al.
Pubblicazione: (2026)
di: Wang, Jingtao, et al.
Pubblicazione: (2026)
Scout-Assisted Planning for Heterogeneous Robot Teams under Partially Known Environments
di: Bui, Hoang-Dung, et al.
Pubblicazione: (2026)
di: Bui, Hoang-Dung, et al.
Pubblicazione: (2026)
Test Before You Deploy: Governing Updates in the LLM Supply Chain
di: Chishti, Mohd Sameen, et al.
Pubblicazione: (2026)
di: Chishti, Mohd Sameen, et al.
Pubblicazione: (2026)
REFRAG: Rethinking RAG based Decoding
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms
di: Deb, Rohan, et al.
Pubblicazione: (2025)
di: Deb, Rohan, et al.
Pubblicazione: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
di: Spiegel, Michal, et al.
Pubblicazione: (2025)
di: Spiegel, Michal, et al.
Pubblicazione: (2025)
Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
di: Aytes, Simon A., et al.
Pubblicazione: (2025)
di: Aytes, Simon A., et al.
Pubblicazione: (2025)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
di: Yuan, Wenhao, et al.
Pubblicazione: (2026)
di: Yuan, Wenhao, et al.
Pubblicazione: (2026)
Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design
di: Yang, Zaifei, et al.
Pubblicazione: (2026)
di: Yang, Zaifei, et al.
Pubblicazione: (2026)
Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
di: Yao, Feiyu, et al.
Pubblicazione: (2026)
di: Yao, Feiyu, et al.
Pubblicazione: (2026)
BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
HULLMI: Human vs LLM identification with explainability
di: Joshi, Prathamesh Dinesh, et al.
Pubblicazione: (2024)
di: Joshi, Prathamesh Dinesh, et al.
Pubblicazione: (2024)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
di: Han, Feijiang, et al.
Pubblicazione: (2025)
di: Han, Feijiang, et al.
Pubblicazione: (2025)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
di: Lin, Wei-Hsiang, et al.
Pubblicazione: (2025)
di: Lin, Wei-Hsiang, et al.
Pubblicazione: (2025)
Sensitivity Meets Sparsity: The Impact of Extremely Sparse Parameter Patterns on Theory-of-Mind of Large Language Models
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
Causal-Aware Generative Adversarial Networks with Reinforcement Learning
di: Nguyen, Tu Anh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Tu Anh Hoang, et al.
Pubblicazione: (2025)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
di: Long, Jikai, et al.
Pubblicazione: (2025)
di: Long, Jikai, et al.
Pubblicazione: (2025)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
di: He, Jiashu, et al.
Pubblicazione: (2026)
di: He, Jiashu, et al.
Pubblicazione: (2026)
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
di: Song, Chuxu, et al.
Pubblicazione: (2026)
di: Song, Chuxu, et al.
Pubblicazione: (2026)
UltraSketchLLM: Saliency-Driven Sketching for Ultra-Low Bit LLM Compression
di: Zou, Sunan, et al.
Pubblicazione: (2025)
di: Zou, Sunan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
di: Zhang, Tianyi, et al.
Pubblicazione: (2024) -
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
di: Yang, Zeyu, et al.
Pubblicazione: (2025) -
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
di: Joshi, Sahil, et al.
Pubblicazione: (2025) -
Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
di: Zhang, Tianyi, et al.
Pubblicazione: (2024) -
SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
di: Joshi, Sahil, et al.
Pubblicazione: (2026)