Saved in:
Bibliographic Details
Main Authors: Song, Zhao, Xie, Shenghao, Zhou, Samson
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.03678
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • This paper studies the computational challenges of large-scale attention-based models in artificial intelligence by utilizing importance sampling methods in the streaming setting. Inspired by the classical definition of the $\ell_2$ sampler and the recent progress of the attention scheme in Large Language Models (LLMs), we propose the definition of the attention sampler. Our approach significantly reduces the computational burden of traditional attention mechanisms. We analyze the effectiveness of the attention sampler from a theoretical perspective, including space and update time. Additionally, our framework exhibits scalability and broad applicability across various model architectures and domains.