StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
Fuente:
arXiv
Saved in:
| Main Authors: | Jaber, Jaber, Jaber, Osama |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
HCLSM: Hierarchical Causal Latent State Machines for Object-Centric World Modeling
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
by: Li, Haorui, et al.
Published: (2026)
by: Li, Haorui, et al.
Published: (2026)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
by: Çöplü, Tolga, et al.
Published: (2023)
by: Çöplü, Tolga, et al.
Published: (2023)
STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts
by: Bamberger, Zachary, et al.
Published: (2026)
by: Bamberger, Zachary, et al.
Published: (2026)
A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design
by: Leoveanu-Condrei, Claudiu
Published: (2025)
by: Leoveanu-Condrei, Claudiu
Published: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
by: Topcu, Burak, et al.
Published: (2026)
by: Topcu, Burak, et al.
Published: (2026)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
The Hidden Attention of Mamba Models
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Dual-Phase Federated Deep Unlearning via Weight-Aware Rollback and Reconstruction
by: Zhou, Changjun, et al.
Published: (2025)
by: Zhou, Changjun, et al.
Published: (2025)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)
by: Atad, Ido Andrew, et al.
Published: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
by: Katraouras, Thomas, et al.
Published: (2026)
by: Katraouras, Thomas, et al.
Published: (2026)
Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
by: Kashyap, Ankit
Published: (2025)
by: Kashyap, Ankit
Published: (2025)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
by: Zhao, Hangyue, et al.
Published: (2026)
by: Zhao, Hangyue, et al.
Published: (2026)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
by: Forchheimer, Robert
Published: (2026)
by: Forchheimer, Robert
Published: (2026)
Memory-Efficient Differentially Private Training with Gradient Random Projection
by: Mulrooney, Alex, et al.
Published: (2025)
by: Mulrooney, Alex, et al.
Published: (2025)
Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data
by: de Campos, Andresa Rodrigues, et al.
Published: (2026)
by: de Campos, Andresa Rodrigues, et al.
Published: (2026)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
by: Kamath, Aditya K, et al.
Published: (2026)
by: Kamath, Aditya K, et al.
Published: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing
by: Huang, Yinghui, et al.
Published: (2025)
by: Huang, Yinghui, et al.
Published: (2025)
TinyML NLP Scheme for Semantic Wireless Sentiment Classification with Privacy Preservation
by: Radwan, Ahmed Y., et al.
Published: (2024)
by: Radwan, Ahmed Y., et al.
Published: (2024)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
by: Karabulut, Burak, et al.
Published: (2026)
by: Karabulut, Burak, et al.
Published: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
by: Zimerman, Itamar, et al.
Published: (2024)
by: Zimerman, Itamar, et al.
Published: (2024)
AdaptiveCoPilot: Design and Testing of a NeuroAdaptive LLM Cockpit Guidance System in both Novice and Expert Pilots
by: Wen, Shaoyue, et al.
Published: (2025)
by: Wen, Shaoyue, et al.
Published: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025)
by: Amini, Ali
Published: (2025)
Representation Fidelity:Auditing Algorithmic Decisions About Humans Using Self-Descriptions
by: Elstner, Theresa, et al.
Published: (2026)
by: Elstner, Theresa, et al.
Published: (2026)
Similar Items
-
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
by: Jaber, Jaber, et al.
Published: (2026) -
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
by: Jaber, Jaber, et al.
Published: (2026) -
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
by: Shakerdargah, Mohammadali, et al.
Published: (2024) -
HCLSM: Hierarchical Causal Latent State Machines for Object-Centric World Modeling
by: Jaber, Jaber, et al.
Published: (2026) -
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)