SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jonathan, Farahini, Nasim, Iuliugin, Evgenii, Vesterlund, Magnus, Häggström, Christian, Wang, Guangtao, Upasani, Shubhangi, Sachdeva, Ayush, Li, Rui, Fu, Faline, Wu, Chen, Siddiqua, Ayesha, Long, John, Zhao, Tuowen, Musaddiq, Matheen, Zeffer, Håkan, Du, Yun, Wang, Mingran, Li, Qinghua, Li, Bo, Thakker, Urmish, Prabhakar, Raghu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Limits of Long-Context Reasoning in Automated Bug Fixing
von: Raju, Ravi, et al.
Veröffentlicht: (2026)
von: Raju, Ravi, et al.
Veröffentlicht: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
von: Koeplinger, David, et al.
Veröffentlicht: (2024)
von: Koeplinger, David, et al.
Veröffentlicht: (2024)
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
von: Prabhakar, Raghu, et al.
Veröffentlicht: (2024)
von: Prabhakar, Raghu, et al.
Veröffentlicht: (2024)
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
von: Raju, Ravi, et al.
Veröffentlicht: (2024)
von: Raju, Ravi, et al.
Veröffentlicht: (2024)
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
von: Zhang, Qizheng, et al.
Veröffentlicht: (2025)
von: Zhang, Qizheng, et al.
Veröffentlicht: (2025)
Proximity to seed sites_Proposal for measurement_DavidVesterlund_10.5281/zenodo.17401322
von: Vesterlund, David
Veröffentlicht: (2025)
von: Vesterlund, David
Veröffentlicht: (2025)
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
Synthetic Document Question Answering in Hungarian
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
Economic Impacts of Salmonella Dublin in Dairy Farms: Panel Evidence From Denmark
von: Dagim Belay, et al.
Veröffentlicht: (2025)
von: Dagim Belay, et al.
Veröffentlicht: (2025)
SambaLingo: Teaching Large Language Models New Languages
von: Csaki, Zoltan, et al.
Veröffentlicht: (2024)
von: Csaki, Zoltan, et al.
Veröffentlicht: (2024)
Original Antigenic Sin in CD4+ T Cells
von: Mingran Zhang, et al.
Veröffentlicht: (2025)
von: Mingran Zhang, et al.
Veröffentlicht: (2025)
Unified Low-Light Traffic Image Enhancement via Multi-Stage Illumination Recovery and Adaptive Noise Suppression
von: Namrah, Siddiqua
Veröffentlicht: (2025)
von: Namrah, Siddiqua
Veröffentlicht: (2025)
Concurrent Multiphysics and Multiscale Topology Optimization for Lightweight Laser-Driven Porous Actuator Systems
von: Ali, Musaddiq Al, et al.
Veröffentlicht: (2024)
von: Ali, Musaddiq Al, et al.
Veröffentlicht: (2024)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)
Surveillance of high‐yield processes using deep learning models
von: Musaddiq Ibrahim, et al.
Veröffentlicht: (2024)
von: Musaddiq Ibrahim, et al.
Veröffentlicht: (2024)
Snapping Actuators with Asymmetric and Sequenced Motion
von: Li, Xin, et al.
Veröffentlicht: (2026)
von: Li, Xin, et al.
Veröffentlicht: (2026)
Scaling Inter-procedural Dataflow Analysis on the Cloud
von: Sun, Zewen, et al.
Veröffentlicht: (2024)
von: Sun, Zewen, et al.
Veröffentlicht: (2024)
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
von: Li, Richie
Veröffentlicht: (2025)
von: Li, Richie
Veröffentlicht: (2025)
TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
Does corporate digitalisation moderate real earnings management?
von: Zhukun Lou, et al.
Veröffentlicht: (2024)
von: Zhukun Lou, et al.
Veröffentlicht: (2024)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
von: Li, Fang
Veröffentlicht: (2025)
von: Li, Fang
Veröffentlicht: (2025)
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
General-purpose Dataflow Model with Neuromorphic Primitives
von: Zhang, Weihao, et al.
Veröffentlicht: (2024)
von: Zhang, Weihao, et al.
Veröffentlicht: (2024)
The Role of Libraries in Lifelong Learning. Final Report of the IFLA Project under the Section for Public Libraries
von: Haggstrom, Britt Marie, Ed.
Veröffentlicht: (2004)
von: Haggstrom, Britt Marie, Ed.
Veröffentlicht: (2004)
Rotation‐Based Snap‐Fit Mechanical Metamaterials
von: Rui Xu, et al.
Veröffentlicht: (2025)
von: Rui Xu, et al.
Veröffentlicht: (2025)
A Trust-Based Malicious RSU Detection Mechanism in Edge-Enabled Vehicular Ad Hoc Networks
von: Siddiqua, Farhana, et al.
Veröffentlicht: (2022)
von: Siddiqua, Farhana, et al.
Veröffentlicht: (2022)
A High Accuracy Symplectic Scheme for Advection Diffusion Reaction Models in Bioseparation
von: Siddiqua, Farjana, et al.
Veröffentlicht: (2025)
von: Siddiqua, Farjana, et al.
Veröffentlicht: (2025)
Statistics in a Backscatter Eddy Viscosity Turbulence Model
von: Pakzad, Ali, et al.
Veröffentlicht: (2024)
von: Pakzad, Ali, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Conventional Host‐Guest Inclusion Complex & Sustainable Resveratrol Loaded Hydroxypropyl‐β‐Cyclodextrin Nanosponges With Improved Therapeutic Potential
von: Ayesha Siddiqua, et al.
Veröffentlicht: (2026)
von: Ayesha Siddiqua, et al.
Veröffentlicht: (2026)
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
von: Zhang, Xingyi, et al.
Veröffentlicht: (2026)
von: Zhang, Xingyi, et al.
Veröffentlicht: (2026)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
Software Defined Wireless Network: The Rise, The Evolution, The Future
von: Shubhangi Saha
Veröffentlicht: (2023)
von: Shubhangi Saha
Veröffentlicht: (2023)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
von: Gu, Hang, et al.
Veröffentlicht: (2026)
von: Gu, Hang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Limits of Long-Context Reasoning in Automated Bug Fixing
von: Raju, Ravi, et al.
Veröffentlicht: (2026) -
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025) -
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026) -
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026) -
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
von: Koeplinger, David, et al.
Veröffentlicht: (2024)