Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
Fuente:
arXiv
Guardado en:
| Autores principales: | Tomczak, Nathaniel, Kuppannagari, Sanmukh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
por: Yang, Shang, et al.
Publicado: (2025)
por: Yang, Shang, et al.
Publicado: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
por: Yang, Hanmei, et al.
Publicado: (2024)
por: Yang, Hanmei, et al.
Publicado: (2024)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
por: Daghero, Francesco, et al.
Publicado: (2025)
por: Daghero, Francesco, et al.
Publicado: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
por: Ma, Bole, et al.
Publicado: (2026)
por: Ma, Bole, et al.
Publicado: (2026)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
por: Jung, Victor J. B., et al.
Publicado: (2024)
por: Jung, Victor J. B., et al.
Publicado: (2024)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
por: Sedlak, Boris, et al.
Publicado: (2025)
por: Sedlak, Boris, et al.
Publicado: (2025)
Democratizing AI: A Comparative Study in Deep Learning Efficiency and Future Trends in Computational Processing
por: Amin, Lisan Al, et al.
Publicado: (2026)
por: Amin, Lisan Al, et al.
Publicado: (2026)
Prompt-Aware Scheduling for Low-Latency LLM Serving
por: Tao, Yiheng, et al.
Publicado: (2025)
por: Tao, Yiheng, et al.
Publicado: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
por: Metere, Alfredo
Publicado: (2025)
por: Metere, Alfredo
Publicado: (2025)
Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
por: Qian, Wenzhuo, et al.
Publicado: (2025)
por: Qian, Wenzhuo, et al.
Publicado: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
por: Yu, Shan, et al.
Publicado: (2025)
por: Yu, Shan, et al.
Publicado: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
por: Abstreiter, Maximilian, et al.
Publicado: (2025)
por: Abstreiter, Maximilian, et al.
Publicado: (2025)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
por: Chu, Ruifan, et al.
Publicado: (2025)
por: Chu, Ruifan, et al.
Publicado: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
por: Li, Xiangchen, et al.
Publicado: (2025)
por: Li, Xiangchen, et al.
Publicado: (2025)
A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies
por: Korndörfer, Jonas H. Müller, et al.
Publicado: (2025)
por: Korndörfer, Jonas H. Müller, et al.
Publicado: (2025)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
por: Yuan, Ziqi, et al.
Publicado: (2025)
por: Yuan, Ziqi, et al.
Publicado: (2025)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
por: Huang, Zixiao, et al.
Publicado: (2025)
por: Huang, Zixiao, et al.
Publicado: (2025)
Compiler-First State Space Duality and Portable $O(1)$ Autoregressive Caching for Inference
por: Santoni, Cosmo
Publicado: (2026)
por: Santoni, Cosmo
Publicado: (2026)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
por: Miao, Xupeng, et al.
Publicado: (2023)
por: Miao, Xupeng, et al.
Publicado: (2023)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
por: Nichols, Daniel, et al.
Publicado: (2026)
por: Nichols, Daniel, et al.
Publicado: (2026)
Syno: Structured Synthesis for Neural Operators
por: Zhuo, Yongqi, et al.
Publicado: (2024)
por: Zhuo, Yongqi, et al.
Publicado: (2024)
TrainMover: An Interruption-Resilient Runtime for ML Training
por: Lao, ChonLam, et al.
Publicado: (2024)
por: Lao, ChonLam, et al.
Publicado: (2024)
FastPersist: Accelerating Model Checkpointing in Deep Learning
por: Wang, Guanhua, et al.
Publicado: (2024)
por: Wang, Guanhua, et al.
Publicado: (2024)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
por: Imani, HamidReza, et al.
Publicado: (2024)
por: Imani, HamidReza, et al.
Publicado: (2024)
Training Time Prediction for Mixed Precision-based Distributed Training
por: Kang, Minchul, et al.
Publicado: (2026)
por: Kang, Minchul, et al.
Publicado: (2026)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
por: Ullah, Irfan, et al.
Publicado: (2026)
por: Ullah, Irfan, et al.
Publicado: (2026)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
por: Singh, Siddharth, et al.
Publicado: (2023)
por: Singh, Siddharth, et al.
Publicado: (2023)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
por: Anik, Md Saidul Hoque, et al.
Publicado: (2024)
por: Anik, Md Saidul Hoque, et al.
Publicado: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
por: Luo, Jiawei, et al.
Publicado: (2026)
por: Luo, Jiawei, et al.
Publicado: (2026)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
por: Bhattacharjee, Arijit, et al.
Publicado: (2025)
por: Bhattacharjee, Arijit, et al.
Publicado: (2025)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
por: Maurya, Avinash, et al.
Publicado: (2024)
por: Maurya, Avinash, et al.
Publicado: (2024)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
por: Helal, Ahmed E., et al.
Publicado: (2025)
por: Helal, Ahmed E., et al.
Publicado: (2025)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
por: Panigrahy, Deepak, et al.
Publicado: (2026)
por: Panigrahy, Deepak, et al.
Publicado: (2026)
Efficient Chromosome Parallelization for Precision Medicine Genomic Workflows
por: Montserrat, Daniel Mas, et al.
Publicado: (2025)
por: Montserrat, Daniel Mas, et al.
Publicado: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
por: Deshmukh, Dhruv, et al.
Publicado: (2025)
por: Deshmukh, Dhruv, et al.
Publicado: (2025)
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
por: Shahbazinia, Amirhossein, et al.
Publicado: (2025)
por: Shahbazinia, Amirhossein, et al.
Publicado: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
por: Titopoulos, Vasileios, et al.
Publicado: (2025)
por: Titopoulos, Vasileios, et al.
Publicado: (2025)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
por: Ding, Shiwei, et al.
Publicado: (2025)
por: Ding, Shiwei, et al.
Publicado: (2025)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
por: Gupta, Ahan, et al.
Publicado: (2026)
por: Gupta, Ahan, et al.
Publicado: (2026)
Ejemplares similares
-
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
por: Yang, Shang, et al.
Publicado: (2025) -
ProTrain: Efficient LLM Training via Memory-Aware Techniques
por: Yang, Hanmei, et al.
Publicado: (2024) -
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
por: Daghero, Francesco, et al.
Publicado: (2025) -
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
por: Ma, Bole, et al.
Publicado: (2026) -
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
por: Jung, Victor J. B., et al.
Publicado: (2024)