Loki: Low-rank Keys for Efficient Sparse Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Singhania, Prajwal, Singh, Siddharth, He, Shwai, Feizi, Soheil, Bhatele, Abhinav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Power Law Guided Dynamic Sifting for Efficient Attention
di: Koley, Nirav, et al.
Pubblicazione: (2025)
di: Koley, Nirav, et al.
Pubblicazione: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
di: Singh, Siddharth, et al.
Pubblicazione: (2023)
Speculating Experts Accelerates Inference for Mixture-of-Experts
di: Madan, Vivan, et al.
Pubblicazione: (2026)
di: Madan, Vivan, et al.
Pubblicazione: (2026)
Understanding and Improving Communication Performance in Multi-node LLM Inference
di: Singhania, Prajwal, et al.
Pubblicazione: (2025)
di: Singhania, Prajwal, et al.
Pubblicazione: (2025)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
di: Ranjan, Aditya K., et al.
Pubblicazione: (2025)
di: Ranjan, Aditya K., et al.
Pubblicazione: (2025)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
di: Singh, Siddharth, et al.
Pubblicazione: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
di: Chegini, Atoosa, et al.
Pubblicazione: (2026)
di: Chegini, Atoosa, et al.
Pubblicazione: (2026)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
di: Costello, Ian J., et al.
Pubblicazione: (2020)
di: Costello, Ian J., et al.
Pubblicazione: (2020)
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
di: Saha, Shoumik, et al.
Pubblicazione: (2025)
How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric Encodings
di: Audia, Samuel, et al.
Pubblicazione: (2025)
di: Audia, Samuel, et al.
Pubblicazione: (2025)
What Matters in Transformers? Not All Attention is Needed
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
di: Wang, Wenxiao, et al.
Pubblicazione: (2025)
di: Wang, Wenxiao, et al.
Pubblicazione: (2025)
Revisiting the Past: Data Unlearning with Model State History
di: Rezaei, Keivan, et al.
Pubblicazione: (2025)
di: Rezaei, Keivan, et al.
Pubblicazione: (2025)
Fourier Low-rank and Sparse Tensor for Efficient Tensor Completion
di: Li, Jingyang, et al.
Pubblicazione: (2025)
di: Li, Jingyang, et al.
Pubblicazione: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
di: McLeish, Sean, et al.
Pubblicazione: (2025)
di: McLeish, Sean, et al.
Pubblicazione: (2025)
Maestro: Joint Graph & Config Optimization for Reliable AI Agents
di: Wang, Wenxiao, et al.
Pubblicazione: (2025)
di: Wang, Wenxiao, et al.
Pubblicazione: (2025)
RePanda: Pandas-powered Tabular Verification and Reasoning
di: Chegini, Atoosa Malemir, et al.
Pubblicazione: (2025)
di: Chegini, Atoosa Malemir, et al.
Pubblicazione: (2025)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2024)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2024)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
di: Geiping, Jonas, et al.
Pubblicazione: (2025)
di: Geiping, Jonas, et al.
Pubblicazione: (2025)
Low-Rank Key Value Attention
di: O'Neill, James, et al.
Pubblicazione: (2026)
di: O'Neill, James, et al.
Pubblicazione: (2026)
FLARE: Fast Low-rank Attention Routing Engine
di: Puri, Vedant, et al.
Pubblicazione: (2025)
di: Puri, Vedant, et al.
Pubblicazione: (2025)
Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
di: Cai, Weilin, et al.
Pubblicazione: (2025)
di: Cai, Weilin, et al.
Pubblicazione: (2025)
QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
di: Wang, Yutong, et al.
Pubblicazione: (2025)
di: Wang, Yutong, et al.
Pubblicazione: (2025)
Information Consistent Pruning: How to Efficiently Search for Sparse Networks?
di: Gharatappeh, Soheil, et al.
Pubblicazione: (2025)
di: Gharatappeh, Soheil, et al.
Pubblicazione: (2025)
LOST: Low-rank and Sparse Pre-training for Large Language Models
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
di: Koirala, Prajwal, et al.
Pubblicazione: (2025)
di: Koirala, Prajwal, et al.
Pubblicazione: (2025)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
di: He, Zhengfu, et al.
Pubblicazione: (2025)
di: He, Zhengfu, et al.
Pubblicazione: (2025)
What do we learn from inverting CLIP models?
di: Kazemi, Hamid, et al.
Pubblicazione: (2024)
di: Kazemi, Hamid, et al.
Pubblicazione: (2024)
Early Stopping for Large Reasoning Models via Confidence Dynamics
di: Hosseini, Parsa, et al.
Pubblicazione: (2026)
di: Hosseini, Parsa, et al.
Pubblicazione: (2026)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
MoEless: Efficient MoE LLM Serving via Serverless Computing
di: Yu, Hanfei, et al.
Pubblicazione: (2026)
di: Yu, Hanfei, et al.
Pubblicazione: (2026)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
di: Shahbazi, Ashkan, et al.
Pubblicazione: (2025)
di: Shahbazi, Ashkan, et al.
Pubblicazione: (2025)
Low-rank Momentum Factorization for Memory Efficient Training
di: Mahdavinia, Pouria, et al.
Pubblicazione: (2025)
di: Mahdavinia, Pouria, et al.
Pubblicazione: (2025)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
di: Ji, Xiaodong, et al.
Pubblicazione: (2025)
di: Ji, Xiaodong, et al.
Pubblicazione: (2025)
CRAUM-Net: Contextual Recursive Attention with Uncertainty Modeling for Salient Object Detection
di: Sagar, Abhinav
Pubblicazione: (2020)
di: Sagar, Abhinav
Pubblicazione: (2020)
Documenti analoghi
-
Power Law Guided Dynamic Sifting for Efficient Attention
di: Koley, Nirav, et al.
Pubblicazione: (2025) -
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
di: Singh, Siddharth, et al.
Pubblicazione: (2023) -
Speculating Experts Accelerates Inference for Mixture-of-Experts
di: Madan, Vivan, et al.
Pubblicazione: (2026) -
Understanding and Improving Communication Performance in Multi-node LLM Inference
di: Singhania, Prajwal, et al.
Pubblicazione: (2025) -
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
di: Chaturvedi, Aman, et al.
Pubblicazione: (2024)