MagicPIG: LSH Sampling for Efficient LLM Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhuoming, Sadhukhan, Ranajoy, Ye, Zihao, Zhou, Yang, Zhang, Jianyu, Nolte, Niklas, Tian, Yuandong, Douze, Matthijs, Bottou, Leon, Jia, Zhihao, Chen, Beidi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory Mosaics
by: Zhang, Jianyu, et al.
Published: (2024)
by: Zhang, Jianyu, et al.
Published: (2024)
Kinetics: Rethinking Test-Time Scaling Laws
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
by: Sadhukhan, Ranajoy, et al.
Published: (2024)
by: Sadhukhan, Ranajoy, et al.
Published: (2024)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
STEM: Scaling Transformers with Embedding Modules
by: Sadhukhan, Ranajoy, et al.
Published: (2026)
by: Sadhukhan, Ranajoy, et al.
Published: (2026)
Memory Mosaics at scale
by: Zhang, Jianyu, et al.
Published: (2025)
by: Zhang, Jianyu, et al.
Published: (2025)
Fine-tuning with Very Large Dropout
by: Zhang, Jianyu, et al.
Published: (2024)
by: Zhang, Jianyu, et al.
Published: (2024)
Machine learning and high dimensional vector search
by: Douze, Matthijs
Published: (2025)
by: Douze, Matthijs
Published: (2025)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
by: Svirschevski, Ruslan, et al.
Published: (2024)
by: Svirschevski, Ruslan, et al.
Published: (2024)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
by: Wang, Huanyu, et al.
Published: (2026)
by: Wang, Huanyu, et al.
Published: (2026)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
These Are Not All the Features You Are Looking For: A Fundamental Bottleneck in Supervised Pretraining
by: Yang, Xingyu Alice, et al.
Published: (2025)
by: Yang, Xingyu Alice, et al.
Published: (2025)
Jackpot: Optimal Budgeted Rejection Sampling for Extreme Actor-Policy Mismatch Reinforcement Learning
by: Chen, Zhuoming, et al.
Published: (2026)
by: Chen, Zhuoming, et al.
Published: (2026)
Efficient Streaming Language Models with Attention Sinks
by: Xiao, Guangxuan, et al.
Published: (2023)
by: Xiao, Guangxuan, et al.
Published: (2023)
A Single Character can Make or Break Your LLM Evals
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Sirius: Contextual Sparsity with Correction for Efficient LLMs
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
by: Chen, Zhuoming, et al.
Published: (2024)
by: Chen, Zhuoming, et al.
Published: (2024)
LoCoCo: Dropping In Convolutions for Long Context Compression
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
by: Cheng, Xinle, et al.
Published: (2025)
by: Cheng, Xinle, et al.
Published: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023)
by: Tian, Yuandong, et al.
Published: (2023)
On the Surprising Effectiveness of Attention Transfer for Vision Transformers
by: Li, Alexander C., et al.
Published: (2024)
by: Li, Alexander C., et al.
Published: (2024)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
by: Luo, Cheng, et al.
Published: (2024)
by: Luo, Cheng, et al.
Published: (2024)
Vector search with small radiuses
by: Szilvasy, Gergely, et al.
Published: (2024)
by: Szilvasy, Gergely, et al.
Published: (2024)
Functional Invariants to Watermark Large Transformers
by: Fernandez, Pierre, et al.
Published: (2023)
by: Fernandez, Pierre, et al.
Published: (2023)
Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
MonarchRT: Efficient Attention for Real-Time Video Generation
by: Agarwal, Krish, et al.
Published: (2026)
by: Agarwal, Krish, et al.
Published: (2026)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
Data-Dependent LSH for the Earth Mover's Distance
by: Jayaram, Rajesh, et al.
Published: (2024)
by: Jayaram, Rajesh, et al.
Published: (2024)
Watermark Anything with Localized Messages
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
Watermarking Makes Language Models Radioactive
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
Lossless Compression of Vector IDs for Approximate Nearest Neighbor Search
by: Severo, Daniel, et al.
Published: (2025)
by: Severo, Daniel, et al.
Published: (2025)
On the LSH Distortion of Ulam and Cayley Similarities
by: Chierichetti, Flavio, et al.
Published: (2026)
by: Chierichetti, Flavio, et al.
Published: (2026)
LSH-DynED: A Dynamic Ensemble Framework with LSH-Based Undersampling for Evolving Multi-Class Imbalanced Classification
by: Abadifard, Soheil, et al.
Published: (2025)
by: Abadifard, Soheil, et al.
Published: (2025)
Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
by: Tian, Yuandong
Published: (2024)
by: Tian, Yuandong
Published: (2024)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
by: Tian, Yuandong
Published: (2025)
by: Tian, Yuandong
Published: (2025)
Implementing and Evaluating E2LSH on Storage
by: Nakanishi, Yu, et al.
Published: (2024)
by: Nakanishi, Yu, et al.
Published: (2024)
Improving LSH via Tensorized Random Projection
by: Verma, Bhisham Dev, et al.
Published: (2024)
by: Verma, Bhisham Dev, et al.
Published: (2024)
BERT-LSH: Reducing Absolute Compute For Attention
by: Li, Zezheng, et al.
Published: (2024)
by: Li, Zezheng, et al.
Published: (2024)
Similar Items
-
Memory Mosaics
by: Zhang, Jianyu, et al.
Published: (2024) -
Kinetics: Rethinking Test-Time Scaling Laws
by: Sadhukhan, Ranajoy, et al.
Published: (2025) -
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
by: Sadhukhan, Ranajoy, et al.
Published: (2024) -
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
by: Zhou, Yang, et al.
Published: (2025) -
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
by: Sun, Hanshi, et al.
Published: (2024)