Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Rodionov, Gleb, Garipov, Roman, Shutova, Alina, Yakushev, George, Schultheis, Erik, Egiazarian, Vage, Sinitsin, Anton, Kuznedelev, Denis, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
by: Shutova, Alina, et al.
Published: (2025)
by: Shutova, Alina, et al.
Published: (2025)
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026)
by: Egiazarian, Vage, et al.
Published: (2026)
Evaluating Memory Structure in LLM Agents
by: Shutova, Alina, et al.
Published: (2026)
by: Shutova, Alina, et al.
Published: (2026)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
by: Yakushev, George, et al.
Published: (2025)
by: Yakushev, George, et al.
Published: (2025)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)
by: Chen, Jiale, et al.
Published: (2025)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
by: Schultheis, Erik, et al.
Published: (2025)
by: Schultheis, Erik, et al.
Published: (2025)
AutoJudge: Judge Decoding Without Manual Annotation
by: Garipov, Roman, et al.
Published: (2025)
by: Garipov, Roman, et al.
Published: (2025)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
by: Egiazarian, Vage, et al.
Published: (2025)
by: Egiazarian, Vage, et al.
Published: (2025)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Reasoning Shift: How Context Silently Shortens LLM Reasoning
by: Rodionov, Gleb
Published: (2026)
by: Rodionov, Gleb
Published: (2026)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)
by: Wu, Diyuan, et al.
Published: (2024)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
by: Yakushev, George, et al.
Published: (2025)
by: Yakushev, George, et al.
Published: (2025)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Graph Neural Networks Gone Hogwild
by: Solodova, Olga, et al.
Published: (2024)
by: Solodova, Olga, et al.
Published: (2024)
Neural Optimal Transport with General Cost Functionals
by: Asadulaev, Arip, et al.
Published: (2022)
by: Asadulaev, Arip, et al.
Published: (2022)
Discrete Neural Algorithmic Reasoning
by: Rodionov, Gleb, et al.
Published: (2024)
by: Rodionov, Gleb, et al.
Published: (2024)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
by: Frantar, Elias, et al.
Published: (2024)
by: Frantar, Elias, et al.
Published: (2024)
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
by: Voronov, Anton, et al.
Published: (2024)
by: Voronov, Anton, et al.
Published: (2024)
Rethinking Optimal Transport in Offline Reinforcement Learning
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
by: Leidinger, Alina, et al.
Published: (2024)
by: Leidinger, Alina, et al.
Published: (2024)
Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
by: Iofinova, Eugenia, et al.
Published: (2025)
by: Iofinova, Eugenia, et al.
Published: (2025)
Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks
by: Pankratov, Sergey, et al.
Published: (2025)
by: Pankratov, Sergey, et al.
Published: (2025)
Simple Opinion Dynamics for No-Regret Learning
by: Lazarsfeld, John, et al.
Published: (2023)
by: Lazarsfeld, John, et al.
Published: (2023)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
by: Iofinova, Eugenia, et al.
Published: (2026)
by: Iofinova, Eugenia, et al.
Published: (2026)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Catching the context: the first genre-determined reading of Francysk Skaryna’s portrait
by: Olga Shutova
Published: (2022)
by: Olga Shutova
Published: (2022)
How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
by: Choenni, Rochelle, et al.
Published: (2023)
by: Choenni, Rochelle, et al.
Published: (2023)
Does Diffusion Beat GAN in Image Super Resolution?
by: Kuznedelev, Denis, et al.
Published: (2024)
by: Kuznedelev, Denis, et al.
Published: (2024)
Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and Beyond
by: Platonov, Oleg, et al.
Published: (2022)
by: Platonov, Oleg, et al.
Published: (2022)
Generating artificial digital image correlation data using physics-guided adversarial networks
by: Melching, David, et al.
Published: (2023)
by: Melching, David, et al.
Published: (2023)
Key Algorithms for Keyphrase Generation: Instruction-Based LLMs for Russian Scientific Keyphrases
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
LitmusKt: Concurrency Stress Testing for Kotlin
by: Lochmelis, Denis, et al.
Published: (2025)
by: Lochmelis, Denis, et al.
Published: (2025)
Similar Items
-
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
by: Shutova, Alina, et al.
Published: (2025) -
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024) -
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026) -
Evaluating Memory Structure in LLM Agents
by: Shutova, Alina, et al.
Published: (2026) -
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)