Approximate Top-$k$ for Increased Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Key, Oscar, Ribar, Luka, Cattaneo, Alberto, Hudlass-Galley, Luke, Orr, Douglas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SparQ Attention: Bandwidth-Efficient LLM Inference
by: Ribar, Luka, et al.
Published: (2023)
by: Ribar, Luka, et al.
Published: (2023)
Optimal Formats for Weight Quantisation
by: Orr, Douglas, et al.
Published: (2025)
by: Orr, Douglas, et al.
Published: (2025)
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
by: L, Yashas Samaga B, et al.
Published: (2024)
by: L, Yashas Samaga B, et al.
Published: (2024)
Top-$k$ Feature Importance Ranking
by: Chen, Yuxi, et al.
Published: (2025)
by: Chen, Yuxi, et al.
Published: (2025)
Top-$k$ Classification and Cardinality-Aware Prediction
by: Mao, Anqi, et al.
Published: (2024)
by: Mao, Anqi, et al.
Published: (2024)
Statistical Models of Top-$k$ Partial Orders
by: Awadelkarim, Amel, et al.
Published: (2024)
by: Awadelkarim, Amel, et al.
Published: (2024)
Cardinality-Aware Set Prediction and Top-$k$ Classification
by: Cortes, Corinna, et al.
Published: (2024)
by: Cortes, Corinna, et al.
Published: (2024)
On the Minimax Regret in Online Ranking with Top-k Feedback
by: Zhang, Mingyuan, et al.
Published: (2023)
by: Zhang, Mingyuan, et al.
Published: (2023)
Differentiable Knapsack and Top-k Operators via Dynamic Programming
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
Why Ask One When You Can Ask $k$? Learning-to-Defer to the Top-$k$ Experts
by: Montreuil, Yannis, et al.
Published: (2025)
by: Montreuil, Yannis, et al.
Published: (2025)
Parallel Layer Normalization for Universal Approximation
by: Ni, Yunhao, et al.
Published: (2025)
by: Ni, Yunhao, et al.
Published: (2025)
Tab-Shapley: Identifying Top-k Tabular Data Quality Insights
by: Padala, Manisha, et al.
Published: (2025)
by: Padala, Manisha, et al.
Published: (2025)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
by: Maskey, Sohir, et al.
Published: (2026)
by: Maskey, Sohir, et al.
Published: (2026)
Foundations of Top-$k$ Decoding For Language Models
by: Noarov, Georgy, et al.
Published: (2025)
by: Noarov, Georgy, et al.
Published: (2025)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Generalized Top-k Mallows Model for Ranked Choices
by: Haddadan, Shahrzad, et al.
Published: (2025)
by: Haddadan, Shahrzad, et al.
Published: (2025)
One-Stage Top-$k$ Learning-to-Defer: Score-Based Surrogates with Theoretical Guarantees
by: Montreuil, Yannis, et al.
Published: (2025)
by: Montreuil, Yannis, et al.
Published: (2025)
TeamFormer: Shallow Parallel Transformers with Progressive Approximation
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Personalized Top-k Set Queries Over Predicted Scores
by: Nia, Sohrab Namazi, et al.
Published: (2025)
by: Nia, Sohrab Namazi, et al.
Published: (2025)
ZETA: Leveraging Z-order Curves for Efficient Top-k Attention
by: Zeng, Qiuhao, et al.
Published: (2025)
by: Zeng, Qiuhao, et al.
Published: (2025)
Optimizing Partial Area Under the Top-k Curve: Theory and Practice
by: Wang, Zitai, et al.
Published: (2022)
by: Wang, Zitai, et al.
Published: (2022)
Antithetic Sampling for Top-k Shapley Identification
by: Kolpaczki, Patrick, et al.
Published: (2025)
by: Kolpaczki, Patrick, et al.
Published: (2025)
Approximate Algorithms For $k$-Sparse Wasserstein Barycenter With Outliers
by: Yang, Qingyuan, et al.
Published: (2024)
by: Yang, Qingyuan, et al.
Published: (2024)
Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators
by: Orr, Will, et al.
Published: (2024)
by: Orr, Will, et al.
Published: (2024)
Composite Goodness-of-fit Tests with Kernels
by: Key, Oscar, et al.
Published: (2021)
by: Key, Oscar, et al.
Published: (2021)
A Faster Generalized Two-Stage Approximate Top-K
by: Samaga, Yashas, et al.
Published: (2025)
by: Samaga, Yashas, et al.
Published: (2025)
Reducing Communication for Split Learning by Randomized Top-k Sparsification
by: Zheng, Fei, et al.
Published: (2023)
by: Zheng, Fei, et al.
Published: (2023)
Regularized Top-$k$: A Bayesian Framework for Gradient Sparsification
by: Bereyhi, Ali, et al.
Published: (2025)
by: Bereyhi, Ali, et al.
Published: (2025)
A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
by: Xiu, Di, et al.
Published: (2025)
by: Xiu, Di, et al.
Published: (2025)
$\mathbb{R}^{2k}$ is Theoretically Large Enough for Embedding-based Top-$k$ Retrieval
by: Wang, Zihao, et al.
Published: (2026)
by: Wang, Zihao, et al.
Published: (2026)
Top-k on a Budget: Adaptive Ranking with Weak and Strong Oracles
by: Oettershagen, Lutz
Published: (2026)
by: Oettershagen, Lutz
Published: (2026)
Optimizing Novelty of Top-k Recommendations using Large Language Models and Reinforcement Learning
by: Sharma, Amit, et al.
Published: (2024)
by: Sharma, Amit, et al.
Published: (2024)
Efficient and Responsible Adaptation of Large Language Models for Robust and Equitable Top-k Recommendations
by: Kaur, Kirandeep, et al.
Published: (2025)
by: Kaur, Kirandeep, et al.
Published: (2025)
Scalable and Precise Patch Robustness Certification for Deep Learning Models with Top-k Predictions
by: Zhou, Qilin, et al.
Published: (2025)
by: Zhou, Qilin, et al.
Published: (2025)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
by: Tzachristas, Georgios, et al.
Published: (2025)
by: Tzachristas, Georgios, et al.
Published: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
by: Zeng, Zhiyuan, et al.
Published: (2024)
by: Zeng, Zhiyuan, et al.
Published: (2024)
Neural Likelihood Approximation for Integer Valued Time Series Data
by: O'Loughlin, Luke, et al.
Published: (2023)
by: O'Loughlin, Luke, et al.
Published: (2023)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights
by: Wen, Qishuai, et al.
Published: (2026)
by: Wen, Qishuai, et al.
Published: (2026)
Similar Items
-
SparQ Attention: Bandwidth-Efficient LLM Inference
by: Ribar, Luka, et al.
Published: (2023) -
Optimal Formats for Weight Quantisation
by: Orr, Douglas, et al.
Published: (2025) -
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
by: L, Yashas Samaga B, et al.
Published: (2024) -
Top-$k$ Feature Importance Ranking
by: Chen, Yuxi, et al.
Published: (2025) -
Top-$k$ Classification and Cardinality-Aware Prediction
by: Mao, Anqi, et al.
Published: (2024)