Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Jongseok, Kim, Sunga, Cheung, Alvin, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)
LapSum -- One Method to Differentiate Them All: Ranking, Sorting and Top-k Selection
von: Struski, Łukasz, et al.
Veröffentlicht: (2025)
von: Struski, Łukasz, et al.
Veröffentlicht: (2025)
Redundancy-Driven Top-$k$ Functional Dependency Discovery
von: Wan, Xiaolong, et al.
Veröffentlicht: (2026)
von: Wan, Xiaolong, et al.
Veröffentlicht: (2026)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Foundations of Top-$k$ Decoding For Language Models
von: Noarov, Georgy, et al.
Veröffentlicht: (2025)
von: Noarov, Georgy, et al.
Veröffentlicht: (2025)
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024)
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024)
$\mathbb{R}^{2k}$ is Theoretically Large Enough for Embedding-based Top-$k$ Retrieval
von: Wang, Zihao, et al.
Veröffentlicht: (2026)
von: Wang, Zihao, et al.
Veröffentlicht: (2026)
Antithetic Sampling for Top-k Shapley Identification
von: Kolpaczki, Patrick, et al.
Veröffentlicht: (2025)
von: Kolpaczki, Patrick, et al.
Veröffentlicht: (2025)
Learning Top-k Subtask Planning Tree based on Discriminative Representation Pre-training for Decision Making
von: Ruan, Jingqing, et al.
Veröffentlicht: (2023)
von: Ruan, Jingqing, et al.
Veröffentlicht: (2023)
Personalized Top-k Set Queries Over Predicted Scores
von: Nia, Sohrab Namazi, et al.
Veröffentlicht: (2025)
von: Nia, Sohrab Namazi, et al.
Veröffentlicht: (2025)
Shiva-DiT: Residual-Based Differentiable Top-$k$ Selection for Efficient Diffusion Transformers
von: Zhang, Jiaji, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaji, et al.
Veröffentlicht: (2026)
ZETA: Leveraging Z-order Curves for Efficient Top-k Attention
von: Zeng, Qiuhao, et al.
Veröffentlicht: (2025)
von: Zeng, Qiuhao, et al.
Veröffentlicht: (2025)
Optimizing Partial Area Under the Top-k Curve: Theory and Practice
von: Wang, Zitai, et al.
Veröffentlicht: (2022)
von: Wang, Zitai, et al.
Veröffentlicht: (2022)
Top Pass: Improve Code Generation by Pass@k-Maximized Code Ranking
von: Lyu, Zhi-Cun, et al.
Veröffentlicht: (2024)
von: Lyu, Zhi-Cun, et al.
Veröffentlicht: (2024)
Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling
von: Yao, Yuxuan, et al.
Veröffentlicht: (2024)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2024)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
CLIP Model for Images to Textual Prompts Based on Top-k Neighbors
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
Denoising the Future: Top-p Distributions for Moving Through Time
von: Marwitz, Florian Andreas, et al.
Veröffentlicht: (2025)
von: Marwitz, Florian Andreas, et al.
Veröffentlicht: (2025)
Unifying and Certifying Top-Quality Planning
von: Katz, Michael, et al.
Veröffentlicht: (2024)
von: Katz, Michael, et al.
Veröffentlicht: (2024)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
von: Tzachristas, Georgios, et al.
Veröffentlicht: (2025)
von: Tzachristas, Georgios, et al.
Veröffentlicht: (2025)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
von: Zhang, Lunjun, et al.
Veröffentlicht: (2026)
von: Zhang, Lunjun, et al.
Veröffentlicht: (2026)
Causality from Bottom to Top: A Survey
von: Weinberg, Abraham Itzhak, et al.
Veröffentlicht: (2024)
von: Weinberg, Abraham Itzhak, et al.
Veröffentlicht: (2024)
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
von: Subbiah, Melanie, et al.
Veröffentlicht: (2025)
von: Subbiah, Melanie, et al.
Veröffentlicht: (2025)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
P$^2$RAG: Efficient Privacy-Preserving RAG Service Supporting Arbitrary Top-$k$ Retrieval
von: Ming, Yulong, et al.
Veröffentlicht: (2026)
von: Ming, Yulong, et al.
Veröffentlicht: (2026)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
von: Gong, Ping, et al.
Veröffentlicht: (2025)
von: Gong, Ping, et al.
Veröffentlicht: (2025)
Breaking the Top-$K$ Barrier: Advancing Top-$K$ Ranking Metrics Optimization in Recommender Systems
von: Yang, Weiqin, et al.
Veröffentlicht: (2025)
von: Yang, Weiqin, et al.
Veröffentlicht: (2025)
DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
von: Jin, Can, et al.
Veröffentlicht: (2025)
von: Jin, Can, et al.
Veröffentlicht: (2025)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
von: Li, Hanchen, et al.
Veröffentlicht: (2025)
von: Li, Hanchen, et al.
Veröffentlicht: (2025)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
von: Cao, Shiyi, et al.
Veröffentlicht: (2026)
von: Cao, Shiyi, et al.
Veröffentlicht: (2026)
On Computing Top-$k$ Simple Shortest Paths from a Single Source
von: D'Emidio, Mattia, et al.
Veröffentlicht: (2025)
von: D'Emidio, Mattia, et al.
Veröffentlicht: (2025)
Tracking Skiers from the Top to the Bottom
von: Dunnhofer, Matteo, et al.
Veröffentlicht: (2023)
von: Dunnhofer, Matteo, et al.
Veröffentlicht: (2023)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and More
von: Lim, Jinwoo, et al.
Veröffentlicht: (2026)
von: Lim, Jinwoo, et al.
Veröffentlicht: (2026)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
von: Kwak, Minseo, et al.
Veröffentlicht: (2026)
von: Kwak, Minseo, et al.
Veröffentlicht: (2026)
CNC-TP: Classifier Nominal Concept Based on Top-Pertinent Attributes
von: Souissi, Yasmine, et al.
Veröffentlicht: (2026)
von: Souissi, Yasmine, et al.
Veröffentlicht: (2026)
The Mercurial Top-Level Ontology of Large Language Models
von: Köhler, Nele, et al.
Veröffentlicht: (2024)
von: Köhler, Nele, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026) -
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025) -
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024) -
LapSum -- One Method to Differentiate Them All: Ranking, Sorting and Top-k Selection
von: Struski, Łukasz, et al.
Veröffentlicht: (2025) -
Redundancy-Driven Top-$k$ Functional Dependency Discovery
von: Wan, Xiaolong, et al.
Veröffentlicht: (2026)