HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | L, Yashas Samaga B, Yerram, Varun, You, Chong, Bhojanapalli, Srinadh, Kumar, Sanjiv, Jain, Prateek, Netrapalli, Praneeth |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Faster Generalized Two-Stage Approximate Top-K
by: Samaga, Yashas, et al.
Published: (2025)
by: Samaga, Yashas, et al.
Published: (2025)
Tandem Transformers for Inference Efficient LLMs
by: S, Aishwarya P, et al.
Published: (2024)
by: S, Aishwarya P, et al.
Published: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Scalable In-context Ranking with Generative Models
by: Gupta, Nilesh, et al.
Published: (2025)
by: Gupta, Nilesh, et al.
Published: (2025)
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
Dual-Encoders for Extreme Multi-Label Classification
by: Gupta, Nilesh, et al.
Published: (2023)
by: Gupta, Nilesh, et al.
Published: (2023)
Spark Transformer: Reactivating Sparsity in FFN and Attention
by: You, Chong, et al.
Published: (2025)
by: You, Chong, et al.
Published: (2025)
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
by: Addepalli, Sravanti, et al.
Published: (2024)
by: Addepalli, Sravanti, et al.
Published: (2024)
A model of errors in transformers
by: Raju, Suvrat, et al.
Published: (2026)
by: Raju, Suvrat, et al.
Published: (2026)
The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks
by: Chizat, Lénaïc, et al.
Published: (2023)
by: Chizat, Lénaïc, et al.
Published: (2023)
Compressing Many-Shots in In-Context Learning
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
Functional Interpolation for Relative Positions Improves Long Context Transformers
by: Li, Shanda, et al.
Published: (2023)
by: Li, Shanda, et al.
Published: (2023)
Time-Reversal Provides Unsupervised Feedback to LLMs
by: Varun, Yerram, et al.
Published: (2024)
by: Varun, Yerram, et al.
Published: (2024)
Autoregressive Ranking: Bridging the Gap Between Dual and Cross Encoders
by: Rozonoyer, Benjamin, et al.
Published: (2026)
by: Rozonoyer, Benjamin, et al.
Published: (2026)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Efficient Language Model Architectures for Differentially Private Federated Learning
by: Ro, Jae Hun, et al.
Published: (2024)
by: Ro, Jae Hun, et al.
Published: (2024)
Approximating the Top Eigenvector in Random Order Streams
by: Kacham, Praneeth, et al.
Published: (2024)
by: Kacham, Praneeth, et al.
Published: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
Second Order Methods for Bandit Optimization and Control
by: Suggala, Arun, et al.
Published: (2024)
by: Suggala, Arun, et al.
Published: (2024)
Inferring Asteroseismic Parameters from Short Observations Using Deep Learning: Application to TESS and K2 Red Giants
by: Ghanghas, Nipun, et al.
Published: (2026)
by: Ghanghas, Nipun, et al.
Published: (2026)
HiPC25 Artifact: Energy-Aware Runtime Resource Harmonizer for Co-running Applications
by: Vanshika Jain, et al.
Published: (2025)
by: Vanshika Jain, et al.
Published: (2025)
Approximate Top-$k$ for Increased Parallelism
by: Key, Oscar, et al.
Published: (2024)
by: Key, Oscar, et al.
Published: (2024)
Approximating Opaque Top-k Queries
by: Chang, Jiwon, et al.
Published: (2025)
by: Chang, Jiwon, et al.
Published: (2025)
Estimating Causal Effects in Gaussian Linear SCMs with Finite Data
by: Maiti, Aurghya, et al.
Published: (2026)
by: Maiti, Aurghya, et al.
Published: (2026)
Top-k Approximate Functional Dependency Discovery
by: Wan, Xiaolong, et al.
Published: (2026)
by: Wan, Xiaolong, et al.
Published: (2026)
FastLane: Efficient Routed Systems for Late-Interaction Retrieval
by: Kumar, Ramnath, et al.
Published: (2026)
by: Kumar, Ramnath, et al.
Published: (2026)
A Methodological Approach Utilising Factor Analysis and Clustering to Evaluate the Strengths and Weaknesses of IPL Players
by: Ganesh Kumar Vadla, et al.
Published: (2025)
by: Ganesh Kumar Vadla, et al.
Published: (2025)
Universal Model Routing for Efficient LLM Inference
by: Jitkrittum, Wittawat, et al.
Published: (2025)
by: Jitkrittum, Wittawat, et al.
Published: (2025)
Impact and Educational Effectiveness of the Graphic Adaptation of Sapiens in Depicting Human Evolution
by: Kumar, Sanjiv
Published: (2022)
by: Kumar, Sanjiv
Published: (2022)
Phrasal Movement in Indian English Poetry
by: Sanjiv Kumar
Published: (2019)
by: Sanjiv Kumar
Published: (2019)
Efficient Algorithms for Complexes of Persistence Modules with Applications
by: Dey, Tamal K., et al.
Published: (2024)
by: Dey, Tamal K., et al.
Published: (2024)
Efficient evaluation of Top-k Skyline queries
by: Marlene Goncalves
Published: (2009)
by: Marlene Goncalves
Published: (2009)
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
by: Vadlapati, Praneeth
Published: (2024)
by: Vadlapati, Praneeth
Published: (2024)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
by: Gong, Ping, et al.
Published: (2025)
by: Gong, Ping, et al.
Published: (2025)
SimpliPy: A Source-Tracking Notional Machine for Simplified Python
by: Jain, Moida Praneeth, et al.
Published: (2025)
by: Jain, Moida Praneeth, et al.
Published: (2025)
Probably Approximately Precision and Recall Learning
by: Cohen, Lee, et al.
Published: (2024)
by: Cohen, Lee, et al.
Published: (2024)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
by: Kwon, Soo Min, et al.
Published: (2026)
by: Kwon, Soo Min, et al.
Published: (2026)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
Faster Algorithms for Schatten-p Low Rank Approximation
by: Kacham, Praneeth, et al.
Published: (2024)
by: Kacham, Praneeth, et al.
Published: (2024)
Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
by: You, Chong, et al.
Published: (2025)
by: You, Chong, et al.
Published: (2025)
Similar Items
-
A Faster Generalized Two-Stage Approximate Top-K
by: Samaga, Yashas, et al.
Published: (2025) -
Tandem Transformers for Inference Efficient LLMs
by: S, Aishwarya P, et al.
Published: (2024) -
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024) -
Scalable In-context Ranking with Generative Models
by: Gupta, Nilesh, et al.
Published: (2025) -
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023)