Towards Lossless Token Pruning in Late-Interaction Retrieval Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zong, Yuxuan, Piwowarski, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models
by: Kankanampati, Yash, et al.
Published: (2026)
by: Kankanampati, Yash, et al.
Published: (2026)
From Tokens to Concepts: Leveraging SAE for SPLADE
by: Zong, Yuxuan, et al.
Published: (2026)
by: Zong, Yuxuan, et al.
Published: (2026)
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
by: Ezerceli, Özay, et al.
Published: (2025)
by: Ezerceli, Özay, et al.
Published: (2025)
Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models
by: Edy, Antoine, et al.
Published: (2026)
by: Edy, Antoine, et al.
Published: (2026)
Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
by: Clavié, Benjamin, et al.
Published: (2024)
by: Clavié, Benjamin, et al.
Published: (2024)
PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
A Theory for Token-Level Harmonization in Retrieval-Augmented Generation
by: Xu, Shicheng, et al.
Published: (2024)
by: Xu, Shicheng, et al.
Published: (2024)
Semi-Parametric Retrieval via Binary Bag-of-Tokens Index
by: Zhou, Jiawei, et al.
Published: (2024)
by: Zhou, Jiawei, et al.
Published: (2024)
Multi-Source Knowledge Pruning for Retrieval-Augmented Generation: A Benchmark and Empirical Study
by: Yu, Shuo, et al.
Published: (2024)
by: Yu, Shuo, et al.
Published: (2024)
OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation
by: Fang, Haoyang, et al.
Published: (2026)
by: Fang, Haoyang, et al.
Published: (2026)
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
JaColBERTv2.5: Optimising Multi-Vector Retrievers to Create State-of-the-Art Japanese Retrievers with Constrained Resources
by: Clavié, Benjamin
Published: (2024)
by: Clavié, Benjamin
Published: (2024)
Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
by: Lewis, Sydney
Published: (2026)
by: Lewis, Sydney
Published: (2026)
Towards Enhancing Linked Data Retrieval in Conversational UIs using Large Language Models
by: Mussa, Omar, et al.
Published: (2024)
by: Mussa, Omar, et al.
Published: (2024)
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
by: Fan, Wenqi, et al.
Published: (2024)
by: Fan, Wenqi, et al.
Published: (2024)
Neural Retrievers are Biased Towards LLM-Generated Content
by: Dai, Sunhao, et al.
Published: (2023)
by: Dai, Sunhao, et al.
Published: (2023)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
by: Killingback, Julian, et al.
Published: (2025)
by: Killingback, Julian, et al.
Published: (2025)
Token-wise Influential Training Data Retrieval for Large Language Models
by: Lin, Huawei, et al.
Published: (2024)
by: Lin, Huawei, et al.
Published: (2024)
Scaling Retrieval-Based Language Models with a Trillion-Token Datastore
by: Shao, Rulin, et al.
Published: (2024)
by: Shao, Rulin, et al.
Published: (2024)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
by: Kim, Seonok
Published: (2026)
by: Kim, Seonok
Published: (2026)
Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm
by: Li, Zihang, et al.
Published: (2026)
by: Li, Zihang, et al.
Published: (2026)
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
by: Shi, Zhengliang, et al.
Published: (2025)
by: Shi, Zhengliang, et al.
Published: (2025)
Towards Fair RAG: On the Impact of Fair Ranking in Retrieval-Augmented Generation
by: Kim, To Eun, et al.
Published: (2024)
by: Kim, To Eun, et al.
Published: (2024)
Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
Generative Retrieval and Alignment Model: A New Paradigm for E-commerce Retrieval
by: Pang, Ming, et al.
Published: (2025)
by: Pang, Ming, et al.
Published: (2025)
Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
by: Tang, Qiaoyu, et al.
Published: (2024)
by: Tang, Qiaoyu, et al.
Published: (2024)
TokenRec: Learning to Tokenize ID for LLM-based Generative Recommendation
by: Qu, Haohao, et al.
Published: (2024)
by: Qu, Haohao, et al.
Published: (2024)
Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models
by: Jeong, Minbyul, et al.
Published: (2024)
by: Jeong, Minbyul, et al.
Published: (2024)
Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies
by: Clavié, Benjamin, et al.
Published: (2026)
by: Clavié, Benjamin, et al.
Published: (2026)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Grounding Language Model with Chunking-Free In-Context Retrieval
by: Qian, Hongjin, et al.
Published: (2024)
by: Qian, Hongjin, et al.
Published: (2024)
PyLate: Flexible Training and Retrieval for Late Interaction Models
by: Chaffin, Antoine, et al.
Published: (2025)
by: Chaffin, Antoine, et al.
Published: (2025)
RARe: Retrieval Augmented Retrieval with In-Context Examples
by: Tejaswi, Atula, et al.
Published: (2024)
by: Tejaswi, Atula, et al.
Published: (2024)
Learning an Effective Premise Retrieval Model for Efficient Mathematical Formalization
by: Tao, Yicheng, et al.
Published: (2025)
by: Tao, Yicheng, et al.
Published: (2025)
GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation
by: Luo, Linhao, et al.
Published: (2025)
by: Luo, Linhao, et al.
Published: (2025)
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval
by: Chi, Yizhou, et al.
Published: (2024)
by: Chi, Yizhou, et al.
Published: (2024)
Thinking Ahead: Prospection-Guided Retrieval of Memory with Language Models
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
In-context Learning with Retrieved Demonstrations for Language Models: A Survey
by: Luo, Man, et al.
Published: (2024)
by: Luo, Man, et al.
Published: (2024)
DocReLM: Mastering Document Retrieval with Language Model
by: Wei, Gengchen, et al.
Published: (2024)
by: Wei, Gengchen, et al.
Published: (2024)
Similar Items
-
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models
by: Kankanampati, Yash, et al.
Published: (2026) -
From Tokens to Concepts: Leveraging SAE for SPLADE
by: Zong, Yuxuan, et al.
Published: (2026) -
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
by: Ezerceli, Özay, et al.
Published: (2025) -
Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models
by: Edy, Antoine, et al.
Published: (2026) -
Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
by: Clavié, Benjamin, et al.
Published: (2024)