From Tokens to Concepts: Leveraging SAE for SPLADE
Fuente:
arXiv
Saved in:
| Main Authors: | Zong, Yuxuan, Vast, Mathias, Van Cooten, Basile, Soulier, Laure, Piwowarski, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simple Domain Adaptation for Sparse Retrievers
by: Vast, Mathias, et al.
Published: (2024)
by: Vast, Mathias, et al.
Published: (2024)
Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders
by: Vast, Mathias, et al.
Published: (2024)
by: Vast, Mathias, et al.
Published: (2024)
Understanding Matching Mechanisms in Cross-Encoders
by: Vast, Mathias, et al.
Published: (2025)
by: Vast, Mathias, et al.
Published: (2025)
Reproducing and Comparing Distillation Techniques for Cross-Encoders
by: Morand, Victor, et al.
Published: (2026)
by: Morand, Victor, et al.
Published: (2026)
MICE: Minimal Interaction Cross-Encoders for efficient Re-ranking
by: Vast, Mathias, et al.
Published: (2026)
by: Vast, Mathias, et al.
Published: (2026)
Towards Lossless Token Pruning in Late-Interaction Retrieval Models
by: Zong, Yuxuan, et al.
Published: (2025)
by: Zong, Yuxuan, et al.
Published: (2025)
CoSPLADE: Contextualizing SPLADE for Conversational Information Retrieval
by: Hai, Nam Le, et al.
Published: (2023)
by: Hai, Nam Le, et al.
Published: (2023)
SPLADE-v3: New baselines for SPLADE
by: Lassance, Carlos, et al.
Published: (2024)
by: Lassance, Carlos, et al.
Published: (2024)
Navigating Uncertainty: Optimizing API Dependency for Hallucination Reduction in Closed-Book Question Answering
by: Erbacher, Pierre, et al.
Published: (2024)
by: Erbacher, Pierre, et al.
Published: (2024)
An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
by: Porco, Aldo, et al.
Published: (2025)
by: Porco, Aldo, et al.
Published: (2025)
Mistral-SPLADE: LLMs for better Learned Sparse Retrieval
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
PAQA: Toward ProActive Open-Retrieval Question Answering
by: Erbacher, Pierre, et al.
Published: (2024)
by: Erbacher, Pierre, et al.
Published: (2024)
The Pre-Training Study of Expanded-SPLADE Models on Web Document Titles
by: Kim, Hiun, et al.
Published: (2026)
by: Kim, Hiun, et al.
Published: (2026)
RAC: Retrieval-Augmented Clarification for Faithful Conversational Search
by: Kebir, Ahmed Rayane, et al.
Published: (2026)
by: Kebir, Ahmed Rayane, et al.
Published: (2026)
QueStER: Query Specification for Generative keyword-based Retrieval
by: Satouf, Arthur, et al.
Published: (2025)
by: Satouf, Arthur, et al.
Published: (2025)
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models
by: Kankanampati, Yash, et al.
Published: (2026)
by: Kankanampati, Yash, et al.
Published: (2026)
Two-Step SPLADE: Simple, Efficient and Effective Approximation of SPLADE
by: Lassance, Carlos, et al.
Published: (2024)
by: Lassance, Carlos, et al.
Published: (2024)
Rational Retrieval Acts: Leveraging Pragmatic Reasoning to Improve Sparse Retrieval
by: Satouf, Arthur, et al.
Published: (2025)
by: Satouf, Arthur, et al.
Published: (2025)
xVLM2Vec: Adapting LVLM-based embedding models to multilinguality using Self-Knowledge Distillation
by: Musacchio, Elio, et al.
Published: (2025)
by: Musacchio, Elio, et al.
Published: (2025)
Clarifying Ambiguities: on the Role of Ambiguity Types in Prompting Methods for Clarification Generation
by: Tang, Anfu, et al.
Published: (2025)
by: Tang, Anfu, et al.
Published: (2025)
HeceTokenizer: A Syllable-Based Tokenization Approach for Turkish Retrieval
by: Gulgonul, Senol
Published: (2026)
by: Gulgonul, Senol
Published: (2026)
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
by: An, Ruize, et al.
Published: (2025)
by: An, Ruize, et al.
Published: (2025)
Contextualization with SPLADE for High Recall Retrieval
by: Yang, Eugene
Published: (2024)
by: Yang, Eugene
Published: (2024)
NevIR: Negation in Neural Information Retrieval
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
PromptLink: Leveraging Large Language Models for Cross-Source Biomedical Concept Linking
by: Xie, Yuzhang, et al.
Published: (2024)
by: Xie, Yuzhang, et al.
Published: (2024)
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
Better Generalizing to Unseen Concepts: An Evaluation Framework and An LLM-Based Auto-Labeled Pipeline for Biomedical Concept Recognition
by: Liu, Shanshan, et al.
Published: (2026)
by: Liu, Shanshan, et al.
Published: (2026)
Multi-Field Adaptive Retrieval
by: Li, Millicent, et al.
Published: (2024)
by: Li, Millicent, et al.
Published: (2024)
Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
by: Lee, Jinhyuk, et al.
Published: (2023)
by: Lee, Jinhyuk, et al.
Published: (2023)
ConExion: Concept Extraction with Large Language Models
by: Norouzi, Ebrahim, et al.
Published: (2025)
by: Norouzi, Ebrahim, et al.
Published: (2025)
Token and Span Classification for Entity Recognition in French Historical Encyclopedias
by: Moncla, Ludovic, et al.
Published: (2025)
by: Moncla, Ludovic, et al.
Published: (2025)
Defending Against Disinformation Attacks in Open-Domain Question Answering
by: Weller, Orion, et al.
Published: (2022)
by: Weller, Orion, et al.
Published: (2022)
FIT-RAG: Black-Box RAG with Factual Information and Token Reduction
by: Mao, Yuren, et al.
Published: (2024)
by: Mao, Yuren, et al.
Published: (2024)
Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
by: Albarede, Lucas, et al.
Published: (2025)
by: Albarede, Lucas, et al.
Published: (2025)
Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
by: Clavié, Benjamin, et al.
Published: (2024)
by: Clavié, Benjamin, et al.
Published: (2024)
Improving Retrieval in Sponsored Search by Leveraging Query Context Signals
by: Mohankumar, Akash Kumar, et al.
Published: (2024)
by: Mohankumar, Akash Kumar, et al.
Published: (2024)
FGR-ColBERT: Identifying Fine-Grained Relevance Tokens During Retrieval
by: Jarolím, Antonín, et al.
Published: (2026)
by: Jarolím, Antonín, et al.
Published: (2026)
An Early FIRST Reproduction and Improvements to Single-Token Decoding for Fast Listwise Reranking
by: Chen, Zijian, et al.
Published: (2024)
by: Chen, Zijian, et al.
Published: (2024)
Leveraging Passage Embeddings for Efficient Listwise Reranking with Large Language Models
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
by: Mousavian, Maryam, et al.
Published: (2025)
by: Mousavian, Maryam, et al.
Published: (2025)
Similar Items
-
Simple Domain Adaptation for Sparse Retrievers
by: Vast, Mathias, et al.
Published: (2024) -
Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders
by: Vast, Mathias, et al.
Published: (2024) -
Understanding Matching Mechanisms in Cross-Encoders
by: Vast, Mathias, et al.
Published: (2025) -
Reproducing and Comparing Distillation Techniques for Cross-Encoders
by: Morand, Victor, et al.
Published: (2026) -
MICE: Minimal Interaction Cross-Encoders for efficient Re-ranking
by: Vast, Mathias, et al.
Published: (2026)