Data, Not Model: Explaining Bias toward LLM Texts in Neural Retrievers
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Wei, Bi, Keping, Cai, Yinqiong, Chen, Wei, Guo, Jiafeng, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Bridging Queries and Tables through Entities in Table Retrieval
by: Li, Da, et al.
Published: (2025)
by: Li, Da, et al.
Published: (2025)
Reproducibility Analysis and Enhancements for Multi-Aspect Dense Retriever with Aspect Learning
by: Bi, Keping, et al.
Published: (2024)
by: Bi, Keping, et al.
Published: (2024)
Tailoring Table Retrieval from a Field-aware Hybrid Matching Perspective
by: Li, Da, et al.
Published: (2025)
by: Li, Da, et al.
Published: (2025)
CIR at the NTCIR-17 ULTRE-2 Task
by: Yu, Lulu, et al.
Published: (2023)
by: Yu, Lulu, et al.
Published: (2023)
Can LLM Annotations Replace User Clicks for Learning to Rank?
by: Yu, Lulu, et al.
Published: (2025)
by: Yu, Lulu, et al.
Published: (2025)
A Multi-Granularity-Aware Aspect Learning Model for Multi-Aspect Dense Retrieval
by: Sun, Xiaojie, et al.
Published: (2023)
by: Sun, Xiaojie, et al.
Published: (2023)
LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Unbiased Learning to Rank with Query-Level Click Propensity Estimation: Beyond Pointwise Observation and Relevance
by: Yu, Lulu, et al.
Published: (2025)
by: Yu, Lulu, et al.
Published: (2025)
Bagging-Based Model Merging for Robust General Text Embeddings
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
Training Dense Retrievers with Multiple Positive Passages
by: Wang, Benben, et al.
Published: (2026)
by: Wang, Benben, et al.
Published: (2026)
Beyond Relevance: Utility-Centric Retrieval in the LLM Era
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality
by: Chen, Jiahan, et al.
Published: (2026)
by: Chen, Jiahan, et al.
Published: (2026)
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
A Comparative Study of Specialized LLMs as Dense Retrievers
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Listwise Generative Retrieval Models via a Sequential Learning Process
by: Tang, Yubao, et al.
Published: (2024)
by: Tang, Yubao, et al.
Published: (2024)
Generative Retrieval Meets Multi-Graded Relevance
by: Tang, Yubao, et al.
Published: (2024)
by: Tang, Yubao, et al.
Published: (2024)
Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Continual Learning for Generative Retrieval over Dynamic Corpora
by: Chen, Jiangui, et al.
Published: (2023)
by: Chen, Jiangui, et al.
Published: (2023)
On the Robustness of Generative Information Retrieval Models
by: Liu, Yu-An, et al.
Published: (2024)
by: Liu, Yu-An, et al.
Published: (2024)
AdversarialCoT: Single-Document Retrieval Poisoning for LLM Reasoning
by: Song, Hongru, et al.
Published: (2026)
by: Song, Hongru, et al.
Published: (2026)
Contextual Dual Learning Algorithm with Listwise Distillation for Unbiased Learning to Rank
by: Yu, Lulu, et al.
Published: (2024)
by: Yu, Lulu, et al.
Published: (2024)
Does Generative Retrieval Overcome the Limitations of Dense Retrieval?
by: Zhang, Yingchen, et al.
Published: (2025)
by: Zhang, Yingchen, et al.
Published: (2025)
On the Scaling of Robustness and Effectiveness in Dense Retrieval
by: Liu, Yu-An, et al.
Published: (2025)
by: Liu, Yu-An, et al.
Published: (2025)
Attack-in-the-Chain: Bootstrapping Large Language Models for Attacks Against Black-box Neural Ranking Models
by: Liu, Yu-An, et al.
Published: (2024)
by: Liu, Yu-An, et al.
Published: (2024)
From Relevance to Utility: Evidence Retrieval with Feedback for Fact Verification
by: Zhang, Hengran, et al.
Published: (2023)
by: Zhang, Hengran, et al.
Published: (2023)
Bootstrapped Pre-training with Dynamic Identifier Prediction for Generative Retrieval
by: Tang, Yubao, et al.
Published: (2024)
by: Tang, Yubao, et al.
Published: (2024)
VeriCite: Towards Reliable Citations in Retrieval-Augmented Generation via Rigorous Verification
by: Qian, Haosheng, et al.
Published: (2025)
by: Qian, Haosheng, et al.
Published: (2025)
A Generative Framework for Personalized Sticker Retrieval
by: Zhou, Changjiang, et al.
Published: (2025)
by: Zhou, Changjiang, et al.
Published: (2025)
RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
by: Guo, Yucan, et al.
Published: (2025)
by: Guo, Yucan, et al.
Published: (2025)
TrustRAG: An Information Assistant with Retrieval Augmented Generation
by: Fan, Yixing, et al.
Published: (2025)
by: Fan, Yixing, et al.
Published: (2025)
Generative Retrieval for Book search
by: Tang, Yubao, et al.
Published: (2025)
by: Tang, Yubao, et al.
Published: (2025)
The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems
by: Song, Hongru, et al.
Published: (2025)
by: Song, Hongru, et al.
Published: (2025)
Robust Neural Information Retrieval: An Adversarial and Out-of-distribution Perspective
by: Liu, Yu-An, et al.
Published: (2024)
by: Liu, Yu-An, et al.
Published: (2024)
Are Large Language Models Good at Utility Judgments?
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Few-shot Link Prediction on N-ary Facts
by: Wei, Jiyao, et al.
Published: (2023)
by: Wei, Jiyao, et al.
Published: (2023)
LifeIR at the NTCIR-18 Lifelog-6 Task
by: Chen, Jiahan, et al.
Published: (2025)
by: Chen, Jiahan, et al.
Published: (2025)
Similar Items
-
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
by: Huang, Wei, et al.
Published: (2025) -
Bridging Queries and Tables through Entities in Table Retrieval
by: Li, Da, et al.
Published: (2025) -
Reproducibility Analysis and Enhancements for Multi-Aspect Dense Retriever with Aspect Learning
by: Bi, Keping, et al.
Published: (2024) -
Tailoring Table Retrieval from a Field-aware Hybrid Matching Perspective
by: Li, Da, et al.
Published: (2025) -
CIR at the NTCIR-17 ULTRE-2 Task
by: Yu, Lulu, et al.
Published: (2023)