Dwell in the Beginning: How Language Models Embed Long Documents for Dense Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Coelho, João, Martins, Bruno, Magalhães, João, Callan, Jamie, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
von: Ning, Jingjie, et al.
Veröffentlicht: (2026)
von: Ning, Jingjie, et al.
Veröffentlicht: (2026)
Aligning Web Query Generation with Ranking Objectives via Direct Preference Optimization
von: Coelho, João, et al.
Veröffentlicht: (2025)
von: Coelho, João, et al.
Veröffentlicht: (2025)
ACER: Automatic Language Model Context Extension via Retrieval
von: Gao, Luyu, et al.
Veröffentlicht: (2024)
von: Gao, Luyu, et al.
Veröffentlicht: (2024)
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
von: Coelho, João, et al.
Veröffentlicht: (2025)
von: Coelho, João, et al.
Veröffentlicht: (2025)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
Improve Dense Passage Retrieval with Entailment Tuning
von: Dai, Lu, et al.
Veröffentlicht: (2024)
von: Dai, Lu, et al.
Veröffentlicht: (2024)
Making Large Language Models Efficient Dense Retrievers
von: Lei, Yibin, et al.
Veröffentlicht: (2025)
von: Lei, Yibin, et al.
Veröffentlicht: (2025)
Revela: Dense Retriever Learning via Language Modeling
von: Cai, Fengyu, et al.
Veröffentlicht: (2025)
von: Cai, Fengyu, et al.
Veröffentlicht: (2025)
ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval
von: Mao, Kelong, et al.
Veröffentlicht: (2024)
von: Mao, Kelong, et al.
Veröffentlicht: (2024)
Recall Them All: Retrieval-Augmented Language Models for Long Object List Extraction from Long Documents
von: Singhania, Sneha, et al.
Veröffentlicht: (2024)
von: Singhania, Sneha, et al.
Veröffentlicht: (2024)
Dense Passage Retrieval: Is it Retrieving?
von: Reichman, Benjamin, et al.
Veröffentlicht: (2024)
von: Reichman, Benjamin, et al.
Veröffentlicht: (2024)
Dense Retrieval for Low Resource Languages -- the Case of Amharic Language
von: Yeshambel, Tilahun, et al.
Veröffentlicht: (2025)
von: Yeshambel, Tilahun, et al.
Veröffentlicht: (2025)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
von: Chen, Jianlyu, et al.
Veröffentlicht: (2025)
von: Chen, Jianlyu, et al.
Veröffentlicht: (2025)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
Interpret and Control Dense Retrieval with Sparse Latent Features
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
Cohort Retrieval using Dense Passage Retrieval
von: Jadhav, Pranav
Veröffentlicht: (2025)
von: Jadhav, Pranav
Veröffentlicht: (2025)
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024)
von: Fang, Yan, et al.
Veröffentlicht: (2024)
BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2024)
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2024)
Translate-Distill: Learning Cross-Language Dense Retrieval by Translation and Distillation
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
Pooling and Semantic Shift: The Fundamental Challenges in Long Text Embedding and Retrieval
von: Gao, Hang, et al.
Veröffentlicht: (2026)
von: Gao, Hang, et al.
Veröffentlicht: (2026)
RaDeR: Reasoning-aware Dense Retrieval Models
von: Das, Debrup, et al.
Veröffentlicht: (2025)
von: Das, Debrup, et al.
Veröffentlicht: (2025)
Summarization-Based Document IDs for Generative Retrieval with Language Models
von: Li, Haoxin, et al.
Veröffentlicht: (2023)
von: Li, Haoxin, et al.
Veröffentlicht: (2023)
History-Aware Conversational Dense Retrieval
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
Less LLM, More Documents: Searching for Improved RAG
von: Ning, Jingjie, et al.
Veröffentlicht: (2025)
von: Ning, Jingjie, et al.
Veröffentlicht: (2025)
LMK > CLS: Landmark Pooling for Dense Embeddings
von: Doshi, Meet, et al.
Veröffentlicht: (2026)
von: Doshi, Meet, et al.
Veröffentlicht: (2026)
Enhancing Dense Retrievers' Robustness with Group-level Reweighting
von: Han, Peixuan, et al.
Veröffentlicht: (2023)
von: Han, Peixuan, et al.
Veröffentlicht: (2023)
Modeling Sequential Sentence Relation to Improve Cross-lingual Dense Retrieval
von: Zhang, Shunyu, et al.
Veröffentlicht: (2023)
von: Zhang, Shunyu, et al.
Veröffentlicht: (2023)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
von: Kobeissi, Amine, et al.
Veröffentlicht: (2026)
von: Kobeissi, Amine, et al.
Veröffentlicht: (2026)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
von: Sinha, Aarush
Veröffentlicht: (2025)
von: Sinha, Aarush
Veröffentlicht: (2025)
PLAID SHIRTTT for Large-Scale Streaming Dense Retrieval
von: Lawrie, Dawn, et al.
Veröffentlicht: (2024)
von: Lawrie, Dawn, et al.
Veröffentlicht: (2024)
Generative Dense Retrieval: Memory Can Be a Burden
von: Yuan, Peiwen, et al.
Veröffentlicht: (2024)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2024)
PairDistill: Pairwise Relevance Distillation for Dense Retrieval
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
GRITHopper: Decomposition-Free Multi-Hop Dense Retrieval
von: Erker, Justus-Jonas, et al.
Veröffentlicht: (2025)
von: Erker, Justus-Jonas, et al.
Veröffentlicht: (2025)
Study on LLMs for Promptagator-Style Dense Retriever Training
von: Gwon, Daniel, et al.
Veröffentlicht: (2025)
von: Gwon, Daniel, et al.
Veröffentlicht: (2025)
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
von: Ma, Yubo, et al.
Veröffentlicht: (2025)
von: Ma, Yubo, et al.
Veröffentlicht: (2025)
DREditor: An Time-efficient Approach for Building a Domain-specific Dense Retrieval Model
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
TempRetriever: Fusion-based Temporal Dense Passage Retrieval for Time-Sensitive Questions
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2025)
Unsupervised Multilingual Dense Retrieval via Generative Pseudo Labeling
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense Retrieval
von: Li, Yongkang, et al.
Veröffentlicht: (2025)
von: Li, Yongkang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
von: Ning, Jingjie, et al.
Veröffentlicht: (2026) -
Aligning Web Query Generation with Ranking Objectives via Direct Preference Optimization
von: Coelho, João, et al.
Veröffentlicht: (2025) -
ACER: Automatic Language Model Context Extension via Retrieval
von: Gao, Luyu, et al.
Veröffentlicht: (2024) -
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
von: Coelho, João, et al.
Veröffentlicht: (2025) -
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
von: Kang, Hao, et al.
Veröffentlicht: (2024)