Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Thakur, Nandan, Ni, Jianmo, Ábrego, Gustavo Hernández, Wieting, John, Lin, Jimmy, Cer, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025)
by: Gwon, Daniel, et al.
Published: (2025)
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024)
by: Gomez, Frank Palma, et al.
Published: (2024)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026)
by: Kuissi, Nathan, et al.
Published: (2026)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)
by: Upadhyay, Shivani, et al.
Published: (2026)
CRISP: Clustering Multi-Vector Representations for Denoising and Pruning
by: Veneroso, João, et al.
Published: (2025)
by: Veneroso, João, et al.
Published: (2025)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
by: Thakur, Nandan, et al.
Published: (2026)
by: Thakur, Nandan, et al.
Published: (2026)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
by: Lin, Jimmy
Published: (2024)
by: Lin, Jimmy
Published: (2024)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2025)
by: Pradeep, Ronak, et al.
Published: (2025)
Leveraging LLMs for Unsupervised Dense Retriever Ranking
by: Khramtsova, Ekaterina, et al.
Published: (2024)
by: Khramtsova, Ekaterina, et al.
Published: (2024)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
Boosting Data Utilization for Multilingual Dense Retrieval
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
The Wisdom of Many Queries: Complexity-Diversity Principle for Dense Retriever Training
by: Feng, Xincan, et al.
Published: (2026)
by: Feng, Xincan, et al.
Published: (2026)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
by: Zhuang, Shengyao, et al.
Published: (2024)
by: Zhuang, Shengyao, et al.
Published: (2024)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
The Overlooked Role of Graded Relevance Thresholds in Multilingual Dense Retrieval
by: Wullach, Tomer, et al.
Published: (2026)
by: Wullach, Tomer, et al.
Published: (2026)
Multilingual Humour-Aware Retrieval with Dense and Re-Ranking Models
by: Arampatzis, Georgios, et al.
Published: (2026)
by: Arampatzis, Georgios, et al.
Published: (2026)
LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
by: Xu, Zhichao, et al.
Published: (2026)
by: Xu, Zhichao, et al.
Published: (2026)
Unifying Adversarial Robustness and Training Across Text Scoring Models
by: Tamber, Manveer Singh, et al.
Published: (2026)
by: Tamber, Manveer Singh, et al.
Published: (2026)
Training Dense Retrievers with Multiple Positive Passages
by: Wang, Benben, et al.
Published: (2026)
by: Wang, Benben, et al.
Published: (2026)
Unsupervised Multilingual Dense Retrieval via Generative Pseudo Labeling
by: Huang, Chao-Wei, et al.
Published: (2024)
by: Huang, Chao-Wei, et al.
Published: (2024)
Scaling Sparse and Dense Retrieval in Decoder-Only LLMs
by: Zeng, Hansi, et al.
Published: (2025)
by: Zeng, Hansi, et al.
Published: (2025)
Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
by: Sharifymoghaddam, Sahel, et al.
Published: (2024)
by: Sharifymoghaddam, Sahel, et al.
Published: (2024)
Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System
by: Balaka, Muhammad Imam Luthfi, et al.
Published: (2025)
by: Balaka, Muhammad Imam Luthfi, et al.
Published: (2025)
MFBE: Leveraging Multi-Field Information of FAQs for Efficient Dense Retrieval
by: Banerjee, Debopriyo, et al.
Published: (2023)
by: Banerjee, Debopriyo, et al.
Published: (2023)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Multilingual Prompts in LLM-Based Recommenders: Performance Across Languages
by: Ozsoy, Makbule Gulcin
Published: (2024)
by: Ozsoy, Makbule Gulcin
Published: (2024)
Language Fairness in Multilingual Information Retrieval
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
Similar Items
-
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025) -
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025) -
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024) -
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026) -
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)