Syntriever: How to Train Your Retriever with Synthetic Data from LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minsang, Baek, Seungjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Large Language Models on Cross-Cultural Values in Connection with Training Methodology
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
by: Kim, Minsang, et al.
Published: (2026)
by: Kim, Minsang, et al.
Published: (2026)
Measuring Sample Importance in Data Pruning for Language Models based on Information Entropy
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Hierarchical Position Embedding of Graphs with Landmarks and Clustering for Link Prediction
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
by: Asawa, Parth, et al.
Published: (2025)
by: Asawa, Parth, et al.
Published: (2025)
How to Train Data-Efficient LLMs
by: Sachdeva, Noveen, et al.
Published: (2024)
by: Sachdeva, Noveen, et al.
Published: (2024)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
Translation of Multifaceted Data without Re-Training of Machine Translation Systems
by: Moon, Hyeonseok, et al.
Published: (2024)
by: Moon, Hyeonseok, et al.
Published: (2024)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
by: Lee, Changhun, et al.
Published: (2025)
by: Lee, Changhun, et al.
Published: (2025)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
by: Hui, Zheng, et al.
Published: (2024)
by: Hui, Zheng, et al.
Published: (2024)
House of Cards: Massive Weights in LLMs
by: Oh, Jaehoon, et al.
Published: (2024)
by: Oh, Jaehoon, et al.
Published: (2024)
Not All Synthetic Data Is Yours to Learn From
by: Alemohammad, Sina, et al.
Published: (2026)
by: Alemohammad, Sina, et al.
Published: (2026)
sDPO: Don't Use Your Data All at Once
by: Kim, Dahyun, et al.
Published: (2024)
by: Kim, Dahyun, et al.
Published: (2024)
Cash or Comfort? How LLMs Value Your Inconvenience
by: Cedro, Mateusz, et al.
Published: (2025)
by: Cedro, Mateusz, et al.
Published: (2025)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
by: Xiong, Zheyang, et al.
Published: (2024)
by: Xiong, Zheyang, et al.
Published: (2024)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
How Persuasive is Your Context?
by: Nguyen, Tu, et al.
Published: (2025)
by: Nguyen, Tu, et al.
Published: (2025)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024)
by: Ziegler, Ingo, et al.
Published: (2024)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
How to Train Your Long-Context Visual Document Model
by: Veselka, Austin
Published: (2026)
by: Veselka, Austin
Published: (2026)
Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
by: Zhao, Siyun, et al.
Published: (2024)
by: Zhao, Siyun, et al.
Published: (2024)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
by: Koo, Hamin, et al.
Published: (2025)
by: Koo, Hamin, et al.
Published: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
by: Kim, Haechan, et al.
Published: (2026)
by: Kim, Haechan, et al.
Published: (2026)
Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
by: Subramanian, Seganrasan, et al.
Published: (2025)
by: Subramanian, Seganrasan, et al.
Published: (2025)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
by: He, Linda, et al.
Published: (2025)
by: He, Linda, et al.
Published: (2025)
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
by: Chen, Zihong, et al.
Published: (2025)
by: Chen, Zihong, et al.
Published: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
by: Bhattacharyya, Pramit, et al.
Published: (2024)
by: Bhattacharyya, Pramit, et al.
Published: (2024)
LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?
by: Tamo, J. Ben, et al.
Published: (2026)
by: Tamo, J. Ben, et al.
Published: (2026)
The Effects of Hallucinations in Synthetic Training Data for Relation Extraction
by: Rogulsky, Steven, et al.
Published: (2024)
by: Rogulsky, Steven, et al.
Published: (2024)
Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function
by: Kwok, Chin Yuen, et al.
Published: (2025)
by: Kwok, Chin Yuen, et al.
Published: (2025)
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Similar Items
-
Exploring Large Language Models on Cross-Cultural Values in Connection with Training Methodology
by: Kim, Minsang, et al.
Published: (2024) -
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
by: Kim, Minsang, et al.
Published: (2024) -
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
by: Kim, Minsang, et al.
Published: (2026) -
Measuring Sample Importance in Data Pruning for Language Models based on Information Entropy
by: Kim, Minsang, et al.
Published: (2024) -
Hierarchical Position Embedding of Graphs with Landmarks and Clustering for Link Prediction
by: Kim, Minsang, et al.
Published: (2024)