Improving LLMs for Machine Translation Using Synthetic Preference Data
Fuente:
arXiv
Saved in:
| Main Authors: | Vajda, Dario, Vreš, Domen, Robnik-Šikonja, Marko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Building a Strong Instruction Language Model for a Less-Resourced Language
by: Vreš, Domen, et al.
Published: (2026)
by: Vreš, Domen, et al.
Published: (2026)
Generative Model for Less-Resourced Language with 1 billion parameters
by: Vreš, Domen, et al.
Published: (2024)
by: Vreš, Domen, et al.
Published: (2024)
Solving Word-Sense Disambiguation and Word-Sense Induction with Dictionary Examples
by: Škvorc, Tadej, et al.
Published: (2025)
by: Škvorc, Tadej, et al.
Published: (2025)
QFS-Composer: Query-focused summarization pipeline for less resourced languages
by: Đuranović, Vuk, et al.
Published: (2026)
by: Đuranović, Vuk, et al.
Published: (2026)
Sarcasm Detection in a Less-Resourced Language
by: Đoković, Lazar, et al.
Published: (2024)
by: Đoković, Lazar, et al.
Published: (2024)
Large language models for folktale type automation based on motifs: Cinderella case study
by: Arčon, Tjaša, et al.
Published: (2025)
by: Arčon, Tjaša, et al.
Published: (2025)
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
by: Klemen, Matej, et al.
Published: (2025)
by: Klemen, Matej, et al.
Published: (2025)
Real-time News Story Identification
by: Škvorc, Tadej, et al.
Published: (2025)
by: Škvorc, Tadej, et al.
Published: (2025)
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
by: Arčon, Tjaša, et al.
Published: (2026)
by: Arčon, Tjaša, et al.
Published: (2026)
Challenges in Explaining Pretrained Clinical Text Classifiers
by: Miok, Kristian, et al.
Published: (2026)
by: Miok, Kristian, et al.
Published: (2026)
TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
by: Miok, Kristian, et al.
Published: (2025)
by: Miok, Kristian, et al.
Published: (2025)
Measuring Catastrophic Forgetting in Cross-Lingual Transfer Paradigms: Exploring Tuning Strategies
by: Koloski, Boshko, et al.
Published: (2023)
by: Koloski, Boshko, et al.
Published: (2023)
Neural spell-checker: Beyond words with synthetic data generation
by: Klemen, Matej, et al.
Published: (2024)
by: Klemen, Matej, et al.
Published: (2024)
Incremental Graph Construction Enables Robust Spectral Clustering of Texts
by: Pranjić, Marko, et al.
Published: (2026)
by: Pranjić, Marko, et al.
Published: (2026)
Code-mixed Sentiment and Hate-speech Prediction
by: Yadav, Anjali, et al.
Published: (2024)
by: Yadav, Anjali, et al.
Published: (2024)
Retrieval-augmented code completion for local projects using large language models
by: Hostnik, Marko, et al.
Published: (2024)
by: Hostnik, Marko, et al.
Published: (2024)
Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
by: Dhawan, Aashish, et al.
Published: (2026)
by: Dhawan, Aashish, et al.
Published: (2026)
Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation
by: Sánchez-Cartagena, Víctor M., et al.
Published: (2024)
by: Sánchez-Cartagena, Víctor M., et al.
Published: (2024)
Review of Natural Language Processing in Pharmacology
by: Trajanov, Dimitar, et al.
Published: (2022)
by: Trajanov, Dimitar, et al.
Published: (2022)
Alleviating Distribution Shift in Synthetic Data for Machine Translation Quality Estimation
by: Geng, Xiang, et al.
Published: (2025)
by: Geng, Xiang, et al.
Published: (2025)
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
by: Gallego, Víctor
Published: (2024)
by: Gallego, Víctor
Published: (2024)
Compensating for Data with Reasoning: Low-Resource Machine Translation with LLMs
by: Frontull, Samuel, et al.
Published: (2025)
by: Frontull, Samuel, et al.
Published: (2025)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
Word Alignment as Preference for Machine Translation
by: Wu, Qiyu, et al.
Published: (2024)
by: Wu, Qiyu, et al.
Published: (2024)
Setting up the Data Printer with Improved English to Ukrainian Machine Translation
by: Paniv, Yurii, et al.
Published: (2024)
by: Paniv, Yurii, et al.
Published: (2024)
Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation
by: Ghassabi, Mehrdad, et al.
Published: (2026)
by: Ghassabi, Mehrdad, et al.
Published: (2026)
Cross-Preference Learning for Sentence-Level and Context-Aware Machine Translation
by: Li, Ying, et al.
Published: (2026)
by: Li, Ying, et al.
Published: (2026)
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
by: Xu, Haoran, et al.
Published: (2024)
by: Xu, Haoran, et al.
Published: (2024)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
Improving Vietnamese-English Medical Machine Translation
by: Vo, Nhu, et al.
Published: (2024)
by: Vo, Nhu, et al.
Published: (2024)
Prompting LLMs: Length Control for Isometric Machine Translation
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
An Empirical Study of In-context Learning in LLMs for Machine Translation
by: Chitale, Pranjal A., et al.
Published: (2024)
by: Chitale, Pranjal A., et al.
Published: (2024)
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
by: Bouthors, Maxime, et al.
Published: (2025)
by: Bouthors, Maxime, et al.
Published: (2025)
Strategies for Improving NL-to-FOL Translation with LLMs: Data Generation, Incremental Fine-Tuning, and Verification
by: Thatikonda, Ramya Keerthy, et al.
Published: (2024)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2024)
Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
Direct Preference Optimization for Neural Machine Translation with Minimum Bayes Risk Decoding
by: Yang, Guangyu, et al.
Published: (2023)
by: Yang, Guangyu, et al.
Published: (2023)
SimulPL: Aligning Human Preferences in Simultaneous Machine Translation
by: Yu, Donglei, et al.
Published: (2025)
by: Yu, Donglei, et al.
Published: (2025)
Configurable Preference Tuning with Rubric-Guided Synthetic Data
by: Gallego, Víctor
Published: (2025)
by: Gallego, Víctor
Published: (2025)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Conditioning LLMs with Emotion in Neural Machine Translation
by: Brazier, Charles, et al.
Published: (2024)
by: Brazier, Charles, et al.
Published: (2024)
Similar Items
-
Building a Strong Instruction Language Model for a Less-Resourced Language
by: Vreš, Domen, et al.
Published: (2026) -
Generative Model for Less-Resourced Language with 1 billion parameters
by: Vreš, Domen, et al.
Published: (2024) -
Solving Word-Sense Disambiguation and Word-Sense Induction with Dictionary Examples
by: Škvorc, Tadej, et al.
Published: (2025) -
QFS-Composer: Query-focused summarization pipeline for less resourced languages
by: Đuranović, Vuk, et al.
Published: (2026) -
Sarcasm Detection in a Less-Resourced Language
by: Đoković, Lazar, et al.
Published: (2024)