Is Context Helpful for Chat Translation Evaluation?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agrawal, Sweta, Farajian, Amin, Fernandes, Patrick, Rei, Ricardo, Martins, André F. T. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Automatic Metrics Assess High-Quality Translations?
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
A Context-aware Framework for Translation-mediated Conversations
von: Pombal, José, et al.
Veröffentlicht: (2024)
von: Pombal, José, et al.
Veröffentlicht: (2024)
Findings of the WMT 2024 Shared Task on Chat Translation
von: Mohammed, Wafaa, et al.
Veröffentlicht: (2024)
von: Mohammed, Wafaa, et al.
Veröffentlicht: (2024)
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2025)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2025)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
von: Fernandes, Patrick, et al.
Veröffentlicht: (2025)
von: Fernandes, Patrick, et al.
Veröffentlicht: (2025)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
von: Farinhas, António, et al.
Veröffentlicht: (2025)
von: Farinhas, António, et al.
Veröffentlicht: (2025)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
von: Alves, Duarte M., et al.
Veröffentlicht: (2024)
von: Alves, Duarte M., et al.
Veröffentlicht: (2024)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
von: Rei, Ricardo, et al.
Veröffentlicht: (2025)
von: Rei, Ricardo, et al.
Veröffentlicht: (2025)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
von: Zaranis, Emmanouil, et al.
Veröffentlicht: (2024)
von: Zaranis, Emmanouil, et al.
Veröffentlicht: (2024)
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2024)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2024)
QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
von: Faria, Gonçalo R. A., et al.
Veröffentlicht: (2024)
von: Faria, Gonçalo R. A., et al.
Veröffentlicht: (2024)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
von: Treviso, Marcos, et al.
Veröffentlicht: (2024)
von: Treviso, Marcos, et al.
Veröffentlicht: (2024)
Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
von: Agrawal, Sweta, et al.
Veröffentlicht: (2023)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2023)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
von: Pombal, José, et al.
Veröffentlicht: (2026)
von: Pombal, José, et al.
Veröffentlicht: (2026)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
Aligning Neural Machine Translation Models: Human Feedback in Training and Inference
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2023)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2023)
EuroLLM: Multilingual Language Models for Europe
von: Martins, Pedro Henrique, et al.
Veröffentlicht: (2024)
von: Martins, Pedro Henrique, et al.
Veröffentlicht: (2024)
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Analyzing Context Contributions in LLM-based Machine Translation
von: Zaranis, Emmanouil, et al.
Veröffentlicht: (2024)
von: Zaranis, Emmanouil, et al.
Veröffentlicht: (2024)
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
von: Rusli, Andre, et al.
Veröffentlicht: (2024)
von: Rusli, Andre, et al.
Veröffentlicht: (2024)
M-Prometheus: A Suite of Open Multilingual LLM Judges
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
Gradable ChatGPT Translation Evaluation
von: Jiao, Hui, et al.
Veröffentlicht: (2024)
von: Jiao, Hui, et al.
Veröffentlicht: (2024)
Does Context Help Mitigate Gender Bias in Neural Machine Translation?
von: Gete, Harritxu, et al.
Veröffentlicht: (2024)
von: Gete, Harritxu, et al.
Veröffentlicht: (2024)
EuroLLM-9B: Technical Report
von: Martins, Pedro Henrique, et al.
Veröffentlicht: (2025)
von: Martins, Pedro Henrique, et al.
Veröffentlicht: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
von: Han, HyoJung, et al.
Veröffentlicht: (2025)
von: Han, HyoJung, et al.
Veröffentlicht: (2025)
Did Translation Models Get More Robust Without Anyone Even Noticing?
von: Peters, Ben, et al.
Veröffentlicht: (2024)
von: Peters, Ben, et al.
Veröffentlicht: (2024)
EuroLLM-22B: Technical Report
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2026)
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2026)
DOCE: Finding the Sweet Spot for Execution-Based Code Generation
von: Li, Haau-Sing, et al.
Veröffentlicht: (2024)
von: Li, Haau-Sing, et al.
Veröffentlicht: (2024)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
von: Agrawal, Aryan, et al.
Veröffentlicht: (2025)
von: Agrawal, Aryan, et al.
Veröffentlicht: (2025)
XAMPLER: Learning to Retrieve Cross-Lingual In-Context Examples
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
How Effective are State Space Models for Machine Translation?
von: Pitorro, Hugo, et al.
Veröffentlicht: (2024)
von: Pitorro, Hugo, et al.
Veröffentlicht: (2024)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
Multi-Dimensional Evaluation of Text Summarization with In-Context Learning
von: Jain, Sameer, et al.
Veröffentlicht: (2023)
von: Jain, Sameer, et al.
Veröffentlicht: (2023)
Evaluating Multilingual Long-Context Models for Retrieval and Reasoning
von: Agrawal, Ameeta, et al.
Veröffentlicht: (2024)
von: Agrawal, Ameeta, et al.
Veröffentlicht: (2024)
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
von: Petty, Jackson, et al.
Veröffentlicht: (2026)
von: Petty, Jackson, et al.
Veröffentlicht: (2026)
What is the Best Way for ChatGPT to Translate Poetry?
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can Automatic Metrics Assess High-Quality Translations?
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024) -
A Context-aware Framework for Translation-mediated Conversations
von: Pombal, José, et al.
Veröffentlicht: (2024) -
Findings of the WMT 2024 Shared Task on Chat Translation
von: Mohammed, Wafaa, et al.
Veröffentlicht: (2024) -
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
von: Ramos, Miguel Moura, et al.
Veröffentlicht: (2025) -
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
von: Fernandes, Patrick, et al.
Veröffentlicht: (2025)