Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
Fuente:
arXiv
Guardado en:
| Autores principales: | Fernandes, Patrick, Agrawal, Sweta, Zaranis, Emmanouil, Martins, André F. T., Neubig, Graham |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Context-aware Framework for Translation-mediated Conversations
por: Pombal, José, et al.
Publicado: (2024)
por: Pombal, José, et al.
Publicado: (2024)
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
por: Zaranis, Emmanouil, et al.
Publicado: (2024)
por: Zaranis, Emmanouil, et al.
Publicado: (2024)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
por: Farinhas, António, et al.
Publicado: (2025)
por: Farinhas, António, et al.
Publicado: (2025)
Analyzing Context Contributions in LLM-based Machine Translation
por: Zaranis, Emmanouil, et al.
Publicado: (2024)
por: Zaranis, Emmanouil, et al.
Publicado: (2024)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
por: Mishra, Ritwik, et al.
Publicado: (2024)
por: Mishra, Ritwik, et al.
Publicado: (2024)
Is Context Helpful for Chat Translation Evaluation?
por: Agrawal, Sweta, et al.
Publicado: (2024)
por: Agrawal, Sweta, et al.
Publicado: (2024)
QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
por: Faria, Gonçalo R. A., et al.
Publicado: (2024)
por: Faria, Gonçalo R. A., et al.
Publicado: (2024)
Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering
por: Barbu, Eduard, et al.
Publicado: (2025)
por: Barbu, Eduard, et al.
Publicado: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
por: Ramos, Miguel Moura, et al.
Publicado: (2025)
por: Ramos, Miguel Moura, et al.
Publicado: (2025)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
por: Sutawika, Lintang, et al.
Publicado: (2026)
por: Sutawika, Lintang, et al.
Publicado: (2026)
On the Calibration of Multilingual Question Answering LLMs
por: Yang, Yahan, et al.
Publicado: (2023)
por: Yang, Yahan, et al.
Publicado: (2023)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
por: Zhao, Siyan, et al.
Publicado: (2025)
por: Zhao, Siyan, et al.
Publicado: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
por: Yeo, Edward, et al.
Publicado: (2025)
por: Yeo, Edward, et al.
Publicado: (2025)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
por: Nyandwi, Jean de Dieu, et al.
Publicado: (2025)
por: Nyandwi, Jean de Dieu, et al.
Publicado: (2025)
Synthetic Multimodal Question Generation
por: Wu, Ian, et al.
Publicado: (2024)
por: Wu, Ian, et al.
Publicado: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
por: Kolawole, Steven, et al.
Publicado: (2024)
por: Kolawole, Steven, et al.
Publicado: (2024)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
por: Bakman, Yavuz, et al.
Publicado: (2025)
por: Bakman, Yavuz, et al.
Publicado: (2025)
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
por: Rohera, Pritika, et al.
Publicado: (2024)
por: Rohera, Pritika, et al.
Publicado: (2024)
Localizing Paragraph Memorization in Language Models
por: Stoehr, Niklas, et al.
Publicado: (2024)
por: Stoehr, Niklas, et al.
Publicado: (2024)
Can Automatic Metrics Assess High-Quality Translations?
por: Agrawal, Sweta, et al.
Publicado: (2024)
por: Agrawal, Sweta, et al.
Publicado: (2024)
Retrieval Augmented Question Answering: When Should LLMs Admit Ignorance?
por: Wang, Dingmin, et al.
Publicado: (2025)
por: Wang, Dingmin, et al.
Publicado: (2025)
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
por: Akpinar, Nil-Jana, et al.
Publicado: (2025)
por: Akpinar, Nil-Jana, et al.
Publicado: (2025)
Structured RAG for Answering Aggregative Questions
por: Koshorek, Omri, et al.
Publicado: (2025)
por: Koshorek, Omri, et al.
Publicado: (2025)
On Mechanistic Circuits for Extractive Question-Answering
por: Basu, Samyadeep, et al.
Publicado: (2025)
por: Basu, Samyadeep, et al.
Publicado: (2025)
ESQA: Event Sequences Question Answering
por: Abdullaeva, Irina, et al.
Publicado: (2024)
por: Abdullaeva, Irina, et al.
Publicado: (2024)
Comprehensive Modeling and Question Answering of Cancer Clinical Practice Guidelines using LLMs
por: Gupta, Bhumika, et al.
Publicado: (2025)
por: Gupta, Bhumika, et al.
Publicado: (2025)
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
por: Ramos, Miguel Moura, et al.
Publicado: (2024)
por: Ramos, Miguel Moura, et al.
Publicado: (2024)
Interpretable LLM-based Table Question Answering
por: Nguyen, Giang, et al.
Publicado: (2024)
por: Nguyen, Giang, et al.
Publicado: (2024)
Explainable Fact-checking through Question Answering
por: Yang, Jing, et al.
Publicado: (2021)
por: Yang, Jing, et al.
Publicado: (2021)
Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
por: Agrawal, Sweta, et al.
Publicado: (2023)
por: Agrawal, Sweta, et al.
Publicado: (2023)
Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
por: Stachura, Damian, et al.
Publicado: (2025)
por: Stachura, Damian, et al.
Publicado: (2025)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
por: Gor, Maharshi, et al.
Publicado: (2024)
por: Gor, Maharshi, et al.
Publicado: (2024)
OWLViz: An Open-World Benchmark for Visual Question Answering
por: Nguyen, Thuy, et al.
Publicado: (2025)
por: Nguyen, Thuy, et al.
Publicado: (2025)
Calibrated Large Language Models for Binary Question Answering
por: Giovannotti, Patrizio, et al.
Publicado: (2024)
por: Giovannotti, Patrizio, et al.
Publicado: (2024)
FoQA: A Faroese Question-Answering Dataset
por: Simonsen, Annika, et al.
Publicado: (2025)
por: Simonsen, Annika, et al.
Publicado: (2025)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
por: Wang, Zhanliang, et al.
Publicado: (2026)
por: Wang, Zhanliang, et al.
Publicado: (2026)
Towards Unsupervised Question Answering System with Multi-level Summarization for Legal Text
por: Prabhu, M Manvith, et al.
Publicado: (2024)
por: Prabhu, M Manvith, et al.
Publicado: (2024)
Interpretable Question Answering with Knowledge Graphs
por: Aneja, Kartikeya, et al.
Publicado: (2025)
por: Aneja, Kartikeya, et al.
Publicado: (2025)
Repetition Improves Language Model Embeddings
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
Ejemplares similares
-
A Context-aware Framework for Translation-mediated Conversations
por: Pombal, José, et al.
Publicado: (2024) -
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
por: Zaranis, Emmanouil, et al.
Publicado: (2024) -
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
por: Farinhas, António, et al.
Publicado: (2025) -
Analyzing Context Contributions in LLM-based Machine Translation
por: Zaranis, Emmanouil, et al.
Publicado: (2024) -
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
por: Mishra, Ritwik, et al.
Publicado: (2024)