A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls
Fuente:
arXiv
Saved in:
| Main Authors: | Shafayat, Sheikh, Yoon, Dongkeun, Jang, Woori, Choi, Jiwoo, Oh, Alice, Jung, Seohyon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Style over Story: Measuring LLM Narrative Preferences via Structured Selection
by: Jung, Donghoon, et al.
Published: (2025)
by: Jung, Donghoon, et al.
Published: (2025)
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
by: Choi, Jiwoo, et al.
Published: (2026)
by: Choi, Jiwoo, et al.
Published: (2026)
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024)
by: Yoon, Dongkeun, et al.
Published: (2024)
Narrative Landscape: Mapping Narrative Dispositions Across LLMs
by: Jung, Donghoon, et al.
Published: (2026)
by: Jung, Donghoon, et al.
Published: (2026)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
by: Kabir, Daeen, et al.
Published: (2025)
by: Kabir, Daeen, et al.
Published: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
A Single Model Ensemble Framework for Neural Machine Translation using Pivot Translation
by: Oh, Seokjin, et al.
Published: (2025)
by: Oh, Seokjin, et al.
Published: (2025)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
Large Language Models for Stemming: Promises, Pitfalls and Failures
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
by: Gerrits, Kyo, et al.
Published: (2026)
by: Gerrits, Kyo, et al.
Published: (2026)
Fluency and Faithfulness in Human and Machine Literary Translation
by: Griebel, Sarah, et al.
Published: (2026)
by: Griebel, Sarah, et al.
Published: (2026)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
by: Xu, Xi, et al.
Published: (2024)
by: Xu, Xi, et al.
Published: (2024)
The Promises and Pitfalls of Using Language Models to Measure Instruction Quality in Education
by: Xu, Paiheng, et al.
Published: (2024)
by: Xu, Paiheng, et al.
Published: (2024)
Evaluating the Consistency of LLM Evaluators
by: Lee, Noah, et al.
Published: (2024)
by: Lee, Noah, et al.
Published: (2024)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
by: Horych, Tomas, et al.
Published: (2024)
by: Horych, Tomas, et al.
Published: (2024)
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
by: Zhang, Ran, et al.
Published: (2026)
by: Zhang, Ran, et al.
Published: (2026)
Multiple References with Meaningful Variations Improve Literary Machine Translation
by: Wu, Si, et al.
Published: (2024)
by: Wu, Si, et al.
Published: (2024)
Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation
by: Ploeger, Esther, et al.
Published: (2024)
by: Ploeger, Esther, et al.
Published: (2024)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
by: Hwang, Yerin, et al.
Published: (2024)
by: Hwang, Yerin, et al.
Published: (2024)
Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation
by: Wang, Longyue, et al.
Published: (2024)
by: Wang, Longyue, et al.
Published: (2024)
Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams
by: Lee, Yukyung, et al.
Published: (2026)
by: Lee, Yukyung, et al.
Published: (2026)
Pitfalls of Evaluating Language Models with Open Benchmarks
by: Hasan, Md. Najib, et al.
Published: (2025)
by: Hasan, Md. Najib, et al.
Published: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
Translation Entropy: A Statistical Framework for Evaluating Translation Systems
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Extending CREAMT: Leveraging Large Language Models for Literary Translation Post-Editing
by: Castaldo, Antonio, et al.
Published: (2025)
by: Castaldo, Antonio, et al.
Published: (2025)
MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment
by: Kim, Junghwan, et al.
Published: (2025)
by: Kim, Junghwan, et al.
Published: (2025)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
by: Chan, Jason, et al.
Published: (2025)
by: Chan, Jason, et al.
Published: (2025)
Conditional Unbalanced Optimal Transport Maps: An Outlier-Robust Framework for Conditional Generative Modeling
by: Yoon, Jiwoo, et al.
Published: (2026)
by: Yoon, Jiwoo, et al.
Published: (2026)
Reasoning Models Better Express Their Confidence
by: Yoon, Dongkeun, et al.
Published: (2025)
by: Yoon, Dongkeun, et al.
Published: (2025)
SAMAS: A Spectrum-Guided Multi-Agent System for Achieving Style Fidelity in Literary Translation
by: Wu, Jingzhuo, et al.
Published: (2026)
by: Wu, Jingzhuo, et al.
Published: (2026)
Similar Items
-
Style over Story: Measuring LLM Narrative Preferences via Structured Selection
by: Jung, Donghoon, et al.
Published: (2025) -
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
by: Choi, Jiwoo, et al.
Published: (2026) -
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024) -
Narrative Landscape: Mapping Narrative Dispositions Across LLMs
by: Jung, Donghoon, et al.
Published: (2026) -
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024)