How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Ran, Zhao, Wei, Eger, Steffen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering
por: Zhang, Ran, et al.
Publicado: (2025)
por: Zhang, Ran, et al.
Publicado: (2025)
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
por: Zhang, Ran, et al.
Publicado: (2026)
por: Zhang, Ran, et al.
Publicado: (2026)
How to Evaluate Coreference in Literary Texts?
por: Duron-Tejedor, Ana-Isabel, et al.
Publicado: (2023)
por: Duron-Tejedor, Ana-Isabel, et al.
Publicado: (2023)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
por: Nadas, Mihai, et al.
Publicado: (2025)
por: Nadas, Mihai, et al.
Publicado: (2025)
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
por: Leiter, Christoph, et al.
Publicado: (2024)
por: Leiter, Christoph, et al.
Publicado: (2024)
General2Specialized LLMs Translation for E-commerce
por: Chen, Kaidi, et al.
Publicado: (2024)
por: Chen, Kaidi, et al.
Publicado: (2024)
Fluency and Faithfulness in Human and Machine Literary Translation
por: Griebel, Sarah, et al.
Publicado: (2026)
por: Griebel, Sarah, et al.
Publicado: (2026)
Automatic Translation Alignment Pipeline for Multilingual Digital Editions of Literary Works
por: Levchenko, Maria
Publicado: (2024)
por: Levchenko, Maria
Publicado: (2024)
Do LLMs Really Struggle at NL-FOL Translation? Revealing their Strengths via a Novel Benchmarking Strategy
por: Brunello, Andrea, et al.
Publicado: (2025)
por: Brunello, Andrea, et al.
Publicado: (2025)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
por: Bonomo, Tommaso, et al.
Publicado: (2025)
por: Bonomo, Tommaso, et al.
Publicado: (2025)
Fake Alignment: Are LLMs Really Aligned Well?
por: Wang, Yixu, et al.
Publicado: (2023)
por: Wang, Yixu, et al.
Publicado: (2023)
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
por: Wu, Xinwei, et al.
Publicado: (2026)
por: Wu, Xinwei, et al.
Publicado: (2026)
Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect
por: Presa, Joao Paulo Cavalcante, et al.
Publicado: (2026)
por: Presa, Joao Paulo Cavalcante, et al.
Publicado: (2026)
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
por: Gerrits, Kyo, et al.
Publicado: (2026)
por: Gerrits, Kyo, et al.
Publicado: (2026)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
por: Koneru, Sai, et al.
Publicado: (2024)
por: Koneru, Sai, et al.
Publicado: (2024)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
por: Rei, Ricardo, et al.
Publicado: (2025)
por: Rei, Ricardo, et al.
Publicado: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
por: Wang, Yicheng, et al.
Publicado: (2024)
por: Wang, Yicheng, et al.
Publicado: (2024)
LitVISTA: A Benchmark for Narrative Orchestration in Literary Text
por: Lu, Mingzhe, et al.
Publicado: (2026)
por: Lu, Mingzhe, et al.
Publicado: (2026)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
por: Zhan, Runzhe, et al.
Publicado: (2025)
por: Zhan, Runzhe, et al.
Publicado: (2025)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
por: Belouadi, Jonas, et al.
Publicado: (2022)
por: Belouadi, Jonas, et al.
Publicado: (2022)
A Perspective on Literary Metaphor in the Context of Generative AI
por: van Heerden, Imke, et al.
Publicado: (2024)
por: van Heerden, Imke, et al.
Publicado: (2024)
Transcending Language Boundaries: Harnessing LLMs for Low-Resource Language Translation
por: Shu, Peng, et al.
Publicado: (2024)
por: Shu, Peng, et al.
Publicado: (2024)
SparQLe: Speech Queries to Text Translation Through LLMs
por: Djanibekov, Amirbek, et al.
Publicado: (2025)
por: Djanibekov, Amirbek, et al.
Publicado: (2025)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
por: Gogoulou, Evangelia, et al.
Publicado: (2025)
por: Gogoulou, Evangelia, et al.
Publicado: (2025)
Translating Under Pressure: Domain-Aware LLMs for Crisis Communication
por: Castaldo, Antonio, et al.
Publicado: (2026)
por: Castaldo, Antonio, et al.
Publicado: (2026)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
por: Wang, Xintao, et al.
Publicado: (2025)
por: Wang, Xintao, et al.
Publicado: (2025)
Is there really a Citation Age Bias in NLP?
por: Nguyen, Hoa, et al.
Publicado: (2024)
por: Nguyen, Hoa, et al.
Publicado: (2024)
Large Language Models "Ad Referendum": How Good Are They at Machine Translation in the Legal Domain?
por: Briva-Iglesias, Vicent, et al.
Publicado: (2024)
por: Briva-Iglesias, Vicent, et al.
Publicado: (2024)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
por: Zhang, Leixin, et al.
Publicado: (2024)
por: Zhang, Leixin, et al.
Publicado: (2024)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
por: Choenni, Rochelle, et al.
Publicado: (2024)
por: Choenni, Rochelle, et al.
Publicado: (2024)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
por: Papi, Sara, et al.
Publicado: (2025)
por: Papi, Sara, et al.
Publicado: (2025)
Long Story Generation via Knowledge Graph and Literary Theory
por: Shi, Ge, et al.
Publicado: (2025)
por: Shi, Ge, et al.
Publicado: (2025)
Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
por: Xu, Fangzhi, et al.
Publicado: (2023)
por: Xu, Fangzhi, et al.
Publicado: (2023)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
por: Larionov, Daniil, et al.
Publicado: (2025)
por: Larionov, Daniil, et al.
Publicado: (2025)
TextQuests: How Good are LLMs at Text-Based Video Games?
por: Phan, Long, et al.
Publicado: (2025)
por: Phan, Long, et al.
Publicado: (2025)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
por: Greisinger, Christian, et al.
Publicado: (2026)
por: Greisinger, Christian, et al.
Publicado: (2026)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
por: Zan, Changtong, et al.
Publicado: (2024)
por: Zan, Changtong, et al.
Publicado: (2024)
Fine-Tuning LLMs for Low-Resource Dialect Translation: The Case of Lebanese
por: Yakhni, Silvana, et al.
Publicado: (2025)
por: Yakhni, Silvana, et al.
Publicado: (2025)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
por: Halder, Deepon, et al.
Publicado: (2025)
por: Halder, Deepon, et al.
Publicado: (2025)
Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs
por: English, William, et al.
Publicado: (2025)
por: English, William, et al.
Publicado: (2025)
Ejemplares similares
-
LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering
por: Zhang, Ran, et al.
Publicado: (2025) -
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
por: Zhang, Ran, et al.
Publicado: (2026) -
How to Evaluate Coreference in Literary Texts?
por: Duron-Tejedor, Ana-Isabel, et al.
Publicado: (2023) -
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
por: Nadas, Mihai, et al.
Publicado: (2025) -
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
por: Leiter, Christoph, et al.
Publicado: (2024)