GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Jianhao, Yan, Pingchuan, Chen, Yulong, Li, Judy, Zhu, Xianchao, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
What Have We Achieved on Non-autoregressive Translation?
von: Li, Yafu, et al.
Veröffentlicht: (2024)
von: Li, Yafu, et al.
Veröffentlicht: (2024)
Gradable ChatGPT Translation Evaluation
von: Jiao, Hui, et al.
Veröffentlicht: (2024)
von: Jiao, Hui, et al.
Veröffentlicht: (2024)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
Creative and Context-Aware Translation of East Asian Idioms with GPT-4
von: Tang, Kenan, et al.
Veröffentlicht: (2024)
von: Tang, Kenan, et al.
Veröffentlicht: (2024)
How Well Do Large Reasoning Models Translate? A Comprehensive Evaluation for Multi-Domain Machine Translation
von: Ye, Yongshi, et al.
Veröffentlicht: (2025)
von: Ye, Yongshi, et al.
Veröffentlicht: (2025)
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation
von: Luo, Jiaming, et al.
Veröffentlicht: (2024)
von: Luo, Jiaming, et al.
Veröffentlicht: (2024)
Quality Estimation Reranking for Document-Level Translation
von: Mrozinski, Krzysztof, et al.
Veröffentlicht: (2025)
von: Mrozinski, Krzysztof, et al.
Veröffentlicht: (2025)
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
von: Jiang, Zhaokun, et al.
Veröffentlicht: (2024)
von: Jiang, Zhaokun, et al.
Veröffentlicht: (2024)
Evaluating the Translation Performance of Large Language Models Based on Euas-20
von: Huang, Yan, et al.
Veröffentlicht: (2024)
von: Huang, Yan, et al.
Veröffentlicht: (2024)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
von: Weigang, Li, et al.
Veröffentlicht: (2025)
von: Weigang, Li, et al.
Veröffentlicht: (2025)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
von: Wang, Shun, et al.
Veröffentlicht: (2024)
von: Wang, Shun, et al.
Veröffentlicht: (2024)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
Chapter Translation complexities
von: Yue, Yan
Veröffentlicht: (2026)
von: Yue, Yan
Veröffentlicht: (2026)
LLM4FS: Leveraging Large Language Models for Feature Selection
von: Li, Jianhao, et al.
Veröffentlicht: (2025)
von: Li, Jianhao, et al.
Veröffentlicht: (2025)
Distinguishing Translations by Human, NMT, and ChatGPT: A Linguistic and Statistical Approach
von: Jiang, Zhaokun, et al.
Veröffentlicht: (2023)
von: Jiang, Zhaokun, et al.
Veröffentlicht: (2023)
Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models
von: Mammen, Priyanka Mary, et al.
Veröffentlicht: (2026)
von: Mammen, Priyanka Mary, et al.
Veröffentlicht: (2026)
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2025)
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2025)
DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains
von: Zhao, Xiying, et al.
Veröffentlicht: (2025)
von: Zhao, Xiying, et al.
Veröffentlicht: (2025)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
von: Liu, Zhongtao, et al.
Veröffentlicht: (2024)
von: Liu, Zhongtao, et al.
Veröffentlicht: (2024)
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
von: Chen, Andong, et al.
Veröffentlicht: (2025)
von: Chen, Andong, et al.
Veröffentlicht: (2025)
Questionnaires for Everyone: Streamlining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation
von: Haavisto, Otso, et al.
Veröffentlicht: (2024)
von: Haavisto, Otso, et al.
Veröffentlicht: (2024)
Can Uniform Meaning Representation Help GPT-4 Translate from Indigenous Languages?
von: Wein, Shira
Veröffentlicht: (2025)
von: Wein, Shira
Veröffentlicht: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
Spatial Speech Translation: Translating Across Space With Binaural Hearables
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Sentiment Analysis Across Languages: Evaluation Before and After Machine Translation to English
von: Kathunia, Aekansh, et al.
Veröffentlicht: (2024)
von: Kathunia, Aekansh, et al.
Veröffentlicht: (2024)
Text Understanding in GPT-4 vs Humans
von: Shultz, Thomas R., et al.
Veröffentlicht: (2024)
von: Shultz, Thomas R., et al.
Veröffentlicht: (2024)
Reconsidering Sentence-Level Sign Language Translation
von: Tanzer, Garrett, et al.
Veröffentlicht: (2024)
von: Tanzer, Garrett, et al.
Veröffentlicht: (2024)
Span-Level Machine Translation Meta-Evaluation
von: Perrella, Stefano, et al.
Veröffentlicht: (2026)
von: Perrella, Stefano, et al.
Veröffentlicht: (2026)
AI-Assisted Human Evaluation of Machine Translation
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2026)
Non-Linear Scoring Model for Translation Quality Evaluation
von: Gladkoff, Serge, et al.
Veröffentlicht: (2025)
von: Gladkoff, Serge, et al.
Veröffentlicht: (2025)
ROC Analysis for Evaluating Translation Quality Estimation Systems
von: Garland, Evelyn Y., et al.
Veröffentlicht: (2026)
von: Garland, Evelyn Y., et al.
Veröffentlicht: (2026)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
Translate-and-Revise: Boosting Large Language Models for Constrained Translation
von: Huang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Huang, Pengcheng, et al.
Veröffentlicht: (2024)
Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
Flexible and Adaptable Summarization via Expertise Separation
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
Consensus-Aligned Neuron Efficient Fine-Tuning Large Language Models for Multi-Domain Machine Translation
von: Jiang, Shuting, et al.
Veröffentlicht: (2026)
von: Jiang, Shuting, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
von: Yan, Jianhao, et al.
Veröffentlicht: (2024) -
What Have We Achieved on Non-autoregressive Translation?
von: Li, Yafu, et al.
Veröffentlicht: (2024) -
Gradable ChatGPT Translation Evaluation
von: Jiao, Hui, et al.
Veröffentlicht: (2024) -
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
von: Yan, Jianhao, et al.
Veröffentlicht: (2024) -
Creative and Context-Aware Translation of East Asian Idioms with GPT-4
von: Tang, Kenan, et al.
Veröffentlicht: (2024)