Evaluating LLMs on Chinese Idiom Translation
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Cai, Dou, Yao, Heineman, David, Wu, Xiaofeng, Xu, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Minimum Bayes Risk Decoding with Multi-Prompt
por: Heineman, David, et al.
Publicado: (2024)
por: Heineman, David, et al.
Publicado: (2024)
Evaluating Large Language Models on Urdu Idiom Translation
por: Khan, Muhammad Farmal, et al.
Publicado: (2025)
por: Khan, Muhammad Farmal, et al.
Publicado: (2025)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
por: Dou, Yao, et al.
Publicado: (2026)
por: Dou, Yao, et al.
Publicado: (2026)
Creative and Context-Aware Translation of East Asian Idioms with GPT-4
por: Tang, Kenan, et al.
Publicado: (2024)
por: Tang, Kenan, et al.
Publicado: (2024)
The Impact of Visual Information in Chinese Characters: Evaluating Large Models' Ability to Recognize and Utilize Radicals
por: Wu, Xiaofeng, et al.
Publicado: (2024)
por: Wu, Xiaofeng, et al.
Publicado: (2024)
It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems
por: Zaitova, Iuliia, et al.
Publicado: (2025)
por: Zaitova, Iuliia, et al.
Publicado: (2025)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
por: Kim, Jisu, et al.
Publicado: (2025)
por: Kim, Jisu, et al.
Publicado: (2025)
Readability-guided Idiom-aware Sentence Simplification (RISS) for Chinese
por: Zhang, Jingshen, et al.
Publicado: (2024)
por: Zhang, Jingshen, et al.
Publicado: (2024)
Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use
por: Fu, Yicheng, et al.
Publicado: (2025)
por: Fu, Yicheng, et al.
Publicado: (2025)
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
por: Rezaeimanesh, Sara, et al.
Publicado: (2024)
por: Rezaeimanesh, Sara, et al.
Publicado: (2024)
Idiom Understanding as a Tool to Measure the Dialect Gap
por: Beauchemin, David, et al.
Publicado: (2025)
por: Beauchemin, David, et al.
Publicado: (2025)
Towards a Path Dependent Account of Category Fluency
por: Heineman, David, et al.
Publicado: (2024)
por: Heineman, David, et al.
Publicado: (2024)
A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality
por: Agarwal, Ishika, et al.
Publicado: (2026)
por: Agarwal, Ishika, et al.
Publicado: (2026)
Idiom Detection in Sorani Kurdish Texts
por: Omer, Skala Kamaran, et al.
Publicado: (2025)
por: Omer, Skala Kamaran, et al.
Publicado: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
por: Zhang, Ran, et al.
Publicado: (2024)
por: Zhang, Ran, et al.
Publicado: (2024)
DualCoTs: Dual Chain-of-Thoughts Prompting for Sentiment Lexicon Expansion of Idioms
por: Niu, Fuqiang, et al.
Publicado: (2024)
por: Niu, Fuqiang, et al.
Publicado: (2024)
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
por: Wu, Xiaofeng, et al.
Publicado: (2025)
por: Wu, Xiaofeng, et al.
Publicado: (2025)
NLP Datasets for Idiom and Figurative Language Tasks
por: Matheny, Blake, et al.
Publicado: (2025)
por: Matheny, Blake, et al.
Publicado: (2025)
DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation
por: Man, Zhibo, et al.
Publicado: (2025)
por: Man, Zhibo, et al.
Publicado: (2025)
Killing Two Flies with One Stone: An Attempt to Break LLMs Using English->Icelandic Idioms and Proper Names
por: Ármannsson, Bjarki, et al.
Publicado: (2024)
por: Ármannsson, Bjarki, et al.
Publicado: (2024)
A Survey of Idiom Datasets for Psycholinguistic and Computational Research
por: Flor, Michael, et al.
Publicado: (2025)
por: Flor, Michael, et al.
Publicado: (2025)
ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs
por: Zheng, Xiang, et al.
Publicado: (2026)
por: Zheng, Xiang, et al.
Publicado: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
por: Cai, Yunna, et al.
Publicado: (2025)
por: Cai, Yunna, et al.
Publicado: (2025)
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
por: Gao, Lirong, et al.
Publicado: (2026)
por: Gao, Lirong, et al.
Publicado: (2026)
Machine Translation Evaluation Benchmark for Wu Chinese: Workflow and Analysis
por: Yu, Hongjian, et al.
Publicado: (2024)
por: Yu, Hongjian, et al.
Publicado: (2024)
AC-EVAL: Evaluating Ancient Chinese Language Understanding in Large Language Models
por: Wei, Yuting, et al.
Publicado: (2024)
por: Wei, Yuting, et al.
Publicado: (2024)
Comparative Study of Multilingual Idioms and Similes in Large Language Models
por: Khoshtab, Paria, et al.
Publicado: (2024)
por: Khoshtab, Paria, et al.
Publicado: (2024)
Benchmarking Machine Translation on Chinese Social Media Texts
por: Zhao, Kaiyan, et al.
Publicado: (2026)
por: Zhao, Kaiyan, et al.
Publicado: (2026)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
por: Wei, Qikai, et al.
Publicado: (2024)
por: Wei, Qikai, et al.
Publicado: (2024)
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving
por: Chen, Andong, et al.
Publicado: (2024)
por: Chen, Andong, et al.
Publicado: (2024)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
por: Sun, Jiaxing, et al.
Publicado: (2024)
por: Sun, Jiaxing, et al.
Publicado: (2024)
General2Specialized LLMs Translation for E-commerce
por: Chen, Kaidi, et al.
Publicado: (2024)
por: Chen, Kaidi, et al.
Publicado: (2024)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
por: Xiao, Kelaiti, et al.
Publicado: (2025)
por: Xiao, Kelaiti, et al.
Publicado: (2025)
Evaluating LLMs on Chinese Topic Constructions: A Research Proposal Inspired by Tian et al. (2024)
por: Yang, Xiaodong
Publicado: (2025)
por: Yang, Xiaodong
Publicado: (2025)
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
por: Dou, Yao, et al.
Publicado: (2025)
por: Dou, Yao, et al.
Publicado: (2025)
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
por: Jiang, Zhenhui, et al.
Publicado: (2024)
por: Jiang, Zhenhui, et al.
Publicado: (2024)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
por: Wang, Shanshan, et al.
Publicado: (2025)
por: Wang, Shanshan, et al.
Publicado: (2025)
When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
por: Sakhawat, Adib, et al.
Publicado: (2026)
por: Sakhawat, Adib, et al.
Publicado: (2026)
A Hard Nut to Crack: Idiom Detection with Conversational Large Language Models
por: Fornaciari, Francesca De Luca, et al.
Publicado: (2024)
por: Fornaciari, Francesca De Luca, et al.
Publicado: (2024)
Are LLMs Effective Backbones for Fine-tuning? An Experimental Investigation of Supervised LLMs on Chinese Short Text Matching
por: Liu, Shulin, et al.
Publicado: (2024)
por: Liu, Shulin, et al.
Publicado: (2024)
Ejemplares similares
-
Improving Minimum Bayes Risk Decoding with Multi-Prompt
por: Heineman, David, et al.
Publicado: (2024) -
Evaluating Large Language Models on Urdu Idiom Translation
por: Khan, Muhammad Farmal, et al.
Publicado: (2025) -
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
por: Dou, Yao, et al.
Publicado: (2026) -
Creative and Context-Aware Translation of East Asian Idioms with GPT-4
por: Tang, Kenan, et al.
Publicado: (2024) -
The Impact of Visual Information in Chinese Characters: Evaluating Large Models' Ability to Recognize and Utilize Radicals
por: Wu, Xiaofeng, et al.
Publicado: (2024)