Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yanzhi, Wang, Cunxiang, Liu, Zeming, Huang, Heyan, Yu, Wenbo, Song, Dawei, Tang, Jie, Guo, Yuhang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRIM: Towards Practical In-Image Multilingual Machine Translation
by: Tian, Yanzhi, et al.
Published: (2025)
by: Tian, Yanzhi, et al.
Published: (2025)
Deterministic Reversible Data Augmentation for Neural Machine Translation
by: Yao, Jiashu, et al.
Published: (2024)
by: Yao, Jiashu, et al.
Published: (2024)
Exploring In-Image Machine Translation with Real-World Background
by: Tian, Yanzhi, et al.
Published: (2025)
by: Tian, Yanzhi, et al.
Published: (2025)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
by: Yao, Jiashu, et al.
Published: (2026)
by: Yao, Jiashu, et al.
Published: (2026)
Literate Execution
by: Bond, Joe, et al.
Published: (2026)
by: Bond, Joe, et al.
Published: (2026)
Literate Tracing
by: Sotoudeh, Matthew
Published: (2025)
by: Sotoudeh, Matthew
Published: (2025)
NoLiMa: Long-Context Evaluation Beyond Literal Matching
by: Modarressi, Ali, et al.
Published: (2025)
by: Modarressi, Ali, et al.
Published: (2025)
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Beyond Literal Token Overlap: Token Alignability for Multilinguality
by: Hämmerl, Katharina, et al.
Published: (2025)
by: Hämmerl, Katharina, et al.
Published: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
by: Gao, Siyuan, et al.
Published: (2025)
by: Gao, Siyuan, et al.
Published: (2025)
DocMEdit: Towards Document-Level Model Editing
by: Zeng, Li, et al.
Published: (2025)
by: Zeng, Li, et al.
Published: (2025)
IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions
by: Hashiloni, Kai Golan, et al.
Published: (2026)
by: Hashiloni, Kai Golan, et al.
Published: (2026)
KOSMOS-2.5: A Multimodal Literate Model
by: Lv, Tengchao, et al.
Published: (2023)
by: Lv, Tengchao, et al.
Published: (2023)
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
by: Cheshmi, Seyyed Saeid, et al.
Published: (2026)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2026)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
by: Yerukola, Akhila, et al.
Published: (2024)
by: Yerukola, Akhila, et al.
Published: (2024)
Writing Skill, Literates and Social Apps
by: Mohd Muzamil Sohil
Published: (2017)
by: Mohd Muzamil Sohil
Published: (2017)
Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
by: Li, Yafu, et al.
Published: (2025)
by: Li, Yafu, et al.
Published: (2025)
Pruning Literals for Highly Efficient Explainability at Word Level
by: Yadav, Rohan Kumar, et al.
Published: (2024)
by: Yadav, Rohan Kumar, et al.
Published: (2024)
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks
by: Yao, Jiashu, et al.
Published: (2024)
by: Yao, Jiashu, et al.
Published: (2024)
Grammar-Aware Literate Generative Mathematical Programming with Compiler-in-the-Loop
by: Rossi, Roberto, et al.
Published: (2026)
by: Rossi, Roberto, et al.
Published: (2026)
Literator
Published: (2013)
Published: (2013)
The Sparse Tsetlin Machine: Sparse Representation with Active Literals
by: Østby, Sebastian, et al.
Published: (2024)
by: Østby, Sebastian, et al.
Published: (2024)
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
by: Das, Sarmistha, et al.
Published: (2026)
by: Das, Sarmistha, et al.
Published: (2026)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
by: Yao, Jiashu, et al.
Published: (2026)
by: Yao, Jiashu, et al.
Published: (2026)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
by: Fayyaz, Mohsen, et al.
Published: (2025)
by: Fayyaz, Mohsen, et al.
Published: (2025)
Natural Language Outlines for Code: Literate Programming in the LLM Era
by: Shi, Kensen, et al.
Published: (2024)
by: Shi, Kensen, et al.
Published: (2024)
Provision of Library Service to Non Literates.
by: Ogunsheye, F. Adetowun
Published: (1980)
by: Ogunsheye, F. Adetowun
Published: (1980)
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
by: Qiu, Yuli, et al.
Published: (2024)
by: Qiu, Yuli, et al.
Published: (2024)
Learning Chain Of Thoughts Prompts for Predicting Entities, Relations, and even Literals on Knowledge Graphs
by: Baci, Alkid, et al.
Published: (2026)
by: Baci, Alkid, et al.
Published: (2026)
Floating or Suggesting Ideas? A Large-Scale Contrastive Analysis of Metaphorical and Literal Verb-Object Constructions
by: Piccirilli, Prisca, et al.
Published: (2026)
by: Piccirilli, Prisca, et al.
Published: (2026)
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
by: Li, Silin, et al.
Published: (2025)
by: Li, Silin, et al.
Published: (2025)
SM-based Semantics for Answer Set Programs Containing Conditional Literals and Arithmetic
by: Hansen, Zachary, et al.
Published: (2025)
by: Hansen, Zachary, et al.
Published: (2025)
Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation
by: Vachhani, Bhavik, et al.
Published: (2026)
by: Vachhani, Bhavik, et al.
Published: (2026)
Metaphor and Literalism in Buddhism
by: Hwang, Soonil
Published: (2025)
by: Hwang, Soonil
Published: (2025)
Becoming Video Literate.
by: Beasley, Augie E.
Published: (1995)
by: Beasley, Augie E.
Published: (1995)
Are You Web Literate?
by: Darrow, Rob
Published: (1999)
by: Darrow, Rob
Published: (1999)
Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
by: Yao, Jiashu, et al.
Published: (2025)
by: Yao, Jiashu, et al.
Published: (2025)
How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling
by: Liu, Yuhang, et al.
Published: (2026)
by: Liu, Yuhang, et al.
Published: (2026)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
by: Yang, Gao, et al.
Published: (2025)
by: Yang, Gao, et al.
Published: (2025)
Similar Items
-
PRIM: Towards Practical In-Image Multilingual Machine Translation
by: Tian, Yanzhi, et al.
Published: (2025) -
Deterministic Reversible Data Augmentation for Neural Machine Translation
by: Yao, Jiashu, et al.
Published: (2024) -
Exploring In-Image Machine Translation with Real-World Background
by: Tian, Yanzhi, et al.
Published: (2025) -
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
by: Yao, Jiashu, et al.
Published: (2026) -
Literate Execution
by: Bond, Joe, et al.
Published: (2026)