JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Leonard, Lensenmayer, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
by: Rusli, Andre, et al.
Published: (2024)
by: Rusli, Andre, et al.
Published: (2024)
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model
by: Hennara, Khalil, et al.
Published: (2025)
by: Hennara, Khalil, et al.
Published: (2025)
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study
by: Yamagishi, Yosuke, et al.
Published: (2026)
by: Yamagishi, Yosuke, et al.
Published: (2026)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
by: Wang, Shouju, et al.
Published: (2026)
by: Wang, Shouju, et al.
Published: (2026)
LITERA: An LLM Based Approach to Latin-to-English Translation
by: Rosu, Paul
Published: (2025)
by: Rosu, Paul
Published: (2025)
The Role of Handling Attributive Nouns in Improving Chinese-To-English Machine Translation
by: Wang, Lisa, et al.
Published: (2024)
by: Wang, Lisa, et al.
Published: (2024)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
by: Jourdan, Fanny, et al.
Published: (2025)
by: Jourdan, Fanny, et al.
Published: (2025)
SemBench: A Universal Semantic Framework for LLM Evaluation
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
by: Zhou, Xiaolin, et al.
Published: (2026)
by: Zhou, Xiaolin, et al.
Published: (2026)
Sentiment Analysis Across Languages: Evaluation Before and After Machine Translation to English
by: Kathunia, Aekansh, et al.
Published: (2024)
by: Kathunia, Aekansh, et al.
Published: (2024)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
by: Liu, Junyu, et al.
Published: (2026)
by: Liu, Junyu, et al.
Published: (2026)
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
by: Luo, Shunyang, et al.
Published: (2026)
by: Luo, Shunyang, et al.
Published: (2026)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution
by: Lin, Leon, et al.
Published: (2025)
by: Lin, Leon, et al.
Published: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
by: Tan, Haoran, et al.
Published: (2025)
by: Tan, Haoran, et al.
Published: (2025)
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
by: Joshi, Nisheeth, et al.
Published: (2025)
by: Joshi, Nisheeth, et al.
Published: (2025)
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
by: Wang, Xinda, et al.
Published: (2025)
by: Wang, Xinda, et al.
Published: (2025)
ChiEngMixBench: Evaluating Large Language Models on Spontaneous and Natural Chinese-English Code-Mixed Generation
by: Yang, Qingyan, et al.
Published: (2026)
by: Yang, Qingyan, et al.
Published: (2026)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair
by: Tran, Hong-Viet, et al.
Published: (2025)
by: Tran, Hong-Viet, et al.
Published: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
by: Nguyen, Bang, et al.
Published: (2026)
by: Nguyen, Bang, et al.
Published: (2026)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
by: Schmidt, Jan-Philipp
Published: (2026)
by: Schmidt, Jan-Philipp
Published: (2026)
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods
by: Yue, Richard, et al.
Published: (2024)
by: Yue, Richard, et al.
Published: (2024)
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
by: Sandan, Isik Baran, et al.
Published: (2025)
by: Sandan, Isik Baran, et al.
Published: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
OckBench: Measuring the Efficiency of LLM Reasoning
by: Du, Zheng, et al.
Published: (2025)
by: Du, Zheng, et al.
Published: (2025)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
by: Zhao, Junjie, et al.
Published: (2026)
by: Zhao, Junjie, et al.
Published: (2026)
JETHICS: Japanese Ethics Understanding Evaluation Dataset
by: Takeshita, Masashi, et al.
Published: (2025)
by: Takeshita, Masashi, et al.
Published: (2025)
LLM Prompt Evaluation for Educational Applications
by: Holmes, Langdon, et al.
Published: (2026)
by: Holmes, Langdon, et al.
Published: (2026)
TreePrompt: Leveraging Hierarchical Few-Shot Example Selection for Improved English-Persian and English-German Translation
by: Kakavand, Ramtin, et al.
Published: (2025)
by: Kakavand, Ramtin, et al.
Published: (2025)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
by: Xiao, Jianfei, et al.
Published: (2026)
by: Xiao, Jianfei, et al.
Published: (2026)
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
by: Rosenbaum, Andy, et al.
Published: (2026)
by: Rosenbaum, Andy, et al.
Published: (2026)
QueEn: A Large Language Model for Quechua-English Translation
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
Improving LLM Abilities in Idiomatic Translation
by: Donthi, Sundesh, et al.
Published: (2024)
by: Donthi, Sundesh, et al.
Published: (2024)
Similar Items
-
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
by: Rusli, Andre, et al.
Published: (2024) -
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model
by: Hennara, Khalil, et al.
Published: (2025) -
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study
by: Yamagishi, Yosuke, et al.
Published: (2026) -
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
by: Wang, Shouju, et al.
Published: (2026) -
LITERA: An LLM Based Approach to Latin-to-English Translation
by: Rosu, Paul
Published: (2025)