HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zheng, Zheng, Mao, Song, Mingyang, Fei, Tianxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents
by: Fei, Tianxiang, et al.
Published: (2026)
by: Fei, Tianxiang, et al.
Published: (2026)
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
by: Xu, Chenning, et al.
Published: (2026)
by: Xu, Chenning, et al.
Published: (2026)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
A Survey of Query Optimization in Large Language Models
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
by: Sun, Mingrui, et al.
Published: (2026)
by: Sun, Mingrui, et al.
Published: (2026)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
by: Yang, Wenjie, et al.
Published: (2025)
by: Yang, Wenjie, et al.
Published: (2025)
A Survey of On-Policy Distillation for Large Language Models
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
by: Xu, Chenning, et al.
Published: (2026)
by: Xu, Chenning, et al.
Published: (2026)
HY-MT1.5 Technical Report
by: Zheng, Mao, et al.
Published: (2025)
by: Zheng, Mao, et al.
Published: (2025)
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
by: Chen, Jialin, et al.
Published: (2025)
by: Chen, Jialin, et al.
Published: (2025)
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026)
by: Yuan, Zhonghang, et al.
Published: (2026)
Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator
by: Liu, Chengyuan, et al.
Published: (2024)
by: Liu, Chengyuan, et al.
Published: (2024)
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
by: Lu, Yujing, et al.
Published: (2025)
by: Lu, Yujing, et al.
Published: (2025)
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
by: Zheng, Mao, et al.
Published: (2026)
by: Zheng, Mao, et al.
Published: (2026)
Hunyuan-MT Technical Report
by: Zheng, Mao, et al.
Published: (2025)
by: Zheng, Mao, et al.
Published: (2025)
FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning
by: Hu, Tianxiang, et al.
Published: (2024)
by: Hu, Tianxiang, et al.
Published: (2024)
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
by: Huang, Chenyang, et al.
Published: (2025)
by: Huang, Chenyang, et al.
Published: (2025)
Domain-Specific Machine Translation to Translate Medicine Brochures in English to Sorani Kurdish
by: Shamal, Mariam, et al.
Published: (2025)
by: Shamal, Mariam, et al.
Published: (2025)
Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
Semantic Prosody in Machine Translation: the English-Chinese Case of Passive Structures
by: Ma, Xinyue, et al.
Published: (2025)
by: Ma, Xinyue, et al.
Published: (2025)
StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models
by: Chen, Yongrui, et al.
Published: (2026)
by: Chen, Yongrui, et al.
Published: (2026)
Towards Better Chinese-centric Neural Machine Translation for Low-resource Languages
by: Li, Bin, et al.
Published: (2022)
by: Li, Bin, et al.
Published: (2022)
On Temperature-Constrained Non-Deterministic Machine Translation: Potential and Evaluation
by: Wang, Weichuan, et al.
Published: (2026)
by: Wang, Weichuan, et al.
Published: (2026)
The Role of Handling Attributive Nouns in Improving Chinese-To-English Machine Translation
by: Wang, Lisa, et al.
Published: (2024)
by: Wang, Lisa, et al.
Published: (2024)
Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation
by: Li, Bryan, et al.
Published: (2025)
by: Li, Bryan, et al.
Published: (2025)
Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example
by: Xia, Wei, et al.
Published: (2024)
by: Xia, Wei, et al.
Published: (2024)
StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
FFN: a Fine-grained Chinese-English Financial Domain Parallel Corpus
by: Fu, Yuxin, et al.
Published: (2024)
by: Fu, Yuxin, et al.
Published: (2024)
Translation via Annotation: A Computational Study of Translating Classical Chinese into Japanese
by: Li, Zilong, et al.
Published: (2025)
by: Li, Zilong, et al.
Published: (2025)
Probing Large Language Models in Reasoning and Translating Complex Linguistic Puzzles
by: Lin, Zheng-Lin, et al.
Published: (2025)
by: Lin, Zheng-Lin, et al.
Published: (2025)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
by: Qiu, Longpeng, et al.
Published: (2025)
by: Qiu, Longpeng, et al.
Published: (2025)
Similar Items
-
CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents
by: Fei, Tianxiang, et al.
Published: (2026) -
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
by: Li, Zheng, et al.
Published: (2025) -
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
by: Xu, Chenning, et al.
Published: (2026) -
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
by: Song, Mingyang, et al.
Published: (2026) -
Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
by: Song, Mingyang, et al.
Published: (2026)