MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
Fuente:
arXiv
Saved in:
| Main Authors: | Anugraha, David, Kuwanto, Garry, Susanto, Lucky, Wijaya, Derry Tanti, Winata, Genta Indra |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
Do Language Models Understand Honorific Systems in Javanese?
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
Could We Have Had Better Multilingual LLMs If English Was Not the Central Language?
by: Diandaru, Ryandito, et al.
Published: (2024)
by: Diandaru, Ryandito, et al.
Published: (2024)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
by: Susanto, Lucky, et al.
Published: (2026)
by: Susanto, Lucky, et al.
Published: (2026)
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
by: Merin, Adril Putra, et al.
Published: (2026)
by: Merin, Adril Putra, et al.
Published: (2026)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
by: Winata, Genta Indra, et al.
Published: (2025)
by: Winata, Genta Indra, et al.
Published: (2025)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
by: Susanto, Lucky, et al.
Published: (2024)
by: Susanto, Lucky, et al.
Published: (2024)
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
Generating Faithful Text From a Knowledge Graph with Noisy Reference Text
by: Hashem, Tahsina, et al.
Published: (2023)
by: Hashem, Tahsina, et al.
Published: (2023)
Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
by: Susanto, Lucky, et al.
Published: (2025)
by: Susanto, Lucky, et al.
Published: (2025)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
MINERS: Multilingual Language Models as Semantic Retrievers
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
by: Cettolo, Mauro, et al.
Published: (2025)
by: Cettolo, Mauro, et al.
Published: (2025)
Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
by: Tang, Zilu, et al.
Published: (2025)
by: Tang, Zilu, et al.
Published: (2025)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025)
by: Khelli, Maria, et al.
Published: (2025)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
by: Zhang, Luke, et al.
Published: (2026)
by: Zhang, Luke, et al.
Published: (2026)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)
by: Magdy, Samar M., et al.
Published: (2026)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
by: Winata, Genta Indra, et al.
Published: (2026)
by: Winata, Genta Indra, et al.
Published: (2026)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
by: Sun, Yirong, et al.
Published: (2024)
by: Sun, Yirong, et al.
Published: (2024)
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information
by: Susanto, Lucky, et al.
Published: (2025)
by: Susanto, Lucky, et al.
Published: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
by: Chakraborty, Amartya, et al.
Published: (2025)
by: Chakraborty, Amartya, et al.
Published: (2025)
Similar Items
-
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024) -
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024) -
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025) -
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025) -
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)