MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Winata, Genta Indra, Anugraha, David, Susanto, Lucky, Kuwanto, Garry, Wijaya, Derry Tanti |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
Do Language Models Understand Honorific Systems in Javanese?
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
by: Winata, Genta Indra, et al.
Published: (2025)
by: Winata, Genta Indra, et al.
Published: (2025)
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
by: Merin, Adril Putra, et al.
Published: (2026)
by: Merin, Adril Putra, et al.
Published: (2026)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Generating Faithful and Salient Text from Multimodal Data
by: Hashem, Tahsina, et al.
Published: (2024)
by: Hashem, Tahsina, et al.
Published: (2024)
Could We Have Had Better Multilingual LLMs If English Was Not the Central Language?
by: Diandaru, Ryandito, et al.
Published: (2024)
by: Diandaru, Ryandito, et al.
Published: (2024)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
by: Susanto, Lucky, et al.
Published: (2026)
by: Susanto, Lucky, et al.
Published: (2026)
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
MINERS: Multilingual Language Models as Semantic Retrievers
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
CTest-Metric: A Unified Framework to Assess Clinical Validity of Metrics for CT Report Generation
by: Sharma, Vanshali, et al.
Published: (2026)
by: Sharma, Vanshali, et al.
Published: (2026)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025)
by: Khelli, Maria, et al.
Published: (2025)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
by: Leiter, Christoph, et al.
Published: (2025)
by: Leiter, Christoph, et al.
Published: (2025)
CRG Score: A Distribution-Aware Clinical Metric for Radiology Report Generation
by: Hamamci, Ibrahim Ethem, et al.
Published: (2025)
by: Hamamci, Ibrahim Ethem, et al.
Published: (2025)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
by: Wada, Yuiga, et al.
Published: (2024)
by: Wada, Yuiga, et al.
Published: (2024)
Generating Faithful Text From a Knowledge Graph with Noisy Reference Text
by: Hashem, Tahsina, et al.
Published: (2023)
by: Hashem, Tahsina, et al.
Published: (2023)
Semantic Textual Similarity Assessment in Chest X-ray Reports Using a Domain-Specific Cosine-Based Metric
by: Picha, Sayeh Gholipour, et al.
Published: (2024)
by: Picha, Sayeh Gholipour, et al.
Published: (2024)
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information
by: Susanto, Lucky, et al.
Published: (2025)
by: Susanto, Lucky, et al.
Published: (2025)
Behind Maya: Building a Multilingual Vision Language Model
by: Alam, Nahid, et al.
Published: (2025)
by: Alam, Nahid, et al.
Published: (2025)
Maya: An Instruction Finetuned Multilingual Multimodal Model
by: Alam, Nahid, et al.
Published: (2024)
by: Alam, Nahid, et al.
Published: (2024)
Explaining Human Preferences via Metrics for Structured 3D Reconstruction
by: Langerman, Jack, et al.
Published: (2025)
by: Langerman, Jack, et al.
Published: (2025)
MetaRank: Task-Aware Metric Selection for Model Transferability Estimation
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
by: Liu, Lu, et al.
Published: (2024)
by: Liu, Lu, et al.
Published: (2024)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
by: Biyyala, Varun, et al.
Published: (2025)
by: Biyyala, Varun, et al.
Published: (2025)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026)
by: Baharoon, Mohammed, et al.
Published: (2026)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
Similar Items
-
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024) -
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024) -
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025) -
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025) -
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)