Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Yari, Amir Hossein, Kulkarni, Kalmit, Khan, Ahmad Raza, Koto, Fajri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)
by: Koto, Fajri
Published: (2024)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
by: Laiyk, Nurkhan, et al.
Published: (2026)
by: Laiyk, Nurkhan, et al.
Published: (2026)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
by: Aisyah, Nurul, et al.
Published: (2025)
by: Aisyah, Nurul, et al.
Published: (2025)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026)
by: Aziz, Rashad, et al.
Published: (2026)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Fine-grained Gender Control in Machine Translation with Large Language Models
by: Lee, Minwoo, et al.
Published: (2024)
by: Lee, Minwoo, et al.
Published: (2024)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
by: Liu, Chen Cecilia, et al.
Published: (2023)
by: Liu, Chen Cecilia, et al.
Published: (2023)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
by: Guo, Yue, et al.
Published: (2023)
by: Guo, Yue, et al.
Published: (2023)
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
by: Jiang, Yanbei, et al.
Published: (2026)
by: Jiang, Yanbei, et al.
Published: (2026)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024)
by: Zhang, Shiyue, et al.
Published: (2024)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
Crosslingual Optimized Metric for Translation Assessment of Indian Languages
by: Ahsan, Arafat, et al.
Published: (2025)
by: Ahsan, Arafat, et al.
Published: (2025)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
by: Tonga, Junior Cedric, et al.
Published: (2026)
by: Tonga, Junior Cedric, et al.
Published: (2026)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
by: Aisyah, Nurul, et al.
Published: (2026)
by: Aisyah, Nurul, et al.
Published: (2026)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
by: Wang, Shun, et al.
Published: (2024)
by: Wang, Shun, et al.
Published: (2024)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
by: Wu, Guojun, et al.
Published: (2024)
by: Wu, Guojun, et al.
Published: (2024)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
Revolutionizing API Documentation through Summarization
by: Naghshzan, AmirHossein, et al.
Published: (2024)
by: Naghshzan, AmirHossein, et al.
Published: (2024)
CASPR: Automated Evaluation Metric for Contrastive Summarization
by: Ananthamurugan, Nirupan, et al.
Published: (2024)
by: Ananthamurugan, Nirupan, et al.
Published: (2024)
Calibrating Model-Based Evaluation Metrics for Summarization
by: Liu, Hongye, et al.
Published: (2026)
by: Liu, Hongye, et al.
Published: (2026)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
by: Bhattacharjee, Soham, et al.
Published: (2025)
by: Bhattacharjee, Soham, et al.
Published: (2025)
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
by: Almheiri, Saeed, et al.
Published: (2026)
by: Almheiri, Saeed, et al.
Published: (2026)
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
by: Sun, Yirong, et al.
Published: (2024)
by: Sun, Yirong, et al.
Published: (2024)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
by: Belouadi, Jonas, et al.
Published: (2022)
by: Belouadi, Jonas, et al.
Published: (2022)
Revisiting the Markov Property for Machine Translation
by: Du, Cunxiao, et al.
Published: (2024)
by: Du, Cunxiao, et al.
Published: (2024)
Similar Items
-
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025) -
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024) -
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026) -
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
by: Laiyk, Nurkhan, et al.
Published: (2026) -
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)