Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hidayat, Naila Shafirni, Kautsar, Muhammad Dehan Al, Wicaksono, Alfan Farizki, Koto, Fajri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
von: Azmi, Muhammad Falensi, et al.
Veröffentlicht: (2025)
von: Azmi, Muhammad Falensi, et al.
Veröffentlicht: (2025)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
von: Hakim, Muhammad Alif Al, et al.
Veröffentlicht: (2026)
von: Hakim, Muhammad Alif Al, et al.
Veröffentlicht: (2026)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
von: Aisyah, Nurul, et al.
Veröffentlicht: (2026)
von: Aisyah, Nurul, et al.
Veröffentlicht: (2026)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
von: Aisyah, Nurul, et al.
Veröffentlicht: (2025)
von: Aisyah, Nurul, et al.
Veröffentlicht: (2025)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
von: Koto, Fajri
Veröffentlicht: (2024)
von: Koto, Fajri
Veröffentlicht: (2024)
Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
Vision Language Models are Confused Tourists
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Pengembangan Model untuk Mendeteksi Kerusakan pada Terumbu Karang dengan Klasifikasi Citra
von: Muhammad, Fadhil, et al.
Veröffentlicht: (2023)
von: Muhammad, Fadhil, et al.
Veröffentlicht: (2023)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2025)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
von: Andrylie, Lyzander Marciano, et al.
Veröffentlicht: (2025)
von: Andrylie, Lyzander Marciano, et al.
Veröffentlicht: (2025)
Unveiling the Influence of Amplifying Language-Specific Neurons
von: Rahmanisa, Inaya, et al.
Veröffentlicht: (2025)
von: Rahmanisa, Inaya, et al.
Veröffentlicht: (2025)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2025)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
von: Aziz, Rashad, et al.
Veröffentlicht: (2026)
von: Aziz, Rashad, et al.
Veröffentlicht: (2026)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2026)
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2026)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
von: Liang, Junhong, et al.
Veröffentlicht: (2026)
von: Liang, Junhong, et al.
Veröffentlicht: (2026)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2026)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2026)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
von: Pranida, Salsabila Zahirah, et al.
Veröffentlicht: (2025)
von: Pranida, Salsabila Zahirah, et al.
Veröffentlicht: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
von: Jiang, Yanbei, et al.
Veröffentlicht: (2026)
von: Jiang, Yanbei, et al.
Veröffentlicht: (2026)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2025)
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2025)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
von: Goloburda, Maiya, et al.
Veröffentlicht: (2025)
von: Goloburda, Maiya, et al.
Veröffentlicht: (2025)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
von: Elsetohy, Alaa, et al.
Veröffentlicht: (2026)
von: Elsetohy, Alaa, et al.
Veröffentlicht: (2026)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks
von: Salido, Eva Sánchez, et al.
Veröffentlicht: (2025)
von: Salido, Eva Sánchez, et al.
Veröffentlicht: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
CMMLU: Measuring massive multitask language understanding in Chinese
von: Li, Haonan, et al.
Veröffentlicht: (2023)
von: Li, Haonan, et al.
Veröffentlicht: (2023)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
von: Liu, Jiayi, et al.
Veröffentlicht: (2026)
von: Liu, Jiayi, et al.
Veröffentlicht: (2026)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
von: Wang, Xunzhi, et al.
Veröffentlicht: (2024)
von: Wang, Xunzhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
von: Azmi, Muhammad Falensi, et al.
Veröffentlicht: (2025) -
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
von: Hakim, Muhammad Alif Al, et al.
Veröffentlicht: (2026) -
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025) -
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
von: Aisyah, Nurul, et al.
Veröffentlicht: (2026) -
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
von: Aisyah, Nurul, et al.
Veröffentlicht: (2025)