Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Koto, Fajri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
von: Aisyah, Nurul, et al.
Veröffentlicht: (2026)
von: Aisyah, Nurul, et al.
Veröffentlicht: (2026)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
von: Aziz, Rashad, et al.
Veröffentlicht: (2026)
von: Aziz, Rashad, et al.
Veröffentlicht: (2026)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
von: Jiang, Yanbei, et al.
Veröffentlicht: (2026)
von: Jiang, Yanbei, et al.
Veröffentlicht: (2026)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2025)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2025)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
von: Hakim, Muhammad Alif Al, et al.
Veröffentlicht: (2026)
von: Hakim, Muhammad Alif Al, et al.
Veröffentlicht: (2026)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2026)
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2026)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
von: Liang, Junhong, et al.
Veröffentlicht: (2026)
von: Liang, Junhong, et al.
Veröffentlicht: (2026)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
von: Aisyah, Nurul, et al.
Veröffentlicht: (2025)
von: Aisyah, Nurul, et al.
Veröffentlicht: (2025)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2026)
von: Tonga, Junior Cedric, et al.
Veröffentlicht: (2026)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
von: Pranida, Salsabila Zahirah, et al.
Veröffentlicht: (2025)
von: Pranida, Salsabila Zahirah, et al.
Veröffentlicht: (2025)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
von: Azmi, Muhammad Falensi, et al.
Veröffentlicht: (2025)
von: Azmi, Muhammad Falensi, et al.
Veröffentlicht: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
von: Hanif, Ikhlasul Akmal, et al.
Veröffentlicht: (2026)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
von: Goloburda, Maiya, et al.
Veröffentlicht: (2025)
von: Goloburda, Maiya, et al.
Veröffentlicht: (2025)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2025)
von: Laiyk, Nurkhan, et al.
Veröffentlicht: (2025)
CMMLU: Measuring massive multitask language understanding in Chinese
von: Li, Haonan, et al.
Veröffentlicht: (2023)
von: Li, Haonan, et al.
Veröffentlicht: (2023)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
von: Elsetohy, Alaa, et al.
Veröffentlicht: (2026)
von: Elsetohy, Alaa, et al.
Veröffentlicht: (2026)
RealChart2Code: Advancing Chart-to-Code Generation with Real Data and Multi-Task Evaluation
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
von: Cruz, Jan Christian Blaise, et al.
Veröffentlicht: (2026)
von: Cruz, Jan Christian Blaise, et al.
Veröffentlicht: (2026)
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2026)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2026)
Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare
von: Lamparth, Max, et al.
Veröffentlicht: (2025)
von: Lamparth, Max, et al.
Veröffentlicht: (2025)
Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Zhu, Jianfeng, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2026)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2026)
Vision Language Models are Confused Tourists
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
von: Togmanov, Mukhammed, et al.
Veröffentlicht: (2025)
von: Togmanov, Mukhammed, et al.
Veröffentlicht: (2025)
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
von: Wei, Lu, et al.
Veröffentlicht: (2025)
von: Wei, Lu, et al.
Veröffentlicht: (2025)
Commonsense Reasoning in Arab Culture
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
CogRAG+: Cognitive-Level Guided Diagnosis and Remediation of Memory and Reasoning Deficiencies in Professional Exam QA
von: Wang, Xudong, et al.
Veröffentlicht: (2026)
von: Wang, Xudong, et al.
Veröffentlicht: (2026)
Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025) -
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
von: Aisyah, Nurul, et al.
Veröffentlicht: (2026) -
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025) -
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025) -
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2025)