IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Hanif, Ikhlasul Akmal, Azmi, Muhammad Falensi, Tjiaranata, Filbert Aurelian, Yulianrifat, Eryawan Presma, Koto, Fajri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026)
by: Aziz, Rashad, et al.
Published: (2026)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)
by: Koto, Fajri
Published: (2024)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
by: Aisyah, Nurul, et al.
Published: (2025)
by: Aisyah, Nurul, et al.
Published: (2025)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
by: Tonga, Junior Cedric, et al.
Published: (2026)
by: Tonga, Junior Cedric, et al.
Published: (2026)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
by: Liu, Chen Cecilia, et al.
Published: (2023)
by: Liu, Chen Cecilia, et al.
Published: (2023)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
by: Hakim, Muhammad Alif Al, et al.
Published: (2026)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
by: Kim, Sean, et al.
Published: (2025)
by: Kim, Sean, et al.
Published: (2025)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
by: Aisyah, Nurul, et al.
Published: (2026)
by: Aisyah, Nurul, et al.
Published: (2026)
IndoBERT-Sentiment: Context-Conditioned Sentiment Classification for Indonesian Text
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
IndoBERT-Relevancy: A Context-Conditioned Relevancy Classifier for Indonesian Text
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
by: Jiang, Yanbei, et al.
Published: (2026)
by: Jiang, Yanbei, et al.
Published: (2026)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
by: Pelosio, Giulio, et al.
Published: (2025)
by: Pelosio, Giulio, et al.
Published: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
by: Choudhary, Mukund, et al.
Published: (2025)
by: Choudhary, Mukund, et al.
Published: (2025)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
by: Yamamoto, Taisei, et al.
Published: (2025)
by: Yamamoto, Taisei, et al.
Published: (2025)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
The Impact of Inference Acceleration on Bias of LLMs
by: Kirsten, Elisabeth, et al.
Published: (2024)
by: Kirsten, Elisabeth, et al.
Published: (2024)
Mitigating Cultural Bias in LLMs via Multi-Agent Cultural Debate
by: Tan, Qian, et al.
Published: (2026)
by: Tan, Qian, et al.
Published: (2026)
Predictive Policing and Corporate Governance-Reframing Business Judgment Rule as a Preventive Framework for Corruption in Indonesian State-Owned Enterprises
by: Adhary, Mahaputra, et al.
Published: (2026)
by: Adhary, Mahaputra, et al.
Published: (2026)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
by: Laiyk, Nurkhan, et al.
Published: (2026)
by: Laiyk, Nurkhan, et al.
Published: (2026)
AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
by: Beux, Yann Le, et al.
Published: (2025)
by: Beux, Yann Le, et al.
Published: (2025)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
Similar Items
-
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025) -
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025) -
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026) -
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
by: Koto, Fajri, et al.
Published: (2024) -
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)