Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
Fuente:
arXiv
Saved in:
| Main Authors: | Hakim, Muhammad Alif Al, Wicaksono, Alfan Farizki, Koto, Fajri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
by: Azmi, Muhammad Falensi, et al.
Published: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
by: Hidayat, Naila Shafirni, et al.
Published: (2025)
Pengembangan Model untuk Mendeteksi Kerusakan pada Terumbu Karang dengan Klasifikasi Citra
by: Muhammad, Fadhil, et al.
Published: (2023)
by: Muhammad, Fadhil, et al.
Published: (2023)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)
by: Koto, Fajri
Published: (2024)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
Unveiling the Influence of Amplifying Language-Specific Neurons
by: Rahmanisa, Inaya, et al.
Published: (2025)
by: Rahmanisa, Inaya, et al.
Published: (2025)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026)
by: Aziz, Rashad, et al.
Published: (2026)
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
by: Aisyah, Nurul, et al.
Published: (2026)
by: Aisyah, Nurul, et al.
Published: (2026)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
by: Liu, Chen Cecilia, et al.
Published: (2023)
by: Liu, Chen Cecilia, et al.
Published: (2023)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
by: Tonga, Junior Cedric, et al.
Published: (2026)
by: Tonga, Junior Cedric, et al.
Published: (2026)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
by: Aisyah, Nurul, et al.
Published: (2025)
by: Aisyah, Nurul, et al.
Published: (2025)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
by: Laiyk, Nurkhan, et al.
Published: (2026)
by: Laiyk, Nurkhan, et al.
Published: (2026)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
by: Goloburda, Maiya, et al.
Published: (2025)
by: Goloburda, Maiya, et al.
Published: (2025)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
by: Tonga, Junior Cedric, et al.
Published: (2025)
by: Tonga, Junior Cedric, et al.
Published: (2025)
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
by: Jiang, Yanbei, et al.
Published: (2026)
by: Jiang, Yanbei, et al.
Published: (2026)
CMMLU: Measuring massive multitask language understanding in Chinese
by: Li, Haonan, et al.
Published: (2023)
by: Li, Haonan, et al.
Published: (2023)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
Language Surgery in Multilingual Large Language Models
by: Lopo, Joanito Agili, et al.
Published: (2025)
by: Lopo, Joanito Agili, et al.
Published: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs
by: Muhamad, Wicaksono Leksono, et al.
Published: (2026)
by: Muhamad, Wicaksono Leksono, et al.
Published: (2026)
Improving Fairness in LLMs Through Testing-Time Adversaries
by: Gregio, Isabela Pereira, et al.
Published: (2025)
by: Gregio, Isabela Pereira, et al.
Published: (2025)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
Hybrid Deep Learning for Legal Text Analysis: Predicting Punishment Durations in Indonesian Court Rulings
by: Ibrahim, Muhammad Amien, et al.
Published: (2024)
by: Ibrahim, Muhammad Amien, et al.
Published: (2024)
Similar Items
-
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
by: Azmi, Muhammad Falensi, et al.
Published: (2025) -
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
by: Hidayat, Naila Shafirni, et al.
Published: (2025) -
Pengembangan Model untuk Mendeteksi Kerusakan pada Terumbu Karang dengan Klasifikasi Citra
by: Muhammad, Fadhil, et al.
Published: (2023) -
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025) -
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)