Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino
Fuente:
arXiv
Salvato in:
| Autori principali: | Montalan, Jann Railey, Ngui, Jian Gang, Leong, Wei Qi, Susanto, Yosephine, Rengarajan, Hamsawardhini, Aji, Alham Fikri, Tjhi, William Chandra |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
di: Susanto, Yosephine, et al.
Pubblicazione: (2025)
di: Susanto, Yosephine, et al.
Pubblicazione: (2025)
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
di: Aung, Thura, et al.
Pubblicazione: (2026)
di: Aung, Thura, et al.
Pubblicazione: (2026)
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025)
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
di: Montalan, Jann Railey, et al.
Pubblicazione: (2025)
di: Montalan, Jann Railey, et al.
Pubblicazione: (2025)
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
di: Tasawong, Panuthep, et al.
Pubblicazione: (2026)
di: Tasawong, Panuthep, et al.
Pubblicazione: (2026)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
di: Aji, Alham Fikri, et al.
Pubblicazione: (2025)
di: Aji, Alham Fikri, et al.
Pubblicazione: (2025)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
di: Attia, Ahmed, et al.
Pubblicazione: (2026)
di: Attia, Ahmed, et al.
Pubblicazione: (2026)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
di: Chevi, Rendi, et al.
Pubblicazione: (2024)
di: Chevi, Rendi, et al.
Pubblicazione: (2024)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
SEA-LION: Southeast Asian Languages in One Network
di: Ng, Raymond, et al.
Pubblicazione: (2025)
di: Ng, Raymond, et al.
Pubblicazione: (2025)
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
di: Mansurov, Jonibek, et al.
Pubblicazione: (2024)
di: Mansurov, Jonibek, et al.
Pubblicazione: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
Sense Representations Are Inducible Interfaces
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2025)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2025)
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models
di: Lyu, Chenyang, et al.
Pubblicazione: (2024)
di: Lyu, Chenyang, et al.
Pubblicazione: (2024)
How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
di: Chevi, Rendi, et al.
Pubblicazione: (2025)
di: Chevi, Rendi, et al.
Pubblicazione: (2025)
Language-Specific Latent Process Hinders Cross-Lingual Performance
di: Lim, Zheng Wei, et al.
Pubblicazione: (2025)
di: Lim, Zheng Wei, et al.
Pubblicazione: (2025)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
di: Imam, Mohamed Fazli, et al.
Pubblicazione: (2025)
di: Imam, Mohamed Fazli, et al.
Pubblicazione: (2025)
Efficient and Interpretable Grammatical Error Correction with Mixture of Experts
di: Qorib, Muhammad Reza, et al.
Pubblicazione: (2024)
di: Qorib, Muhammad Reza, et al.
Pubblicazione: (2024)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
di: Elshabrawy, Ahmed, et al.
Pubblicazione: (2024)
di: Elshabrawy, Ahmed, et al.
Pubblicazione: (2024)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
di: Susanto, Lucky, et al.
Pubblicazione: (2026)
di: Susanto, Lucky, et al.
Pubblicazione: (2026)
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2025)
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2025)
Multilinguality as Sense Adaptation
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
di: Hudi, Frederikus, et al.
Pubblicazione: (2025)
di: Hudi, Frederikus, et al.
Pubblicazione: (2025)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
di: Nur'aini, Khumaisa, et al.
Pubblicazione: (2026)
di: Nur'aini, Khumaisa, et al.
Pubblicazione: (2026)
Predicting the Order of Upcoming Tokens Improves Language Modeling
di: Zuhri, Zayd M. K., et al.
Pubblicazione: (2025)
di: Zuhri, Zayd M. K., et al.
Pubblicazione: (2025)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2026)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2026)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
di: Zuhri, Zayd M. K., et al.
Pubblicazione: (2025)
di: Zuhri, Zayd M. K., et al.
Pubblicazione: (2025)
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language
di: Susanto, Lucky, et al.
Pubblicazione: (2024)
di: Susanto, Lucky, et al.
Pubblicazione: (2024)
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2025)
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2025)
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
di: Wu, Minghao, et al.
Pubblicazione: (2023)
di: Wu, Minghao, et al.
Pubblicazione: (2023)
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
di: Zuhri, Zayd Muhammad Kawakibi, et al.
Pubblicazione: (2024)
di: Zuhri, Zayd Muhammad Kawakibi, et al.
Pubblicazione: (2024)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2024)
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2024)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2026)
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2026)
QLESS: A Quantized Approach for Data Valuation and Selection in Large Language Model Fine-Tuning
di: Ananta, Moses, et al.
Pubblicazione: (2025)
di: Ananta, Moses, et al.
Pubblicazione: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2024)
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2024)
IteRABRe: Iterative Recovery-Aided Block Reduction
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2025)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2025)
ThaiCoref: Thai Coreference Resolution Dataset
di: Trakuekul, Pontakorn, et al.
Pubblicazione: (2024)
di: Trakuekul, Pontakorn, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
di: Susanto, Yosephine, et al.
Pubblicazione: (2025) -
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
di: Aung, Thura, et al.
Pubblicazione: (2026) -
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025) -
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
di: Montalan, Jann Railey, et al.
Pubblicazione: (2025) -
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)