Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatti, Hunzalah Hassan, Alam, Firoj |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026)
OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
by: Mousi, Basel, et al.
Published: (2024)
by: Mousi, Basel, et al.
Published: (2024)
ThatiAR: Subjectivity Detection in Arabic News Sentences
by: Suwaileh, Reem, et al.
Published: (2024)
by: Suwaileh, Reem, et al.
Published: (2024)
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
by: Ali, Zien Sheikh, et al.
Published: (2026)
by: Ali, Zien Sheikh, et al.
Published: (2026)
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
by: Hasan, Md. Arid, et al.
Published: (2024)
by: Hasan, Md. Arid, et al.
Published: (2024)
Propaganda to Hate: A Multimodal Analysis of Arabic Memes with Multi-Agent LLMs
by: Alam, Firoj, et al.
Published: (2024)
by: Alam, Firoj, et al.
Published: (2024)
NativQA Framework: Enabling LLMs and VLMs with Native, Local, and Everyday Knowledge
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
LAraBench: Benchmarking Arabic AI with Large Language Models
by: Abdelali, Ahmed, et al.
Published: (2023)
by: Abdelali, Ahmed, et al.
Published: (2023)
Large Language Models for Propaganda Span Annotation
by: Hasanain, Maram, et al.
Published: (2023)
by: Hasanain, Maram, et al.
Published: (2023)
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
by: Ali, Zien Sheikh, et al.
Published: (2026)
by: Ali, Zien Sheikh, et al.
Published: (2026)
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content
by: Kmainasi, Mohamed Bayan, et al.
Published: (2024)
by: Kmainasi, Mohamed Bayan, et al.
Published: (2024)
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
by: Hasan, Md Arid, et al.
Published: (2025)
by: Hasan, Md Arid, et al.
Published: (2025)
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
Native vs Non-Native Language Prompting: A Comparative Analysis
by: Kmainasi, Mohamed Bayan, et al.
Published: (2024)
by: Kmainasi, Mohamed Bayan, et al.
Published: (2024)
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking
by: Dalvi, Fahim, et al.
Published: (2023)
by: Dalvi, Fahim, et al.
Published: (2023)
PropXplain: Can LLMs Enable Explainable Propaganda Detection?
by: Hasanain, Maram, et al.
Published: (2025)
by: Hasanain, Maram, et al.
Published: (2025)
GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge
by: Chowdhury, Shammur Absar, et al.
Published: (2024)
by: Chowdhury, Shammur Absar, et al.
Published: (2024)
ArAIEval Shared Task: Propagandistic Techniques Detection in Unimodal and Multimodal Arabic Content
by: Hasanain, Maram, et al.
Published: (2024)
by: Hasanain, Maram, et al.
Published: (2024)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
by: De Mel, Yomal, et al.
Published: (2024)
by: De Mel, Yomal, et al.
Published: (2024)
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
by: Munakata, Satoshi, et al.
Published: (2024)
by: Munakata, Satoshi, et al.
Published: (2024)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
by: Kaiser, Daniel, et al.
Published: (2026)
by: Kaiser, Daniel, et al.
Published: (2026)
DEM: Distribution Edited Model for Training with Mixed Data Distributions
by: Ram, Dhananjay, et al.
Published: (2024)
by: Ram, Dhananjay, et al.
Published: (2024)
Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs
by: Brown, Nik Bear
Published: (2024)
by: Brown, Nik Bear
Published: (2024)
Uncovering Uncertainty in Transformer Inference
by: Brothers, Greyson, et al.
Published: (2024)
by: Brothers, Greyson, et al.
Published: (2024)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
by: Li, Yangyang
Published: (2025)
by: Li, Yangyang
Published: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
ArMeme: Propagandistic Content in Arabic Memes
by: Alam, Firoj, et al.
Published: (2024)
by: Alam, Firoj, et al.
Published: (2024)
The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
MemeLens: Multilingual Multitask VLMs for Memes
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
by: Guan, Xin, et al.
Published: (2024)
by: Guan, Xin, et al.
Published: (2024)
Adaptive Engram Memory System for Indonesian Language Model: Generative AI Based on TOBA LM for Batak and Minang Language
by: Situngkir, Hokky, et al.
Published: (2026)
by: Situngkir, Hokky, et al.
Published: (2026)
Understanding and Improving Information Preservation in Prompt Compression for LLMs
by: Łajewska, Weronika, et al.
Published: (2025)
by: Łajewska, Weronika, et al.
Published: (2025)
PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models
by: Tikhonov, Alexey
Published: (2024)
by: Tikhonov, Alexey
Published: (2024)
Syllabic Agglutinative Tokenizations for Indonesian LLM: A Study from Gasing Literacy Learning System
by: Situngkir, H., et al.
Published: (2026)
by: Situngkir, H., et al.
Published: (2026)
Similar Items
-
CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025) -
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026) -
OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
by: Alam, Firoj, et al.
Published: (2025) -
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
by: Mousi, Basel, et al.
Published: (2024) -
ThatiAR: Subjectivity Detection in Arabic News Sentences
by: Suwaileh, Reem, et al.
Published: (2024)