KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
Fuente:
arXiv
Guardado en:
| Autores principales: | Togmanov, Mukhammed, Mukhituly, Nurdaulet, Turmakhan, Diana, Mansurov, Jonibek, Goloburda, Maiya, Sakip, Akhmed, Xie, Zhuohan, Wang, Yuxia, Syzdykov, Bekassyl, Laiyk, Nurkhan, Aji, Alham Fikri, Kochmar, Ekaterina, Nakov, Preslav, Koto, Fajri |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
por: Goloburda, Maiya, et al.
Publicado: (2025)
por: Goloburda, Maiya, et al.
Publicado: (2025)
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
por: Mansurov, Jonibek, et al.
Publicado: (2024)
por: Mansurov, Jonibek, et al.
Publicado: (2024)
Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting
por: Koto, Fajri, et al.
Publicado: (2025)
por: Koto, Fajri, et al.
Publicado: (2025)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
por: Laiyk, Nurkhan, et al.
Publicado: (2025)
por: Laiyk, Nurkhan, et al.
Publicado: (2025)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
por: Laiyk, Nurkhan, et al.
Publicado: (2026)
por: Laiyk, Nurkhan, et al.
Publicado: (2026)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
por: Goloburda, Maiya, et al.
Publicado: (2026)
por: Goloburda, Maiya, et al.
Publicado: (2026)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
por: Liang, Junhong, et al.
Publicado: (2026)
por: Liang, Junhong, et al.
Publicado: (2026)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
por: Koto, Fajri
Publicado: (2024)
por: Koto, Fajri
Publicado: (2024)
Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
por: Wang, Yuxia, et al.
Publicado: (2025)
por: Wang, Yuxia, et al.
Publicado: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
por: Sakip, Akhmed, et al.
Publicado: (2026)
por: Sakip, Akhmed, et al.
Publicado: (2026)
KazParC: Kazakh Parallel Corpus for Machine Translation
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
KazQAD: Kazakh Open-Domain Question Answering Dataset
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human
por: Wang, Yuxia, et al.
Publicado: (2025)
por: Wang, Yuxia, et al.
Publicado: (2025)
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
por: Yeshpanov, Rustem, et al.
Publicado: (2024)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
por: Elsetohy, Alaa, et al.
Publicado: (2026)
por: Elsetohy, Alaa, et al.
Publicado: (2026)
PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?
por: Petukhova, Kseniia, et al.
Publicado: (2024)
por: Petukhova, Kseniia, et al.
Publicado: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
por: Abilbekov, Adal, et al.
Publicado: (2024)
por: Abilbekov, Adal, et al.
Publicado: (2024)
KazByte: Adapting Qwen models to Kazakh via Byte-level Adapter
por: Akylzhanov, Rauan
Publicado: (2026)
por: Akylzhanov, Rauan
Publicado: (2026)
SPIRIT: Patching Speech Language Models against Jailbreak Attacks
por: Djanibekov, Amirbek, et al.
Publicado: (2025)
por: Djanibekov, Amirbek, et al.
Publicado: (2025)
PetKaz at SemEval-2024 Task 3: Advancing Emotion Classification with an LLM for Emotion-Cause Pair Extraction in Conversations
por: Kazakov, Roman, et al.
Publicado: (2024)
por: Kazakov, Roman, et al.
Publicado: (2024)
Vision Language Models are Confused Tourists
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
por: Tonga, Junior Cedric, et al.
Publicado: (2025)
por: Tonga, Junior Cedric, et al.
Publicado: (2025)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
por: Vashurin, Roman, et al.
Publicado: (2025)
por: Vashurin, Roman, et al.
Publicado: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
por: Yari, Amir Hossein, et al.
Publicado: (2026)
por: Yari, Amir Hossein, et al.
Publicado: (2026)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
por: Yari, Amir Hossein, et al.
Publicado: (2025)
por: Yari, Amir Hossein, et al.
Publicado: (2025)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
por: Koto, Fajri, et al.
Publicado: (2024)
por: Koto, Fajri, et al.
Publicado: (2024)
Sycophancy Hides Linearly in the Attention Heads
por: Genadi, Rifo, et al.
Publicado: (2026)
por: Genadi, Rifo, et al.
Publicado: (2026)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
por: Altakrori, Malik H., et al.
Publicado: (2025)
por: Altakrori, Malik H., et al.
Publicado: (2025)
NP-complete Problems can be Solved and Verified in Polynomial Time
por: Syzdykov, Mirzakhmet
Publicado: (2025)
por: Syzdykov, Mirzakhmet
Publicado: (2025)
Generalization and Relation of Probabilistic Models to Finite Automata
por: Syzdykov, Mirzakhmet
Publicado: (2025)
por: Syzdykov, Mirzakhmet
Publicado: (2025)
Proof of Millennium Theorem "P versus NP"
por: Syzdykov, Mirzakhmet
Publicado: (2023)
por: Syzdykov, Mirzakhmet
Publicado: (2023)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
Crosslingual Reasoning through Test-Time Scaling
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
por: Aji, Alham Fikri, et al.
Publicado: (2025)
por: Aji, Alham Fikri, et al.
Publicado: (2025)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
por: Attia, Ahmed, et al.
Publicado: (2026)
por: Attia, Ahmed, et al.
Publicado: (2026)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
por: Chevi, Rendi, et al.
Publicado: (2024)
por: Chevi, Rendi, et al.
Publicado: (2024)
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
por: Aziz, Rashad, et al.
Publicado: (2026)
por: Aziz, Rashad, et al.
Publicado: (2026)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
Language Surgery in Multilingual Large Language Models
por: Lopo, Joanito Agili, et al.
Publicado: (2025)
por: Lopo, Joanito Agili, et al.
Publicado: (2025)
Ejemplares similares
-
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
por: Goloburda, Maiya, et al.
Publicado: (2025) -
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
por: Mansurov, Jonibek, et al.
Publicado: (2024) -
Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting
por: Koto, Fajri, et al.
Publicado: (2025) -
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
por: Laiyk, Nurkhan, et al.
Publicado: (2025) -
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
por: Laiyk, Nurkhan, et al.
Publicado: (2026)