Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
Fuente:
arXiv
Guardado en:
| Autores principales: | Aisyah, Nurul, Kautsar, Muhammad Dehan Al, Hidayat, Arif, Chowdhury, Raqib, Koto, Fajri |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
por: Aisyah, Nurul, et al.
Publicado: (2026)
por: Aisyah, Nurul, et al.
Publicado: (2026)
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
por: Azmi, Muhammad Falensi, et al.
Publicado: (2025)
por: Azmi, Muhammad Falensi, et al.
Publicado: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces
por: Koto, Fajri, et al.
Publicado: (2024)
por: Koto, Fajri, et al.
Publicado: (2024)
Vision Language Models are Confused Tourists
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
por: Almheiri, Saeed, et al.
Publicado: (2025)
por: Almheiri, Saeed, et al.
Publicado: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
por: Koto, Fajri
Publicado: (2024)
por: Koto, Fajri
Publicado: (2024)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
por: Hanif, Ikhlasul Akmal, et al.
Publicado: (2026)
por: Hanif, Ikhlasul Akmal, et al.
Publicado: (2026)
Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages
por: Cahyawijaya, Samuel, et al.
Publicado: (2024)
por: Cahyawijaya, Samuel, et al.
Publicado: (2024)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
por: Hakim, Muhammad Alif Al, et al.
Publicado: (2026)
por: Hakim, Muhammad Alif Al, et al.
Publicado: (2026)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
por: Yari, Amir Hossein, et al.
Publicado: (2025)
por: Yari, Amir Hossein, et al.
Publicado: (2025)
Language Surgery in Multilingual Large Language Models
por: Lopo, Joanito Agili, et al.
Publicado: (2025)
por: Lopo, Joanito Agili, et al.
Publicado: (2025)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
por: Yari, Amir Hossein, et al.
Publicado: (2025)
por: Yari, Amir Hossein, et al.
Publicado: (2025)
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
por: Laiyk, Nurkhan, et al.
Publicado: (2026)
por: Laiyk, Nurkhan, et al.
Publicado: (2026)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
por: Koto, Fajri, et al.
Publicado: (2024)
por: Koto, Fajri, et al.
Publicado: (2024)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
por: Pranida, Salsabila Zahirah, et al.
Publicado: (2025)
por: Pranida, Salsabila Zahirah, et al.
Publicado: (2025)
Large Malaysian Language Model Based on Mistral for Enhanced Local Language Understanding
por: Zolkepli, Husein, et al.
Publicado: (2024)
por: Zolkepli, Husein, et al.
Publicado: (2024)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
por: Aziz, Rashad, et al.
Publicado: (2026)
por: Aziz, Rashad, et al.
Publicado: (2026)
Analyzing Large Language Models for Classroom Discussion Assessment
por: Tran, Nhat, et al.
Publicado: (2024)
por: Tran, Nhat, et al.
Publicado: (2024)
MaLLaM -- Malaysia Large Language Model
por: Zolkepli, Husein, et al.
Publicado: (2024)
por: Zolkepli, Husein, et al.
Publicado: (2024)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
por: Laiyk, Nurkhan, et al.
Publicado: (2025)
por: Laiyk, Nurkhan, et al.
Publicado: (2025)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
por: Liu, Chen Cecilia, et al.
Publicado: (2023)
por: Liu, Chen Cecilia, et al.
Publicado: (2023)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
por: Liang, Junhong, et al.
Publicado: (2026)
por: Liang, Junhong, et al.
Publicado: (2026)
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2025)
LLMs as Cultural Archives: Cultural Commonsense Knowledge Graph Extraction
por: Tonga, Junior Cedric, et al.
Publicado: (2026)
por: Tonga, Junior Cedric, et al.
Publicado: (2026)
Mordal: Automated Pretrained Model Selection for Vision Language Models
por: He, Shiqi, et al.
Publicado: (2025)
por: He, Shiqi, et al.
Publicado: (2025)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
por: Togmanov, Mukhammed, et al.
Publicado: (2025)
por: Togmanov, Mukhammed, et al.
Publicado: (2025)
Evaluating Large Language Models in Analysing Classroom Dialogue
por: Long, Yun, et al.
Publicado: (2024)
por: Long, Yun, et al.
Publicado: (2024)
Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations
por: Zolkepli, Husein, et al.
Publicado: (2024)
por: Zolkepli, Husein, et al.
Publicado: (2024)
Generalists vs. Specialists: Evaluating Large Language Models for Urdu
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
Investigating Cultural Alignment of Large Language Models
por: AlKhamissi, Badr, et al.
Publicado: (2024)
por: AlKhamissi, Badr, et al.
Publicado: (2024)
Stuttering-Aware Automatic Speech Recognition for Indonesian Language
por: Muhammad, Fadhil, et al.
Publicado: (2026)
por: Muhammad, Fadhil, et al.
Publicado: (2026)
xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
por: Yu, Qingchen, et al.
Publicado: (2024)
por: Yu, Qingchen, et al.
Publicado: (2024)
Evaluating Large Language Models on Urdu Idiom Translation
por: Khan, Muhammad Farmal, et al.
Publicado: (2025)
por: Khan, Muhammad Farmal, et al.
Publicado: (2025)
Design and Application of Multimodal Large Language Model Based System for End to End Automation of Accident Dataset Generation
por: Chowdhury, MD Thamed Bin Zaman, et al.
Publicado: (2025)
por: Chowdhury, MD Thamed Bin Zaman, et al.
Publicado: (2025)
Using Large Language Models for Automated Grading of Student Writing about Science
por: Impey, Chris, et al.
Publicado: (2024)
por: Impey, Chris, et al.
Publicado: (2024)
Ejemplares similares
-
Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia
por: Aisyah, Nurul, et al.
Publicado: (2026) -
What Do Indonesians Really Need from Language Technology? A Nationwide Survey
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025) -
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2025) -
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
por: Azmi, Muhammad Falensi, et al.
Publicado: (2025) -
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)