Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bsharat, Sondos Mahmoud, Ranjan, Mukul, Myrzakhan, Aidar, Liu, Jiacheng, Guo, Bowei, Tang, Shengkun, Liu, Zhuang, Li, Yuanzhi, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023)
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
von: Ranjan, Mukul, et al.
Veröffentlicht: (2026)
von: Ranjan, Mukul, et al.
Veröffentlicht: (2026)
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
von: KJ, Sankalp, et al.
Veröffentlicht: (2025)
von: KJ, Sankalp, et al.
Veröffentlicht: (2025)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
von: Plaza, Irene, et al.
Veröffentlicht: (2024)
von: Plaza, Irene, et al.
Veröffentlicht: (2024)
A Survey on Diffusion Language Models
von: Li, Tianyi, et al.
Veröffentlicht: (2025)
von: Li, Tianyi, et al.
Veröffentlicht: (2025)
Bi-Mamba: Towards Accurate 1-Bit State Space Models
von: Tang, Shengkun, et al.
Veröffentlicht: (2024)
von: Tang, Shengkun, et al.
Veröffentlicht: (2024)
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
von: Shang, Xinyi, et al.
Veröffentlicht: (2026)
von: Shang, Xinyi, et al.
Veröffentlicht: (2026)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
von: Altakrori, Malik H., et al.
Veröffentlicht: (2025)
von: Altakrori, Malik H., et al.
Veröffentlicht: (2025)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
Are We Done with MMLU?
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
von: Xu, Weikai, et al.
Veröffentlicht: (2025)
von: Xu, Weikai, et al.
Veröffentlicht: (2025)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
von: Wang, Wentian, et al.
Veröffentlicht: (2024)
von: Wang, Wentian, et al.
Veröffentlicht: (2024)
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2024)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2024)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge
von: Abdelaal, Ali, et al.
Veröffentlicht: (2026)
von: Abdelaal, Ali, et al.
Veröffentlicht: (2026)
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
von: Pramodya, Ashmari, et al.
Veröffentlicht: (2025)
von: Pramodya, Ashmari, et al.
Veröffentlicht: (2025)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama
von: Etori, Naome A., et al.
Veröffentlicht: (2025)
von: Etori, Naome A., et al.
Veröffentlicht: (2025)
Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence
von: He, Bowei
Veröffentlicht: (2026)
von: He, Bowei
Veröffentlicht: (2026)
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Prompt Mining for Language-based Human Mobility Forecasting
von: Xue, Hao, et al.
Veröffentlicht: (2024)
von: Xue, Hao, et al.
Veröffentlicht: (2024)
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
von: Guo, Bowei, et al.
Veröffentlicht: (2025)
von: Guo, Bowei, et al.
Veröffentlicht: (2025)
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
von: Huang, Kun, et al.
Veröffentlicht: (2025)
von: Huang, Kun, et al.
Veröffentlicht: (2025)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
von: Ma, Shengkun, et al.
Veröffentlicht: (2025)
von: Ma, Shengkun, et al.
Veröffentlicht: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2026)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
von: Yüksel, Arda, et al.
Veröffentlicht: (2024)
von: Yüksel, Arda, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024) -
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023) -
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026) -
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025) -
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)