MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Ghahroodi, Omid, Hemmat, Arshia, Nouri, Marzia, Hosseini, Seyed Mohammad Hadi, Dastgheib, Doratossadat, Sanian, Mohammad Vali, Sahebi, Alireza, Zohrabi, Reihaneh, Rohban, Mohammad Hossein, Asgari, Ehsaneddin, Baghshah, Mahdieh Soleymani |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
by: Ghahroodi, Omid, et al.
Published: (2024)
by: Ghahroodi, Omid, et al.
Published: (2024)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
by: Zohrabi, Reihaneh, et al.
Published: (2025)
by: Zohrabi, Reihaneh, et al.
Published: (2025)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
ELAB: Extensive LLM Alignment Benchmark in Persian Language
by: Pourbahman, Zahra, et al.
Published: (2025)
by: Pourbahman, Zahra, et al.
Published: (2025)
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
by: Ahmadpour, Mohammadjavad, et al.
Published: (2025)
by: Ahmadpour, Mohammadjavad, et al.
Published: (2025)
Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
by: Kasaei, Seyed Amir, et al.
Published: (2026)
by: Kasaei, Seyed Amir, et al.
Published: (2026)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
by: Mehri, Faridoun, et al.
Published: (2024)
by: Mehri, Faridoun, et al.
Published: (2024)
Analyzing CLIP's Performance Limitations in Multi-Object Scenarios: A Controlled High-Resolution Study
by: Abbasi, Reza, et al.
Published: (2025)
by: Abbasi, Reza, et al.
Published: (2025)
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
by: Abbasi, Reza, et al.
Published: (2025)
by: Abbasi, Reza, et al.
Published: (2025)
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
by: Cekinmez, Jasin, et al.
Published: (2025)
by: Cekinmez, Jasin, et al.
Published: (2025)
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
by: Sameti, Mohammad Hossein, et al.
Published: (2026)
by: Sameti, Mohammad Hossein, et al.
Published: (2026)
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples
by: Mirzaei, Hossein, et al.
Published: (2025)
by: Mirzaei, Hossein, et al.
Published: (2025)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
by: Sameti, Mohammad Hossein, et al.
Published: (2025)
by: Sameti, Mohammad Hossein, et al.
Published: (2025)
CER: Confidence Enhanced Reasoning in LLMs
by: Razghandi, Ali, et al.
Published: (2025)
by: Razghandi, Ali, et al.
Published: (2025)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
Persian Musical Instruments Classification Using Polyphonic Data Augmentation
by: Esfangereh, Diba Hadi, et al.
Published: (2025)
by: Esfangereh, Diba Hadi, et al.
Published: (2025)
Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey
by: Shahhosseini, Fatemeh, et al.
Published: (2025)
by: Shahhosseini, Fatemeh, et al.
Published: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
by: Mozafari, Mohammad, et al.
Published: (2024)
by: Mozafari, Mohammad, et al.
Published: (2024)
T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2025)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2025)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
by: Hosseini, Seyed Mohsen, et al.
Published: (2024)
by: Hosseini, Seyed Mohsen, et al.
Published: (2024)
Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
The Midas Touch in Gaze vs. Hand Pointing: Modality-Specific Failure Modes and Implications for XR Interfaces
by: Dastgheib, Mohammad, et al.
Published: (2026)
by: Dastgheib, Mohammad, et al.
Published: (2026)
Unmasking the Factual-Conceptual Gap in Persian Language Models
by: Sakhaeirad, Alireza, et al.
Published: (2026)
by: Sakhaeirad, Alireza, et al.
Published: (2026)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
by: Abdollahi, Ali, et al.
Published: (2024)
by: Abdollahi, Ali, et al.
Published: (2024)
3D-Guided Scalable Flow Matching for Generating Volumetric Tissue Spatial Transcriptomics from Serial Histology
by: Sanian, Mohammad Vali, et al.
Published: (2025)
by: Sanian, Mohammad Vali, et al.
Published: (2025)
Similar Items
-
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
by: Ghahroodi, Omid, et al.
Published: (2024) -
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026) -
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
by: Zohrabi, Reihaneh, et al.
Published: (2025) -
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025) -
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
by: Marioriyad, Arash, et al.
Published: (2026)