ELAB: Extensive LLM Alignment Benchmark in Persian Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pourbahman, Zahra, Rajabi, Fatemeh, Sadeghi, Mohammadhossein, Ghahroodi, Omid, Bakhshaei, Somaye, Amini, Arash, Kazemi, Reza, Baghshah, Mahdieh Soleymani |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
FaMTEB: Massive Text Embedding Benchmark in Persian Language
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
von: Ahmadpour, Mohammadjavad, et al.
Veröffentlicht: (2025)
von: Ahmadpour, Mohammadjavad, et al.
Veröffentlicht: (2025)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey
von: Shahhosseini, Fatemeh, et al.
Veröffentlicht: (2025)
von: Shahhosseini, Fatemeh, et al.
Veröffentlicht: (2025)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2026)
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2026)
Persian Musical Instruments Classification Using Polyphonic Data Augmentation
von: Esfangereh, Diba Hadi, et al.
Veröffentlicht: (2025)
von: Esfangereh, Diba Hadi, et al.
Veröffentlicht: (2025)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
von: Paqaleh, Mohammad Mahdi Samiei, et al.
Veröffentlicht: (2025)
von: Paqaleh, Mohammad Mahdi Samiei, et al.
Veröffentlicht: (2025)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
von: Rezaei, Parham, et al.
Veröffentlicht: (2025)
von: Rezaei, Parham, et al.
Veröffentlicht: (2025)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
von: Hasani, Hosein, et al.
Veröffentlicht: (2026)
von: Hasani, Hosein, et al.
Veröffentlicht: (2026)
Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
von: Shahabadi, Hossein, et al.
Veröffentlicht: (2025)
von: Shahabadi, Hossein, et al.
Veröffentlicht: (2025)
PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems
von: Sedghiyeh, Nima, et al.
Veröffentlicht: (2025)
von: Sedghiyeh, Nima, et al.
Veröffentlicht: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!
von: Marioriyad, Arash, et al.
Veröffentlicht: (2024)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2024)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
von: Hosseini, Seyed Mohsen, et al.
Veröffentlicht: (2024)
von: Hosseini, Seyed Mohsen, et al.
Veröffentlicht: (2024)
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2025)
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
von: Marioriyad, Arash, et al.
Veröffentlicht: (2024)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2024)
LLM-Agent-Controller: A Universal Multi-Agent Large Language Model System as a Control Engineer
von: Zahedifar, Rasoul, et al.
Veröffentlicht: (2025)
von: Zahedifar, Rasoul, et al.
Veröffentlicht: (2025)
OPSD: an Offensive Persian Social media Dataset and its baseline evaluations
von: Safayani, Mehran, et al.
Veröffentlicht: (2024)
von: Safayani, Mehran, et al.
Veröffentlicht: (2024)
Hakim: Farsi Text Embedding Model
von: Sarmadi, Mehran, et al.
Veröffentlicht: (2025)
von: Sarmadi, Mehran, et al.
Veröffentlicht: (2025)
The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
von: Samiei, Mahdi, et al.
Veröffentlicht: (2025)
von: Samiei, Mahdi, et al.
Veröffentlicht: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
von: Mirbagheri, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Mirbagheri, Mohammad Reza, et al.
Veröffentlicht: (2025)
PersianRAG: A Retrieval-Augmented Generation System for Persian Language
von: Hosseini, Hossein, et al.
Veröffentlicht: (2024)
von: Hosseini, Hossein, et al.
Veröffentlicht: (2024)
ComAlign: Compositional Alignment in Vision-Language Models
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark
von: Kalahroodi, Mohammad Javad Ranjbar, et al.
Veröffentlicht: (2026)
von: Kalahroodi, Mohammad Javad Ranjbar, et al.
Veröffentlicht: (2026)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
von: Mehri, Faridoun, et al.
Veröffentlicht: (2024)
von: Mehri, Faridoun, et al.
Veröffentlicht: (2024)
PreND: Enhancing Intrinsic Motivation in Reinforcement Learning through Pre-trained Network Distillation
von: Davoodabadi, Mohammadamin, et al.
Veröffentlicht: (2024)
von: Davoodabadi, Mohammadamin, et al.
Veröffentlicht: (2024)
Classification of Breast Cancer Histopathology Images using a Modified Supervised Contrastive Learning Method
von: Sani, Matina Mahdizadeh, et al.
Veröffentlicht: (2024)
von: Sani, Matina Mahdizadeh, et al.
Veröffentlicht: (2024)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024)
Leveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model
von: Ghassabi, Mehrdad, et al.
Veröffentlicht: (2025)
von: Ghassabi, Mehrdad, et al.
Veröffentlicht: (2025)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026) -
FaMTEB: Massive Text Embedding Benchmark in Persian Language
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025) -
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024) -
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025) -
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)