The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Marioriyad, Arash, Ghahroodi, Omid, Asgari, Ehsaneddin, Rohban, Mohammad Hossein, Baghshah, Mahdieh Soleymani |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
por: Marioriyad, Arash, et al.
Publicado: (2025)
por: Marioriyad, Arash, et al.
Publicado: (2025)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
por: Marioriyad, Arash, et al.
Publicado: (2026)
por: Marioriyad, Arash, et al.
Publicado: (2026)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
por: Ghahroodi, Omid, et al.
Publicado: (2024)
por: Ghahroodi, Omid, et al.
Publicado: (2024)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
por: Marioriyad, Arash, et al.
Publicado: (2024)
por: Marioriyad, Arash, et al.
Publicado: (2024)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
por: Rezaei, Parham, et al.
Publicado: (2025)
por: Rezaei, Parham, et al.
Publicado: (2025)
Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
por: Marioriyad, Arash, et al.
Publicado: (2025)
por: Marioriyad, Arash, et al.
Publicado: (2025)
Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey
por: Shahhosseini, Fatemeh, et al.
Publicado: (2025)
por: Shahhosseini, Fatemeh, et al.
Publicado: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2025)
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2025)
Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
por: Kasaei, Seyed Amir, et al.
Publicado: (2026)
por: Kasaei, Seyed Amir, et al.
Publicado: (2026)
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!
por: Marioriyad, Arash, et al.
Publicado: (2024)
por: Marioriyad, Arash, et al.
Publicado: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
por: Abbasi, Reza, et al.
Publicado: (2024)
por: Abbasi, Reza, et al.
Publicado: (2024)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
ELAB: Extensive LLM Alignment Benchmark in Persian Language
por: Pourbahman, Zahra, et al.
Publicado: (2025)
por: Pourbahman, Zahra, et al.
Publicado: (2025)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
por: Sameti, Mohammad Hossein, et al.
Publicado: (2025)
por: Sameti, Mohammad Hossein, et al.
Publicado: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
por: Abbasi, Reza, et al.
Publicado: (2024)
por: Abbasi, Reza, et al.
Publicado: (2024)
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
por: Cekinmez, Jasin, et al.
Publicado: (2025)
por: Cekinmez, Jasin, et al.
Publicado: (2025)
Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
por: Shahabadi, Hossein, et al.
Publicado: (2025)
por: Shahabadi, Hossein, et al.
Publicado: (2025)
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
por: Ghahroodi, Omid, et al.
Publicado: (2025)
por: Ghahroodi, Omid, et al.
Publicado: (2025)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
por: Zahraei, Pardis Sadat, et al.
Publicado: (2025)
por: Zahraei, Pardis Sadat, et al.
Publicado: (2025)
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
por: Sameti, Mohammad Hossein, et al.
Publicado: (2026)
por: Sameti, Mohammad Hossein, et al.
Publicado: (2026)
Persian Musical Instruments Classification Using Polyphonic Data Augmentation
por: Esfangereh, Diba Hadi, et al.
Publicado: (2025)
por: Esfangereh, Diba Hadi, et al.
Publicado: (2025)
EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
por: Nazar, Nizi, et al.
Publicado: (2025)
por: Nazar, Nizi, et al.
Publicado: (2025)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2024)
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2024)
Analyzing CLIP's Performance Limitations in Multi-Object Scenarios: A Controlled High-Resolution Study
por: Abbasi, Reza, et al.
Publicado: (2025)
por: Abbasi, Reza, et al.
Publicado: (2025)
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
por: Abbasi, Reza, et al.
Publicado: (2025)
por: Abbasi, Reza, et al.
Publicado: (2025)
TuringQ: Benchmarking AI Comprehension in Theory of Computation
por: Zahraei, Pardis Sadat, et al.
Publicado: (2024)
por: Zahraei, Pardis Sadat, et al.
Publicado: (2024)
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2025)
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2025)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
por: Zohrabi, Reihaneh, et al.
Publicado: (2025)
por: Zohrabi, Reihaneh, et al.
Publicado: (2025)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
por: Asgari, Ehsaneddin, et al.
Publicado: (2025)
por: Asgari, Ehsaneddin, et al.
Publicado: (2025)
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
por: Ahmadpour, Mohammadjavad, et al.
Publicado: (2025)
por: Ahmadpour, Mohammadjavad, et al.
Publicado: (2025)
Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
por: Dijujin, Negin Hashemi, et al.
Publicado: (2025)
por: Dijujin, Negin Hashemi, et al.
Publicado: (2025)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
por: Mofayezi, Mohammadreza, et al.
Publicado: (2024)
por: Mofayezi, Mohammadreza, et al.
Publicado: (2024)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
por: Belmadani, Ikram, et al.
Publicado: (2026)
por: Belmadani, Ikram, et al.
Publicado: (2026)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
por: Hosseini, Seyed Mohsen, et al.
Publicado: (2024)
por: Hosseini, Seyed Mohsen, et al.
Publicado: (2024)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
por: Mehri, Faridoun, et al.
Publicado: (2024)
por: Mehri, Faridoun, et al.
Publicado: (2024)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
por: Taghanaki, Saeid Asgari, et al.
Publicado: (2024)
por: Taghanaki, Saeid Asgari, et al.
Publicado: (2024)
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
por: Paqaleh, Mohammad Mahdi Samiei, et al.
Publicado: (2025)
Ejemplares similares
-
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
por: Marioriyad, Arash, et al.
Publicado: (2025) -
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
por: Marioriyad, Arash, et al.
Publicado: (2026) -
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
por: Ghahroodi, Omid, et al.
Publicado: (2024) -
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
por: Marioriyad, Arash, et al.
Publicado: (2024) -
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
por: Rezaei, Parham, et al.
Publicado: (2025)