LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bavaresco, Anna, Bernardi, Raffaella, Bertolazzi, Leonardo, Elliott, Desmond, Fernández, Raquel, Gatt, Albert, Ghaleb, Esam, Giulianelli, Mario, Hanna, Michael, Koller, Alexander, Martins, André F. T., Mondorf, Philipp, Neplenbroek, Vera, Pezzelle, Sandro, Plank, Barbara, Schlangen, David, Suglia, Alessandro, Surikuchi, Aditya K, Takmaz, Ece, Testoni, Alberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
von: Takmaz, Ece, et al.
Veröffentlicht: (2024)
von: Takmaz, Ece, et al.
Veröffentlicht: (2024)
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2025)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2025)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
von: Horst, Nicola, et al.
Veröffentlicht: (2025)
von: Horst, Nicola, et al.
Veröffentlicht: (2025)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
von: Mazzaccara, Davide, et al.
Veröffentlicht: (2024)
von: Mazzaccara, Davide, et al.
Veröffentlicht: (2024)
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
von: Widhoelzl, Hanna-Sophia, et al.
Veröffentlicht: (2024)
von: Widhoelzl, Hanna-Sophia, et al.
Veröffentlicht: (2024)
Vision-Language Models Align with Human Neural Representations in Concept Processing
von: Bavaresco, Anna, et al.
Veröffentlicht: (2024)
von: Bavaresco, Anna, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
von: Bavaresco, A., et al.
Veröffentlicht: (2024)
von: Bavaresco, A., et al.
Veröffentlicht: (2024)
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
Teaching Small Language Models to Learn Logic through Meta-Learning
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
von: Orth, Jasmin, et al.
Veröffentlicht: (2025)
von: Orth, Jasmin, et al.
Veröffentlicht: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
von: Rabern, Brian, et al.
Veröffentlicht: (2026)
von: Rabern, Brian, et al.
Veröffentlicht: (2026)
A Dialogue Game for Eliciting Balanced Collaboration
von: Jeknić, Isidora, et al.
Veröffentlicht: (2024)
von: Jeknić, Isidora, et al.
Veröffentlicht: (2024)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
von: Liu, Lanmiao, et al.
Veröffentlicht: (2026)
von: Liu, Lanmiao, et al.
Veröffentlicht: (2026)
SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
von: Liu, Lanmiao, et al.
Veröffentlicht: (2025)
von: Liu, Lanmiao, et al.
Veröffentlicht: (2025)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
von: Ghaleb, Esam, et al.
Veröffentlicht: (2025)
von: Ghaleb, Esam, et al.
Veröffentlicht: (2025)
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
von: Cheng, Xinyuan, et al.
Veröffentlicht: (2026)
von: Cheng, Xinyuan, et al.
Veröffentlicht: (2026)
O clima ético das organizações e a temática do meio ambiente
von: Marco Bertolazzi
Veröffentlicht: (2011)
von: Marco Bertolazzi
Veröffentlicht: (2011)
Are formal and functional linguistic mechanisms dissociated in language models?
von: Hanna, Michael, et al.
Veröffentlicht: (2025)
von: Hanna, Michael, et al.
Veröffentlicht: (2025)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance
von: Huang, Yunchong, et al.
Veröffentlicht: (2026)
von: Huang, Yunchong, et al.
Veröffentlicht: (2026)
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models
von: Nakajima, Kumiko, et al.
Veröffentlicht: (2026)
von: Nakajima, Kumiko, et al.
Veröffentlicht: (2026)
Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST!
von: Wildenburg, Frank, et al.
Veröffentlicht: (2024)
von: Wildenburg, Frank, et al.
Veröffentlicht: (2024)
They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse
von: Paci, Walter, et al.
Veröffentlicht: (2025)
von: Paci, Walter, et al.
Veröffentlicht: (2025)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
von: Chen, Xinyi, et al.
Veröffentlicht: (2023)
von: Chen, Xinyi, et al.
Veröffentlicht: (2023)
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025) -
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025) -
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
von: Takmaz, Ece, et al.
Veröffentlicht: (2024) -
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024) -
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)