Evaluating The Impact of Stimulus Quality in Investigations of LLM Language Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pistotti, Timothy, Brown, Jason, Witbrock, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Gaps in the APS: Direct Minimal Pair Analysis in LLM Syntactic Assessments
von: Pistotti, Timothy, et al.
Veröffentlicht: (2025)
von: Pistotti, Timothy, et al.
Veröffentlicht: (2025)
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
von: Bao, Qiming, et al.
Veröffentlicht: (2026)
von: Bao, Qiming, et al.
Veröffentlicht: (2026)
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
von: Bao, Qiming, et al.
Veröffentlicht: (2025)
von: Bao, Qiming, et al.
Veröffentlicht: (2025)
Using Large Language Models for the Interpretation of Building Regulations
von: Fuchs, Stefan, et al.
Veröffentlicht: (2024)
von: Fuchs, Stefan, et al.
Veröffentlicht: (2024)
A Survey of Pun Generation: Datasets, Evaluations and Methodologies
von: Su, Yuchen, et al.
Veröffentlicht: (2025)
von: Su, Yuchen, et al.
Veröffentlicht: (2025)
Where Are We? Evaluating LLM Performance on African Languages
von: Adebara, Ife, et al.
Veröffentlicht: (2025)
von: Adebara, Ife, et al.
Veröffentlicht: (2025)
Investigating the Impact of Data Selection Strategies on Language Model Performance
von: Gu, Jiayao, et al.
Veröffentlicht: (2025)
von: Gu, Jiayao, et al.
Veröffentlicht: (2025)
Large Language Models as Common-Sense Heuristics
von: Borro, Andrey, et al.
Veröffentlicht: (2025)
von: Borro, Andrey, et al.
Veröffentlicht: (2025)
Psychology-Driven Enhancement of Humour Translation
von: Su, Yuchen, et al.
Veröffentlicht: (2025)
von: Su, Yuchen, et al.
Veröffentlicht: (2025)
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models
von: Yang, Xiulin, et al.
Veröffentlicht: (2026)
von: Yang, Xiulin, et al.
Veröffentlicht: (2026)
Impact of Model Size on Fine-tuned LLM Performance in Data-to-Text Generation: A State-of-the-Art Investigation
von: Mahapatra, Joy, et al.
Veröffentlicht: (2024)
von: Mahapatra, Joy, et al.
Veröffentlicht: (2024)
Test Set Quality in Multilingual LLM Evaluation
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Counterfactual Causal Inference in Natural Language with Large Language Models
von: Gendron, Gaël, et al.
Veröffentlicht: (2024)
von: Gendron, Gaël, et al.
Veröffentlicht: (2024)
Evaluating the Impact of Compression Techniques on Task-Specific Performance of Large Language Models
von: Khanal, Bishwash, et al.
Veröffentlicht: (2024)
von: Khanal, Bishwash, et al.
Veröffentlicht: (2024)
An Investigation of Linguistic Biases in LLM-Based Recommendations
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2026)
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2026)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
Language Barriers: Evaluating Cross-Lingual Performance of CNN and Transformer Architectures for Speech Quality Estimation
von: Wardah, Wafaa, et al.
Veröffentlicht: (2025)
von: Wardah, Wafaa, et al.
Veröffentlicht: (2025)
Large Language Models Are Not Strong Abstract Reasoners
von: Gendron, Gaël, et al.
Veröffentlicht: (2023)
von: Gendron, Gaël, et al.
Veröffentlicht: (2023)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
von: Kurz, Simon, et al.
Veröffentlicht: (2024)
von: Kurz, Simon, et al.
Veröffentlicht: (2024)
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
von: Martinez, Matias
Veröffentlicht: (2024)
von: Martinez, Matias
Veröffentlicht: (2024)
Arabic Dataset for LLM Safeguard Evaluation
von: Ashraf, Yasser, et al.
Veröffentlicht: (2024)
von: Ashraf, Yasser, et al.
Veröffentlicht: (2024)
Investigating the Impact of Rationales for LLMs on Natural Language Understanding
von: Shi, Wenhang, et al.
Veröffentlicht: (2025)
von: Shi, Wenhang, et al.
Veröffentlicht: (2025)
Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models
von: Su, Yuchen, et al.
Veröffentlicht: (2026)
von: Su, Yuchen, et al.
Veröffentlicht: (2026)
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
von: Bao, Qiming, et al.
Veröffentlicht: (2023)
Investigating the Effectiveness of HyperTuning via Gisting
von: Phang, Jason
Veröffentlicht: (2024)
von: Phang, Jason
Veröffentlicht: (2024)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
von: Lee, Dongryeol, et al.
Veröffentlicht: (2024)
von: Lee, Dongryeol, et al.
Veröffentlicht: (2024)
Evaluating the Diversity and Quality of LLM Generated Content
von: Shypula, Alexander, et al.
Veröffentlicht: (2025)
von: Shypula, Alexander, et al.
Veröffentlicht: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
von: Casola, Silvia, et al.
Veröffentlicht: (2025)
von: Casola, Silvia, et al.
Veröffentlicht: (2025)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
von: Sajith, Aryan, et al.
Veröffentlicht: (2024)
von: Sajith, Aryan, et al.
Veröffentlicht: (2024)
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
von: AlQadi, Leen, et al.
Veröffentlicht: (2026)
von: AlQadi, Leen, et al.
Veröffentlicht: (2026)
Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models
von: Roy, Saumya
Veröffentlicht: (2025)
von: Roy, Saumya
Veröffentlicht: (2025)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
von: Bao, Qiming, et al.
Veröffentlicht: (2022)
von: Bao, Qiming, et al.
Veröffentlicht: (2022)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025)
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025)
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
von: Chen, Yi-Pei, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Pei, et al.
Veröffentlicht: (2024)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
von: Min, Hyangsuk, et al.
Veröffentlicht: (2025)
von: Min, Hyangsuk, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring Gaps in the APS: Direct Minimal Pair Analysis in LLM Syntactic Assessments
von: Pistotti, Timothy, et al.
Veröffentlicht: (2025) -
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models
von: Bao, Qiming, et al.
Veröffentlicht: (2023) -
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
von: Bao, Qiming, et al.
Veröffentlicht: (2026) -
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning
von: Bao, Qiming, et al.
Veröffentlicht: (2023) -
Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
von: Bao, Qiming, et al.
Veröffentlicht: (2025)