Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Waldis, Andreas, Perlitz, Yotam, Choshen, Leshem, Hou, Yufang, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026)
by: Waldis, Andreas, et al.
Published: (2026)
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
by: Waldis, Andreas, et al.
Published: (2023)
by: Waldis, Andreas, et al.
Published: (2023)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
by: Falk, Neele, et al.
Published: (2024)
by: Falk, Neele, et al.
Published: (2024)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023)
by: Perlitz, Yotam, et al.
Published: (2023)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
by: Habba, Eliya, et al.
Published: (2026)
by: Habba, Eliya, et al.
Published: (2026)
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets
by: Schiller, Benjamin, et al.
Published: (2022)
by: Schiller, Benjamin, et al.
Published: (2022)
Aligned Probing: Relating Toxic Behavior and Model Internals
by: Waldis, Andreas, et al.
Published: (2025)
by: Waldis, Andreas, et al.
Published: (2025)
Transforming Scholarly Landscapes: Influence of Large Language Models on Academic Fields beyond Computer Science
by: Pramanick, Aniket, et al.
Published: (2024)
by: Pramanick, Aniket, et al.
Published: (2024)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Systematic Task Exploration with LLMs: A Study in Citation Text Generation
by: Şahinuç, Furkan, et al.
Published: (2024)
by: Şahinuç, Furkan, et al.
Published: (2024)
Missci: Reconstructing Fallacies in Misrepresented Science
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
by: Ashury-Tahan, Shir, et al.
Published: (2025)
by: Ashury-Tahan, Shir, et al.
Published: (2025)
The Nature of NLP: Analyzing Contributions in NLP Papers
by: Pramanick, Aniket, et al.
Published: (2024)
by: Pramanick, Aniket, et al.
Published: (2024)
ClaimFlow: Tracing the Evolution of Scientific Claims in NLP
by: Pramanick, Aniket, et al.
Published: (2026)
by: Pramanick, Aniket, et al.
Published: (2026)
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
by: Şahinuç, Furkan, et al.
Published: (2024)
by: Şahinuç, Furkan, et al.
Published: (2024)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
A Pipeline to Assess Merging Methods via Behavior and Internals
by: Sigrist, Yutaro, et al.
Published: (2025)
by: Sigrist, Yutaro, et al.
Published: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
by: Akyürek, Afra Feyza, et al.
Published: (2024)
by: Akyürek, Afra Feyza, et al.
Published: (2024)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
Attribute or Abstain: Large Language Models as Long Document Assistants
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
by: Paul, Indraneil, et al.
Published: (2024)
by: Paul, Indraneil, et al.
Published: (2024)
DAPR: A Benchmark on Document-Aware Passage Retrieval
by: Wang, Kexin, et al.
Published: (2023)
by: Wang, Kexin, et al.
Published: (2023)
Naturally Occurring Feedback is Common, Extractable and Useful
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Cultural Learning-Based Culture Adaptation of Language Models
by: Liu, Chen Cecilia, et al.
Published: (2025)
by: Liu, Chen Cecilia, et al.
Published: (2025)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)
by: Dycke, Nils, et al.
Published: (2025)
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification
by: Bates, Luke, et al.
Published: (2023)
by: Bates, Luke, et al.
Published: (2023)
Citation Failure: Definition, Analysis and Efficient Mitigation
by: Buchmann, Jan, et al.
Published: (2025)
by: Buchmann, Jan, et al.
Published: (2025)
Patches of Nonlinearity: Instruction Vectors in Large Language Models
by: Bigoulaeva, Irina, et al.
Published: (2026)
by: Bigoulaeva, Irina, et al.
Published: (2026)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
by: Bates, Luke, et al.
Published: (2025)
by: Bates, Luke, et al.
Published: (2025)
Similar Items
-
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026) -
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
by: Waldis, Andreas, et al.
Published: (2023) -
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024) -
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
by: Falk, Neele, et al.
Published: (2024) -
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
by: Waldis, Andreas, et al.
Published: (2024)