Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
Fuente:
arXiv
Saved in:
| Main Authors: | Weissweiler, Leonie, Mahowald, Kyle, Goldberg, Adele |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A suite of LMs comprehend puzzle statements as well as humans
by: Goldberg, Adele E, et al.
Published: (2025)
by: Goldberg, Adele E, et al.
Published: (2025)
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
by: Boguraev, Sasha, et al.
Published: (2024)
by: Boguraev, Sasha, et al.
Published: (2024)
Constructions are Revealed in Word Distributions
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
by: Yao, Qing, et al.
Published: (2025)
by: Yao, Qing, et al.
Published: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
by: Boguraev, Sasha, et al.
Published: (2026)
by: Boguraev, Sasha, et al.
Published: (2026)
How Linguistics Learned to Stop Worrying and Love the Language Models
by: Futrell, Richard, et al.
Published: (2025)
by: Futrell, Richard, et al.
Published: (2025)
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
by: Jumelet, Jaap, et al.
Published: (2025)
by: Jumelet, Jaap, et al.
Published: (2025)
BabyLM's First Constructions: Causal probing provides a signal of learning
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
For Generated Text, Is NLI-Neutral Text the Best Text?
by: Mersinias, Michail, et al.
Published: (2023)
by: Mersinias, Michail, et al.
Published: (2023)
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
by: Mortensen, David R., et al.
Published: (2024)
by: Mortensen, David R., et al.
Published: (2024)
You Can't Fight in Here! This is BBS!
by: Futrell, Richard, et al.
Published: (2026)
by: Futrell, Richard, et al.
Published: (2026)
Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
by: Misra, Kanishka, et al.
Published: (2024)
by: Misra, Kanishka, et al.
Published: (2024)
Are Language Models More Like Libraries or Like Librarians? Bibliotechnism, the Novel Reference Problem, and the Attitudes of LLMs
by: Lederman, Harvey, et al.
Published: (2024)
by: Lederman, Harvey, et al.
Published: (2024)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models
by: Drucker, Daniel, et al.
Published: (2026)
by: Drucker, Daniel, et al.
Published: (2026)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
by: Hofmann, Valentin, et al.
Published: (2024)
by: Hofmann, Valentin, et al.
Published: (2024)
Participle-Prepended Nominals Have Lower Entropy Than Nominals Appended After the Participle
by: Denlinger, Kristie, et al.
Published: (2024)
by: Denlinger, Kristie, et al.
Published: (2024)
France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
by: Boguraev, Sasha, et al.
Published: (2026)
by: Boguraev, Sasha, et al.
Published: (2026)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
by: Scivetti, Wesley, et al.
Published: (2026)
by: Scivetti, Wesley, et al.
Published: (2026)
SYNTHEVAL: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists
by: Zhao, Raoyuan, et al.
Published: (2024)
by: Zhao, Raoyuan, et al.
Published: (2024)
Convergence and Divergence of Language Models under Different Random Seeds
by: Fehlauer, Finlay, et al.
Published: (2025)
by: Fehlauer, Finlay, et al.
Published: (2025)
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
by: Boguraev, Sasha, et al.
Published: (2025)
by: Boguraev, Sasha, et al.
Published: (2025)
Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
by: Misra, Kanishka, et al.
Published: (2024)
by: Misra, Kanishka, et al.
Published: (2024)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
by: Rozner, Josh, et al.
Published: (2021)
by: Rozner, Josh, et al.
Published: (2021)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs
by: Rakshit, Supantho, et al.
Published: (2025)
by: Rakshit, Supantho, et al.
Published: (2025)
On Language Models' Sensitivity to Suspicious Coincidences
by: Padmanabhan, Sriram, et al.
Published: (2025)
by: Padmanabhan, Sriram, et al.
Published: (2025)
Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
by: Govindarajan, Venkata S, et al.
Published: (2023)
by: Govindarajan, Venkata S, et al.
Published: (2023)
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces
by: Ranganathan, Jwalanthi, et al.
Published: (2025)
by: Ranganathan, Jwalanthi, et al.
Published: (2025)
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
by: Gritta, Milan, et al.
Published: (2024)
by: Gritta, Milan, et al.
Published: (2024)
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias
by: Govindarajan, Venkata S, et al.
Published: (2023)
by: Govindarajan, Venkata S, et al.
Published: (2023)
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
Do they mean 'us'? Interpreting Referring Expressions in Intergroup Bias
by: Govindarajan, Venkata S, et al.
Published: (2024)
by: Govindarajan, Venkata S, et al.
Published: (2024)
For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies
by: Cuneo, Nicole, et al.
Published: (2025)
by: Cuneo, Nicole, et al.
Published: (2025)
Are BabyLMs Second Language Learners?
by: Edman, Lukas, et al.
Published: (2024)
by: Edman, Lukas, et al.
Published: (2024)
Language models align with human judgments on key grammatical constructions
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
by: Nemitz, Jonathan, et al.
Published: (2026)
by: Nemitz, Jonathan, et al.
Published: (2026)
Similar Items
-
A suite of LMs comprehend puzzle statements as well as humans
by: Goldberg, Adele E, et al.
Published: (2025) -
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
by: Boguraev, Sasha, et al.
Published: (2024) -
Constructions are Revealed in Word Distributions
by: Rozner, Joshua, et al.
Published: (2025) -
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
by: Yao, Qing, et al.
Published: (2025) -
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)