Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Momen, Omar, Sitter, Emilie, Herrmann, Berenike, Zarrieß, Sina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
by: Momen, Omar, et al.
Published: (2026)
by: Momen, Omar, et al.
Published: (2026)
Mehr Text! Mehr Daten! Das Kreativitätsduell als transmodales Game with a Purpose zur Datenerhebung für die textwissenschaftliche Kreativitätsforschung.
by: Sitter, Emilie, et al.
Published: (2026)
by: Sitter, Emilie, et al.
Published: (2026)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis
by: Zimmerman, Fred
Published: (2026)
by: Zimmerman, Fred
Published: (2026)
Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
by: Padovani, Francesca, et al.
Published: (2025)
by: Padovani, Francesca, et al.
Published: (2025)
The Gray Area: Characterizing Moderator Disagreement on Reddit
by: Alipour, Shayan, et al.
Published: (2026)
by: Alipour, Shayan, et al.
Published: (2026)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
The InviTE Corpus: Annotating Invectives in Tudor English Texts for Computational Modeling
by: Spliethoff, Sophie, et al.
Published: (2025)
by: Spliethoff, Sophie, et al.
Published: (2025)
DEPLAIN: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification
by: Stodden, Regina, et al.
Published: (2023)
by: Stodden, Regina, et al.
Published: (2023)
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language
by: Tanaka-Ishii, Kumiko
Published: (2026)
by: Tanaka-Ishii, Kumiko
Published: (2026)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
Child-directed speech facilitates production, not comprehension, in BabyLMs
by: Bunzeck, Bastian, et al.
Published: (2026)
by: Bunzeck, Bastian, et al.
Published: (2026)
Resilience through Scene Context in Visual Referring Expression Generation
by: Junker, Simeon, et al.
Published: (2024)
by: Junker, Simeon, et al.
Published: (2024)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
by: Sieker, Judith, et al.
Published: (2026)
by: Sieker, Judith, et al.
Published: (2026)
SceneGram: Conceptualizing and Describing Tangrams in Scene Context
by: Junker, Simeon, et al.
Published: (2025)
by: Junker, Simeon, et al.
Published: (2025)
Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generation
by: Lippmann, Philip, et al.
Published: (2025)
by: Lippmann, Philip, et al.
Published: (2025)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
by: Zhang, Ming, et al.
Published: (2026)
by: Zhang, Ming, et al.
Published: (2026)
Cost-aware LLM-based Online Dataset Annotation
by: Elumar, Eray Can, et al.
Published: (2025)
by: Elumar, Eray Can, et al.
Published: (2025)
A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews
by: Achkar, Pierre, et al.
Published: (2026)
by: Achkar, Pierre, et al.
Published: (2026)
Generalized Score Matching: Bridging $f$-Divergence and Statistical Estimation Under Correlated Noise
by: Shen, Yirong, et al.
Published: (2025)
by: Shen, Yirong, et al.
Published: (2025)
Efficient Scientific Full Text Classification: The Case of EICAT Impact Assessments
by: Brinner, Marc Felix, et al.
Published: (2025)
by: Brinner, Marc Felix, et al.
Published: (2025)
The Medical Metaphors Corpus (MCC)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
Linguistic Structure Induction from Language Models
by: Momen, Omar
Published: (2024)
by: Momen, Omar
Published: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Beyond Single-Dimension Novelty: How Combinations of Theory, Method, and Results-based Novelty Shape Scientific Impact
by: Zhao, Yi, et al.
Published: (2026)
by: Zhao, Yi, et al.
Published: (2026)
Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval
by: Knappich, Valentin, et al.
Published: (2026)
by: Knappich, Valentin, et al.
Published: (2026)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
by: Bozorgkhoo, Amirhossein, et al.
Published: (2026)
by: Bozorgkhoo, Amirhossein, et al.
Published: (2026)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
by: da Silva, Italo Luis, et al.
Published: (2025)
by: da Silva, Italo Luis, et al.
Published: (2025)
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
by: Sieker, Judith, et al.
Published: (2025)
by: Sieker, Judith, et al.
Published: (2025)
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus
by: Gui, Honghao, et al.
Published: (2024)
by: Gui, Honghao, et al.
Published: (2024)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas
by: Bunzeck, Bastian, et al.
Published: (2024)
by: Bunzeck, Bastian, et al.
Published: (2024)
Similar Items
-
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
by: Momen, Omar, et al.
Published: (2026) -
Mehr Text! Mehr Daten! Das Kreativitätsduell als transmodales Game with a Purpose zur Datenerhebung für die textwissenschaftliche Kreativitätsforschung.
by: Sitter, Emilie, et al.
Published: (2026) -
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
by: Brinner, Marc, et al.
Published: (2025) -
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
by: Brinner, Marc, et al.
Published: (2025) -
Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis
by: Zimmerman, Fred
Published: (2026)