Towards a Similarity-adjusted Surprisal Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Meister, Clara, Giulianelli, Mario, Pimentel, Tiago |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Testing the Predictions of Surprisal Theory in 11 Languages
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)
How to Compute the Probability of a Word
by: Pimentel, Tiago, et al.
Published: (2024)
by: Pimentel, Tiago, et al.
Published: (2024)
Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue
by: Utting, Tom, et al.
Published: (2026)
by: Utting, Tom, et al.
Published: (2026)
What Language is This? Ask Your Tokenizer
by: Meister, Clara, et al.
Published: (2026)
by: Meister, Clara, et al.
Published: (2026)
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
by: Tsipidi, Eleftheria, et al.
Published: (2024)
by: Tsipidi, Eleftheria, et al.
Published: (2024)
Locally Typical Sampling
by: Meister, Clara, et al.
Published: (2022)
by: Meister, Clara, et al.
Published: (2022)
On the Efficacy of Sampling Adapters
by: Meister, Clara, et al.
Published: (2023)
by: Meister, Clara, et al.
Published: (2023)
Analyzing Wrap-Up Effects through an Information-Theoretic Lens
by: Meister, Clara, et al.
Published: (2022)
by: Meister, Clara, et al.
Published: (2022)
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025)
by: Lesci, Pietro, et al.
Published: (2025)
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models
by: Frisch, Ivar, et al.
Published: (2024)
by: Frisch, Ivar, et al.
Published: (2024)
Structure-Conditional Minimum Bayes Risk Decoding
by: Eikema, Bryan, et al.
Published: (2025)
by: Eikema, Bryan, et al.
Published: (2025)
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?
by: Gay, Matteo, et al.
Published: (2026)
by: Gay, Matteo, et al.
Published: (2026)
On the Proper Treatment of Units in Surprisal Theory
by: Kiegeland, Samuel, et al.
Published: (2026)
by: Kiegeland, Samuel, et al.
Published: (2026)
Generalized Measures of Anticipation and Responsivity in Online Language Processing
by: Giulianelli, Mario, et al.
Published: (2024)
by: Giulianelli, Mario, et al.
Published: (2024)
Back to Bytes: Revisiting Tokenization Through UTF-8
by: Moryossef, Amit, et al.
Published: (2025)
by: Moryossef, Amit, et al.
Published: (2025)
Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
by: Nair, Sathvik, et al.
Published: (2026)
by: Nair, Sathvik, et al.
Published: (2026)
FunQA: Towards Surprising Video Comprehension
by: Xie, Binzhu, et al.
Published: (2023)
by: Xie, Binzhu, et al.
Published: (2023)
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
by: Wilcox, Ethan Gotlieb, et al.
Published: (2025)
by: Wilcox, Ethan Gotlieb, et al.
Published: (2025)
Formal Aspects of Language Modeling
by: Cotterell, Ryan, et al.
Published: (2023)
by: Cotterell, Ryan, et al.
Published: (2023)
On the Proper Treatment of Tokenization in Psycholinguistics
by: Giulianelli, Mario, et al.
Published: (2024)
by: Giulianelli, Mario, et al.
Published: (2024)
Tokenisation is NP-Complete
by: Whittington, Philip, et al.
Published: (2024)
by: Whittington, Philip, et al.
Published: (2024)
Convergence and Divergence of Language Models under Different Random Seeds
by: Fehlauer, Finlay, et al.
Published: (2025)
by: Fehlauer, Finlay, et al.
Published: (2025)
Investigating Critical Period Effects in Language Acquisition through Neural Language Models
by: Constantinescu, Ionut, et al.
Published: (2024)
by: Constantinescu, Ionut, et al.
Published: (2024)
Local and Global Decoding in Text Generation
by: Gareev, Daniel, et al.
Published: (2024)
by: Gareev, Daniel, et al.
Published: (2024)
Information Locality as an Inductive Bias for Neural Language Models
by: Someya, Taiga, et al.
Published: (2025)
by: Someya, Taiga, et al.
Published: (2025)
Probing for Reading Times
by: Tsipidi, Eleftheria, et al.
Published: (2026)
by: Tsipidi, Eleftheria, et al.
Published: (2026)
A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior
by: Re, Francesco Ignazio, et al.
Published: (2025)
by: Re, Francesco Ignazio, et al.
Published: (2025)
Diversidade linguística e inclusão digital: desafios para uma ia brasileira
by: Freitag, Raquel Meister Ko
Published: (2024)
by: Freitag, Raquel Meister Ko
Published: (2024)
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
The Harmonic Structure of Information Contours
by: Tsipidi, Eleftheria, et al.
Published: (2025)
by: Tsipidi, Eleftheria, et al.
Published: (2025)
Speakers Fill Lexical Semantic Gaps with Context
by: Pimentel, Tiago, et al.
Published: (2020)
by: Pimentel, Tiago, et al.
Published: (2020)
The Role of $n$-gram Smoothing in the Age of Neural Networks
by: Malagutti, Luca, et al.
Published: (2024)
by: Malagutti, Luca, et al.
Published: (2024)
Expect the Unexpected? Testing the Surprisal of Salient Entities
by: Lin, Jessica, et al.
Published: (2026)
by: Lin, Jessica, et al.
Published: (2026)
Probing for the Usage of Grammatical Number
by: Lasri, Karim, et al.
Published: (2022)
by: Lasri, Karim, et al.
Published: (2022)
Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor
by: Ma, Yuxi, et al.
Published: (2026)
by: Ma, Yuxi, et al.
Published: (2026)
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
by: Momen, Omar, et al.
Published: (2026)
by: Momen, Omar, et al.
Published: (2026)
Glitter: Visualizing Lexical Surprisal for Readability in Administrative Texts
by: Černý, Jan, et al.
Published: (2026)
by: Černý, Jan, et al.
Published: (2026)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
by: Momentè, Filippo, et al.
Published: (2025)
by: Momentè, Filippo, et al.
Published: (2025)
What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
by: Yadavalli, Aditya, et al.
Published: (2025)
by: Yadavalli, Aditya, et al.
Published: (2025)
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
by: Oh, Byung-Doh, et al.
Published: (2024)
by: Oh, Byung-Doh, et al.
Published: (2024)
Similar Items
-
Testing the Predictions of Surprisal Theory in 11 Languages
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023) -
How to Compute the Probability of a Word
by: Pimentel, Tiago, et al.
Published: (2024) -
Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue
by: Utting, Tom, et al.
Published: (2026) -
What Language is This? Ask Your Tokenizer
by: Meister, Clara, et al.
Published: (2026) -
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
by: Tsipidi, Eleftheria, et al.
Published: (2024)