On the Proper Treatment of Tokenization in Psycholinguistics
Fuente:
arXiv
Saved in:
| Main Authors: | Giulianelli, Mario, Malagutti, Luca, Gastaldi, Juan Luis, DuSell, Brian, Vieira, Tim, Cotterell, Ryan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Foundations of Tokenization: Statistical and Computational Concerns
by: Gastaldi, Juan Luis, et al.
Published: (2024)
by: Gastaldi, Juan Luis, et al.
Published: (2024)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025)
by: DuSell, Brian, et al.
Published: (2025)
Algorithms for Weighted Pushdown Automata
by: Butoi, Alexandra, et al.
Published: (2022)
by: Butoi, Alexandra, et al.
Published: (2022)
Information Locality as an Inductive Bias for Neural Language Models
by: Someya, Taiga, et al.
Published: (2025)
by: Someya, Taiga, et al.
Published: (2025)
Language Models over Canonical Byte-Pair Encodings
by: Vieira, Tim, et al.
Published: (2025)
by: Vieira, Tim, et al.
Published: (2025)
On the Proper Treatment of Units in Surprisal Theory
by: Kiegeland, Samuel, et al.
Published: (2026)
by: Kiegeland, Samuel, et al.
Published: (2026)
Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns
by: DuSell, Brian, et al.
Published: (2023)
by: DuSell, Brian, et al.
Published: (2023)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
Generalized Measures of Anticipation and Responsivity in Online Language Processing
by: Giulianelli, Mario, et al.
Published: (2024)
by: Giulianelli, Mario, et al.
Published: (2024)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
A Formal Perspective on Byte-Pair Encoding
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
On the Efficacy of Sampling Adapters
by: Meister, Clara, et al.
Published: (2023)
by: Meister, Clara, et al.
Published: (2023)
The Role of $n$-gram Smoothing in the Age of Neural Networks
by: Malagutti, Luca, et al.
Published: (2024)
by: Malagutti, Luca, et al.
Published: (2024)
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
by: Tsipidi, Eleftheria, et al.
Published: (2024)
by: Tsipidi, Eleftheria, et al.
Published: (2024)
A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior
by: Re, Francesco Ignazio, et al.
Published: (2025)
by: Re, Francesco Ignazio, et al.
Published: (2025)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025)
by: Amini, Afra, et al.
Published: (2025)
Probing for Reading Times
by: Tsipidi, Eleftheria, et al.
Published: (2026)
by: Tsipidi, Eleftheria, et al.
Published: (2026)
The Harmonic Structure of Information Contours
by: Tsipidi, Eleftheria, et al.
Published: (2025)
by: Tsipidi, Eleftheria, et al.
Published: (2025)
LanguaShrink: Reducing Token Overhead with Psycholinguistics
by: Liang, Xuechen, et al.
Published: (2024)
by: Liang, Xuechen, et al.
Published: (2024)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Syntactic Control of Language Models by Posterior Inference
by: Xefteri, Vicky, et al.
Published: (2025)
by: Xefteri, Vicky, et al.
Published: (2025)
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models
by: Frisch, Ivar, et al.
Published: (2024)
by: Frisch, Ivar, et al.
Published: (2024)
Transducing Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2026)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2026)
Prefix Parsing is Just Parsing
by: Pasti, Clemente, et al.
Published: (2026)
by: Pasti, Clemente, et al.
Published: (2026)
Towards a Similarity-adjusted Surprisal Theory
by: Meister, Clara, et al.
Published: (2024)
by: Meister, Clara, et al.
Published: (2024)
Structure-Conditional Minimum Bayes Risk Decoding
by: Eikema, Bryan, et al.
Published: (2025)
by: Eikema, Bryan, et al.
Published: (2025)
Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue
by: Utting, Tom, et al.
Published: (2026)
by: Utting, Tom, et al.
Published: (2026)
Structured Voronoi Sampling
by: Amini, Afra, et al.
Published: (2023)
by: Amini, Afra, et al.
Published: (2023)
Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields
by: Cotterell, Ryan, et al.
Published: (2024)
by: Cotterell, Ryan, et al.
Published: (2024)
Characterizing the Expressivity of Local Attention in Transformers
by: Li, Jiaoda, et al.
Published: (2026)
by: Li, Jiaoda, et al.
Published: (2026)
Exact Hard Monotonic Attention for Character-Level Transduction
by: Wu, Shijie, et al.
Published: (2019)
by: Wu, Shijie, et al.
Published: (2019)
Cross-lingual, Character-Level Neural Morphological Tagging
by: Cotterell, Ryan, et al.
Published: (2017)
by: Cotterell, Ryan, et al.
Published: (2017)
Characterizing the Expressivity of Fixed-Precision Transformer Language Models
by: Li, Jiaoda, et al.
Published: (2025)
by: Li, Jiaoda, et al.
Published: (2025)
How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
by: Clark, Christian, et al.
Published: (2025)
by: Clark, Christian, et al.
Published: (2025)
Efficiently Computing Susceptibility to Context in Language Models
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
Automating the Analysis of Parsing Algorithms (and other Dynamic Programs)
by: Vieira, Tim, et al.
Published: (2025)
by: Vieira, Tim, et al.
Published: (2025)
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
by: Vargas, Francisco, et al.
Published: (2020)
by: Vargas, Francisco, et al.
Published: (2020)
Similar Items
-
The Foundations of Tokenization: Statistical and Computational Concerns
by: Gastaldi, Juan Luis, et al.
Published: (2024) -
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024) -
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025) -
Algorithms for Weighted Pushdown Automata
by: Butoi, Alexandra, et al.
Published: (2022) -
Information Locality as an Inductive Bias for Neural Language Models
by: Someya, Taiga, et al.
Published: (2025)