How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Disen, Shi, Freda |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
by: Wang, Haojin, et al.
Published: (2025)
by: Wang, Haojin, et al.
Published: (2025)
Learning Language Structures through Grounding
by: Shi, Freda
Published: (2024)
by: Shi, Freda
Published: (2024)
LingGym: How Far Are LLMs from Thinking Like Field Linguists?
by: Yang, Changbing, et al.
Published: (2025)
by: Yang, Changbing, et al.
Published: (2025)
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
by: Kalamkar, Prathamesh, et al.
Published: (2025)
by: Kalamkar, Prathamesh, et al.
Published: (2025)
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
by: Kezar, Lee, et al.
Published: (2025)
by: Kezar, Lee, et al.
Published: (2025)
Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery
by: Jiang, Yifan, et al.
Published: (2026)
by: Jiang, Yifan, et al.
Published: (2026)
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
by: Zhao, Qinyu, et al.
Published: (2024)
by: Zhao, Qinyu, et al.
Published: (2024)
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
by: Abagyan, Diana, et al.
Published: (2025)
by: Abagyan, Diana, et al.
Published: (2025)
Logical forms complement probability in understanding language model (and human) performance
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Cross-Linguistic Transcription and Phonological Representation in the Huìtóngguǎnxì Huáyíyìyǔ
by: Kim, Ji-eun
Published: (2026)
by: Kim, Ji-eun
Published: (2026)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026)
by: Kraus, Oliver, et al.
Published: (2026)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
by: Suvarna, Ashima, et al.
Published: (2024)
by: Suvarna, Ashima, et al.
Published: (2024)
How Language Directions Align with Token Geometry in Multilingual LLMs
by: Kim, JaeSeong, et al.
Published: (2025)
by: Kim, JaeSeong, et al.
Published: (2025)
How does a Language-Specific Tokenizer affect LLMs?
by: Seo, Jean, et al.
Published: (2025)
by: Seo, Jean, et al.
Published: (2025)
How Proficient Are Large Language Models in Formal Languages? An In-Depth Insight for Knowledge Base Question Answering
by: Liu, Jinxin, et al.
Published: (2024)
by: Liu, Jinxin, et al.
Published: (2024)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
by: Jiang, Yifan, et al.
Published: (2026)
by: Jiang, Yifan, et al.
Published: (2026)
Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
by: Zhao, Yize, et al.
Published: (2025)
by: Zhao, Yize, et al.
Published: (2025)
How Important Is Tokenization in French Medical Masked Language Models?
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Blessing of Multilinguality: A Systematic Analysis of Multilingual In-Context Learning
by: Tu, Yilei, et al.
Published: (2025)
by: Tu, Yilei, et al.
Published: (2025)
Phonology Recognition in American Sign Language
by: Tavella, Federico, et al.
Published: (2021)
by: Tavella, Federico, et al.
Published: (2021)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
by: Arakelyan, Erik
Published: (2025)
by: Arakelyan, Erik
Published: (2025)
LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?
by: Tamo, J. Ben, et al.
Published: (2026)
by: Tamo, J. Ben, et al.
Published: (2026)
How Large Language Models Balance Internal Knowledge with User and Document Assertions
by: Li, Shuowei, et al.
Published: (2026)
by: Li, Shuowei, et al.
Published: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
by: Lan, Disen, et al.
Published: (2025)
by: Lan, Disen, et al.
Published: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
by: Wu, Juncheng, et al.
Published: (2026)
by: Wu, Juncheng, et al.
Published: (2026)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Number Cookbook: Number Understanding of Language Models and How to Improve It
by: Yang, Haotong, et al.
Published: (2024)
by: Yang, Haotong, et al.
Published: (2024)
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
by: Lin, Liuhao, et al.
Published: (2025)
by: Lin, Liuhao, et al.
Published: (2025)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
Perception of Phonological Assimilation by Neural Speech Recognition Models
by: Pouw, Charlotte, et al.
Published: (2024)
by: Pouw, Charlotte, et al.
Published: (2024)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
by: Madureira, Brielen, et al.
Published: (2022)
by: Madureira, Brielen, et al.
Published: (2022)
Large Language Models for Predictive Analysis: How Far Are They?
by: Chen, Qin, et al.
Published: (2025)
by: Chen, Qin, et al.
Published: (2025)
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
by: Bertolazzi, Leonardo, et al.
Published: (2025)
by: Bertolazzi, Leonardo, et al.
Published: (2025)
How Susceptible are Large Language Models to Ideological Manipulation?
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
Improved Representation Steering for Language Models
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?
by: Sun, Kai, et al.
Published: (2023)
by: Sun, Kai, et al.
Published: (2023)
Similar Items
-
Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
by: Wang, Haojin, et al.
Published: (2025) -
Learning Language Structures through Grounding
by: Shi, Freda
Published: (2024) -
LingGym: How Far Are LLMs from Thinking Like Field Linguists?
by: Yang, Changbing, et al.
Published: (2025) -
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
by: Kalamkar, Prathamesh, et al.
Published: (2025) -
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
by: Zhang, Zheyuan, et al.
Published: (2024)