The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cosma, Adrian, Ruseti, Stefan, Radoi, Emilian, Dascalu, Mihai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Language Models with homotokens Leads to Delayed Overfitting
by: Cosma, Adrian, et al.
Published: (2026)
by: Cosma, Adrian, et al.
Published: (2026)
How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
Value-Aware Numerical Representations for Transformer Language Models
by: Dutulescu, Andreea, et al.
Published: (2026)
by: Dutulescu, Andreea, et al.
Published: (2026)
A Retrieval-Based Approach to Medical Procedure Matching in Romanian
by: Niculae, Andrei, et al.
Published: (2025)
by: Niculae, Andrei, et al.
Published: (2025)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
by: Cosma, Adrian, et al.
Published: (2026)
by: Cosma, Adrian, et al.
Published: (2026)
Neural Grammatical Error Correction for Romanian
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
RoMath: A Mathematical Reasoning Benchmark in Romanian
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian
by: Niculae, Andrei, et al.
Published: (2025)
by: Niculae, Andrei, et al.
Published: (2025)
On Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition
by: Cosma, Adrian, et al.
Published: (2025)
by: Cosma, Adrian, et al.
Published: (2025)
The Paradox of Motion: Evidence for Spurious Correlations in Skeleton-based Gait Recognition Models
by: Cătrună, Andy, et al.
Published: (2024)
by: Cătrună, Andy, et al.
Published: (2024)
Spatial Colour Mixing Illusions as a Perception Stress Test for Vision-Language Models
by: Basoc, Nicoleta-Nina, et al.
Published: (2026)
by: Basoc, Nicoleta-Nina, et al.
Published: (2026)
MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
by: Cǎtrunǎ, Andy, et al.
Published: (2025)
by: Cǎtrunǎ, Andy, et al.
Published: (2025)
CrossGaze: A Strong Method for 3D Gaze Estimation in the Wild
by: Cătrună, Andy, et al.
Published: (2024)
by: Cătrună, Andy, et al.
Published: (2024)
GaitPT: Skeletons Are All You Need For Gait Recognition
by: Catruna, Andy, et al.
Published: (2023)
by: Catruna, Andy, et al.
Published: (2023)
Database-Agnostic Gait Enrollment using SetTransformers
by: Basoc, Nicoleta, et al.
Published: (2025)
by: Basoc, Nicoleta, et al.
Published: (2025)
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
by: Masala, Mihai, et al.
Published: (2026)
by: Masala, Mihai, et al.
Published: (2026)
An In-Vitro Study on Cross-Lingual Generalization in Language Models
by: Cosma, Adrian
Published: (2026)
by: Cosma, Adrian
Published: (2026)
Aligning Actions and Walking to LLM-Generated Textual Descriptions
by: Chivereanu, Radu, et al.
Published: (2024)
by: Chivereanu, Radu, et al.
Published: (2024)
Gait Recognition from Highly Compressed Videos
by: Niculae, Andrei, et al.
Published: (2024)
by: Niculae, Andrei, et al.
Published: (2024)
RoCode: A Dataset for Measuring Code Intelligence from Problem Definitions in Romanian
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
by: Yuan, Xinfeng, et al.
Published: (2024)
by: Yuan, Xinfeng, et al.
Published: (2024)
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
by: Yang, Zhipeng, et al.
Published: (2026)
by: Yang, Zhipeng, et al.
Published: (2026)
Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
by: Xu, Zhu, et al.
Published: (2024)
by: Xu, Zhu, et al.
Published: (2024)
Large Language Models Lack Understanding of Character Composition of Words
by: Shin, Andrew, et al.
Published: (2024)
by: Shin, Andrew, et al.
Published: (2024)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
Improving Language and Modality Transfer in Translation by Character-level Modeling
by: Tsiamas, Ioannis, et al.
Published: (2025)
by: Tsiamas, Ioannis, et al.
Published: (2025)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering
by: McNichols, Hunter, et al.
Published: (2026)
by: McNichols, Hunter, et al.
Published: (2026)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
CharacterBench: Benchmarking Character Customization of Large Language Models
by: Zhou, Jinfeng, et al.
Published: (2024)
by: Zhou, Jinfeng, et al.
Published: (2024)
Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space
by: Rom, Aviad, et al.
Published: (2024)
by: Rom, Aviad, et al.
Published: (2024)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025)
by: Dolga, Rares, et al.
Published: (2025)
OpenLLM-Ro -- Technical Report on Open-source Romanian LLMs
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Evaluating Language Model Character Traits
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
Dialectal and Low-Resource Machine Translation for Aromanian
by: Jerpelea, Alexandru-Iulius, et al.
Published: (2024)
by: Jerpelea, Alexandru-Iulius, et al.
Published: (2024)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
From Tokens To Agents: A Researcher's Guide To Understanding Large Language Models
by: Barolo, Daniele
Published: (2026)
by: Barolo, Daniele
Published: (2026)
Similar Items
-
Training Language Models with homotokens Leads to Delayed Overfitting
by: Cosma, Adrian, et al.
Published: (2026) -
How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics
by: Cosma, Adrian, et al.
Published: (2024) -
Value-Aware Numerical Representations for Transformer Language Models
by: Dutulescu, Andreea, et al.
Published: (2026) -
A Retrieval-Based Approach to Medical Procedure Matching in Romanian
by: Niculae, Andrei, et al.
Published: (2025) -
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
by: Cosma, Adrian, et al.
Published: (2026)