How Do Language Models Acquire Character-Level Information?
Fuente:
arXiv
Saved in:
| Main Authors: | Sato, Soma, Sasano, Ryohei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
by: Sato, Soma, et al.
Published: (2024)
by: Sato, Soma, et al.
Published: (2024)
On Representational Dissociation of Language and Arithmetic in Large Language Models
by: Kisako, Riku, et al.
Published: (2025)
by: Kisako, Riku, et al.
Published: (2025)
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
by: Utami, Nabelanita, et al.
Published: (2026)
by: Utami, Nabelanita, et al.
Published: (2026)
Ruri: Japanese General Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2024)
by: Tsukagoshi, Hayato, et al.
Published: (2024)
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2025)
by: Tsukagoshi, Hayato, et al.
Published: (2025)
FrameEOL: Semantic Frame Induction using Causal Language Models
by: Yano, Chihiro, et al.
Published: (2025)
by: Yano, Chihiro, et al.
Published: (2025)
When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression
by: Kisako, Riku, et al.
Published: (2026)
by: Kisako, Riku, et al.
Published: (2026)
CiMaTe: Citation Count Prediction Effectively Leveraging the Main Text
by: Hirako, Jun, et al.
Published: (2024)
by: Hirako, Jun, et al.
Published: (2024)
Verifying Claims About Metaphors with Large-Scale Automatic Metaphor Identification
by: Aono, Kotaro, et al.
Published: (2024)
by: Aono, Kotaro, et al.
Published: (2024)
Are Social Sentiments Inherent in LLMs? An Empirical Study on Extraction of Inter-demographic Sentiments
by: Tanaka, Kunitomo, et al.
Published: (2024)
by: Tanaka, Kunitomo, et al.
Published: (2024)
Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
by: Sugiura, Naoya, et al.
Published: (2025)
by: Sugiura, Naoya, et al.
Published: (2025)
Sentence Representations via Gaussian Embedding
by: Yoda, Shohei, et al.
Published: (2023)
by: Yoda, Shohei, et al.
Published: (2023)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
by: Chang, Hoyeon, et al.
Published: (2024)
by: Chang, Hoyeon, et al.
Published: (2024)
To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
by: Ishizuki, Yukiko, et al.
Published: (2024)
by: Ishizuki, Yukiko, et al.
Published: (2024)
Simplifying Translations for Children: Iterative Simplification Considering Age of Acquisition with LLMs
by: Oshika, Masashi, et al.
Published: (2024)
by: Oshika, Masashi, et al.
Published: (2024)
WikiSplit++: Easy Data Refinement for Split and Rephrase
by: Tsukagoshi, Hayato, et al.
Published: (2024)
by: Tsukagoshi, Hayato, et al.
Published: (2024)
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
by: Yang, Zhipeng, et al.
Published: (2026)
by: Yang, Zhipeng, et al.
Published: (2026)
CharacterBench: Benchmarking Character Customization of Large Language Models
by: Zhou, Jinfeng, et al.
Published: (2024)
by: Zhou, Jinfeng, et al.
Published: (2024)
Acquiring Bidirectionality via Large and Small Language Models
by: Goto, Takumi, et al.
Published: (2024)
by: Goto, Takumi, et al.
Published: (2024)
Evaluating Language Model Character Traits
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
How Do Multilingual Language Models Remember Facts?
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Acquiring Common Chinese Emotional Events Using Large Language Model
by: Wang, Ya, et al.
Published: (2025)
by: Wang, Ya, et al.
Published: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
by: Yuan, Xinfeng, et al.
Published: (2024)
by: Yuan, Xinfeng, et al.
Published: (2024)
How Do Language Models Compose Functions?
by: Khandelwal, Apoorv, et al.
Published: (2025)
by: Khandelwal, Apoorv, et al.
Published: (2025)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
Understanding the Ability of LLMs to Handle Character-Level Perturbation
by: Zhuo, Anyuan, et al.
Published: (2025)
by: Zhuo, Anyuan, et al.
Published: (2025)
Exact Hard Monotonic Attention for Character-Level Transduction
by: Wu, Shijie, et al.
Published: (2019)
by: Wu, Shijie, et al.
Published: (2019)
Cross-lingual, Character-Level Neural Morphological Tagging
by: Cotterell, Ryan, et al.
Published: (2017)
by: Cotterell, Ryan, et al.
Published: (2017)
Hard Non-Monotonic Attention for Character-Level Transduction
by: Wu, Shijie, et al.
Published: (2018)
by: Wu, Shijie, et al.
Published: (2018)
SpeLLM: Character-Level Multi-Head Decoding
by: Ben-Artzy, Amit, et al.
Published: (2025)
by: Ben-Artzy, Amit, et al.
Published: (2025)
How Well Do Large Language Models Disambiguate Swedish Words?
by: Johansson, Richard
Published: (2024)
by: Johansson, Richard
Published: (2024)
Character-Level Chinese Dependency Parsing via Modeling Latent Intra-Word Structure
by: Hou, Yang, et al.
Published: (2024)
by: Hou, Yang, et al.
Published: (2024)
Do Language Models Update their Forecasts with New Information?
by: Yuan, Zhangdie, et al.
Published: (2025)
by: Yuan, Zhangdie, et al.
Published: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
by: Ou, Yixin, et al.
Published: (2025)
by: Ou, Yixin, et al.
Published: (2025)
Large Language Models Lack Understanding of Character Composition of Words
by: Shin, Andrew, et al.
Published: (2024)
by: Shin, Andrew, et al.
Published: (2024)
Improving Language and Modality Transfer in Translation by Character-level Modeling
by: Tsiamas, Ioannis, et al.
Published: (2025)
by: Tsiamas, Ioannis, et al.
Published: (2025)
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
by: Ponwitayarat, Wuttikorn, et al.
Published: (2025)
by: Ponwitayarat, Wuttikorn, et al.
Published: (2025)
(How) Do Language Models Track State?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Similar Items
-
Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples
by: Sato, Soma, et al.
Published: (2024) -
On Representational Dissociation of Language and Arithmetic in Large Language Models
by: Kisako, Riku, et al.
Published: (2025) -
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
by: Utami, Nabelanita, et al.
Published: (2026) -
Ruri: Japanese General Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2024) -
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
by: Tsukagoshi, Hayato, et al.
Published: (2025)