Saved in:
| Main Authors: | Sakata, Masaki, Heinzerling, Benjamin, Yokoi, Sho, Ito, Takumi, Inui, Kentaro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.02701 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Representations of Hierarchical Concepts in Language Models
by: Sakata, Masaki, et al.
Published: (2026)
by: Sakata, Masaki, et al.
Published: (2026)
Monotonic Representation of Numeric Properties in Language Models
by: Heinzerling, Benjamin, et al.
Published: (2024)
by: Heinzerling, Benjamin, et al.
Published: (2024)
Representational Analysis of Binding in Language Models
by: Dai, Qin, et al.
Published: (2024)
by: Dai, Qin, et al.
Published: (2024)
Cell-Based Representation of Relational Binding in Language Models
by: Dai, Qin, et al.
Published: (2026)
by: Dai, Qin, et al.
Published: (2026)
The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models
by: Takahashi, Ryosuke, et al.
Published: (2024)
by: Takahashi, Ryosuke, et al.
Published: (2024)
TopK Language Models
by: Takahashi, Ryosuke, et al.
Published: (2025)
by: Takahashi, Ryosuke, et al.
Published: (2025)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
by: Kobayashi, Goro, et al.
Published: (2023)
by: Kobayashi, Goro, et al.
Published: (2023)
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
by: Hara, Tomomasa, et al.
Published: (2026)
by: Hara, Tomomasa, et al.
Published: (2026)
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference
by: Kamoda, Go, et al.
Published: (2025)
by: Kamoda, Go, et al.
Published: (2025)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation
by: Brassard, Ana, et al.
Published: (2024)
by: Brassard, Ana, et al.
Published: (2024)
Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki
by: Sasaki, Mutsumi, et al.
Published: (2025)
by: Sasaki, Mutsumi, et al.
Published: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025)
by: Nwadike, Munachiso, et al.
Published: (2025)
How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders
by: Inaba, Tatsuro, et al.
Published: (2025)
by: Inaba, Tatsuro, et al.
Published: (2025)
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
Syntactic Learnability of Echo State Neural Language Models at Scale
by: Ueda, Ryo, et al.
Published: (2025)
by: Ueda, Ryo, et al.
Published: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
by: Takishita, Sho, et al.
Published: (2025)
by: Takishita, Sho, et al.
Published: (2025)
Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted
by: Shimmei, Machi, et al.
Published: (2025)
by: Shimmei, Machi, et al.
Published: (2025)
Large Language Models Are Human-Like Internally
by: Kuribayashi, Tatsuki, et al.
Published: (2025)
by: Kuribayashi, Tatsuki, et al.
Published: (2025)
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
by: Deguchi, Hiroyuki, et al.
Published: (2025)
by: Deguchi, Hiroyuki, et al.
Published: (2025)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
by: Shing, Makoto, et al.
Published: (2025)
by: Shing, Makoto, et al.
Published: (2025)
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems
by: Sato, Shiki, et al.
Published: (2024)
by: Sato, Shiki, et al.
Published: (2024)
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
by: Yano, Kazuki, et al.
Published: (2025)
by: Yano, Kazuki, et al.
Published: (2025)
SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
by: Yoneda, Masataka, et al.
Published: (2026)
by: Yoneda, Masataka, et al.
Published: (2026)
How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
by: Chevi, Rendi, et al.
Published: (2025)
by: Chevi, Rendi, et al.
Published: (2025)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
by: Doan, Nhi Hoai, et al.
Published: (2025)
by: Doan, Nhi Hoai, et al.
Published: (2025)
Zipfian Whitening
by: Yokoi, Sho, et al.
Published: (2024)
by: Yokoi, Sho, et al.
Published: (2024)
Subspace Representations for Soft Set Operations and Sentence Similarities
by: Ishibashi, Yoichi, et al.
Published: (2022)
by: Ishibashi, Yoichi, et al.
Published: (2022)
First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model Reasoning
by: Aoki, Yoichi, et al.
Published: (2024)
by: Aoki, Yoichi, et al.
Published: (2024)
J-UniMorph: Japanese Morphological Annotation through the Universal Feature Schema
by: Matsuzaki, Kosuke, et al.
Published: (2024)
by: Matsuzaki, Kosuke, et al.
Published: (2024)
FinchGPT: a Transformer based language model for birdsong analysis
by: Kobayashi, Kosei, et al.
Published: (2025)
by: Kobayashi, Kosei, et al.
Published: (2025)
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
by: Kishino, Ryo, et al.
Published: (2024)
by: Kishino, Ryo, et al.
Published: (2024)
Nodes Are Early, Edges Are Late: Probing Diagram Representations in Large Vision-Language Models
by: Yoshida, Haruto, et al.
Published: (2026)
by: Yoshida, Haruto, et al.
Published: (2026)
Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)
Reducing the Cost: Cross-Prompt Pre-Finetuning for Short Answer Scoring
by: Funayama, Hiroaki, et al.
Published: (2024)
by: Funayama, Hiroaki, et al.
Published: (2024)
Similar Items
-
Linear Representations of Hierarchical Concepts in Language Models
by: Sakata, Masaki, et al.
Published: (2026) -
Monotonic Representation of Numeric Properties in Language Models
by: Heinzerling, Benjamin, et al.
Published: (2024) -
Representational Analysis of Binding in Language Models
by: Dai, Qin, et al.
Published: (2024) -
Cell-Based Representation of Relational Binding in Language Models
by: Dai, Qin, et al.
Published: (2026) -
The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models
by: Takahashi, Ryosuke, et al.
Published: (2024)