How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Inaba, Tatsuro, Kamoda, Go, Inui, Kentaro, Isonuma, Masaru, Miyao, Yusuke, Oseki, Yohei, Heinzerling, Benjamin, Takagi, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference
von: Kamoda, Go, et al.
Veröffentlicht: (2025)
von: Kamoda, Go, et al.
Veröffentlicht: (2025)
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
Exclusive Unlearning
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2026)
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2026)
TopK Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2025)
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2025)
The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2024)
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2024)
Monotonic Representation of Numeric Properties in Language Models
von: Heinzerling, Benjamin, et al.
Veröffentlicht: (2024)
von: Heinzerling, Benjamin, et al.
Veröffentlicht: (2024)
Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
von: Harada, Yuto, et al.
Veröffentlicht: (2025)
von: Harada, Yuto, et al.
Veröffentlicht: (2025)
Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2025)
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2025)
Representational Analysis of Binding in Language Models
von: Dai, Qin, et al.
Veröffentlicht: (2024)
von: Dai, Qin, et al.
Veröffentlicht: (2024)
Cell-Based Representation of Relational Binding in Language Models
von: Dai, Qin, et al.
Veröffentlicht: (2026)
von: Dai, Qin, et al.
Veröffentlicht: (2026)
Large Language Models Are Human-Like Internally
von: Kuribayashi, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kuribayashi, Tatsuki, et al.
Veröffentlicht: (2025)
Unlearning Traces the Influential Training Data of Language Models
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
Linear Representations of Hierarchical Concepts in Language Models
von: Sakata, Masaki, et al.
Veröffentlicht: (2026)
von: Sakata, Masaki, et al.
Veröffentlicht: (2026)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
Instability in Downstream Task Performance During LLM Pretraining
von: Nishida, Yuto, et al.
Veröffentlicht: (2025)
von: Nishida, Yuto, et al.
Veröffentlicht: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
von: Ye, Mengyu, et al.
Veröffentlicht: (2025)
von: Ye, Mengyu, et al.
Veröffentlicht: (2025)
What's New in My Data? Novelty Exploration via Contrastive Generation
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
On Entity Identification in Language Models
von: Sakata, Masaki, et al.
Veröffentlicht: (2025)
von: Sakata, Masaki, et al.
Veröffentlicht: (2025)
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation
von: Brassard, Ana, et al.
Veröffentlicht: (2024)
von: Brassard, Ana, et al.
Veröffentlicht: (2024)
Is Structure Dependence Shaped for Efficient Communication?: A Case Study on Coordination
von: Kajikawa, Kohei, et al.
Veröffentlicht: (2024)
von: Kajikawa, Kohei, et al.
Veröffentlicht: (2024)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
von: Shen, Zhewen, et al.
Veröffentlicht: (2024)
von: Shen, Zhewen, et al.
Veröffentlicht: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
von: Haga, Akari, et al.
Veröffentlicht: (2024)
von: Haga, Akari, et al.
Veröffentlicht: (2024)
Composition, Attention, or Both?
von: Yoshida, Ryo, et al.
Veröffentlicht: (2022)
von: Yoshida, Ryo, et al.
Veröffentlicht: (2022)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
von: Doi, Tomoki, et al.
Veröffentlicht: (2025)
von: Doi, Tomoki, et al.
Veröffentlicht: (2025)
Comprehensive Evaluation of Large Language Models for Topic Modeling
von: Doi, Tomoki, et al.
Veröffentlicht: (2024)
von: Doi, Tomoki, et al.
Veröffentlicht: (2024)
On the Acquisition of Shared Grammatical Representations in Bilingual Language Models
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
von: El-Shangiti, Ahmed Oumar, et al.
Veröffentlicht: (2024)
von: El-Shangiti, Ahmed Oumar, et al.
Veröffentlicht: (2024)
Repetition Neurons: How Do Language Models Produce Repetitions?
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
Plan Optimization to Bilingual Dictionary Induction for Low-Resource Language Families
von: Nasution, Arbi Haza, et al.
Veröffentlicht: (2020)
von: Nasution, Arbi Haza, et al.
Veröffentlicht: (2020)
How Lexical is Bilingual Lexicon Induction?
von: Kohli, Harsh, et al.
Veröffentlicht: (2024)
von: Kohli, Harsh, et al.
Veröffentlicht: (2024)
Bilingual Literacy Curricula for Bilingual Deaf Students
von: Rachel Friedman Narr, et al.
Veröffentlicht: (2026)
von: Rachel Friedman Narr, et al.
Veröffentlicht: (2026)
Bilingual Colombia: What does It Mean to Be Bilingual within the Framework of the National Plan of Bilingualism?
von: Carmen Helena Guerrero
Veröffentlicht: (2008)
von: Carmen Helena Guerrero
Veröffentlicht: (2008)
A Comparative Analysis of LLM Memorization at Statistical and Internal Levels: Cross-Model Commonalities and Model-Specific Signatures
von: Chen, Bowen, et al.
Veröffentlicht: (2026)
von: Chen, Bowen, et al.
Veröffentlicht: (2026)
Understanding Internal Representations of Recommendation Models with Sparse Autoencoders
von: Wang, Jiayin, et al.
Veröffentlicht: (2024)
von: Wang, Jiayin, et al.
Veröffentlicht: (2024)
Insights on Bilingualism and Bilingual Education: A Sociolinguistic Perspective
von: Iván Ricardo Miranda Montenegro
Veröffentlicht: (2012)
von: Iván Ricardo Miranda Montenegro
Veröffentlicht: (2012)
Bilingual Europe
von: Bloemendal, Jan
Veröffentlicht: (2019)
von: Bloemendal, Jan
Veröffentlicht: (2019)
Position: No Retroactive Cure for Infringement during Training
von: Utsunomiya, Satoru, et al.
Veröffentlicht: (2026)
von: Utsunomiya, Satoru, et al.
Veröffentlicht: (2026)
Towards Transfer Unlearning: Empirical Evidence of Cross-Domain Bias Mitigation
von: Lu, Huimin, et al.
Veröffentlicht: (2024)
von: Lu, Huimin, et al.
Veröffentlicht: (2024)
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
von: Lu, Huimin, et al.
Veröffentlicht: (2025)
von: Lu, Huimin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference
von: Kamoda, Go, et al.
Veröffentlicht: (2025) -
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025) -
Exclusive Unlearning
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2026) -
TopK Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2025) -
The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models
von: Takahashi, Ryosuke, et al.
Veröffentlicht: (2024)