On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
Fuente:
arXiv
Saved in:
| Main Authors: | Kojima, Takeshi, Okimura, Itsuki, Iwasawa, Yusuke, Yanaka, Hitomi, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025)
by: Gambardella, Andrew, et al.
Published: (2025)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
by: Uchiyama, Fumiya, et al.
Published: (2024)
by: Uchiyama, Fumiya, et al.
Published: (2024)
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
by: Harada, Keno, et al.
Published: (2025)
by: Harada, Keno, et al.
Published: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
by: Li, Rongzhi, et al.
Published: (2026)
by: Li, Rongzhi, et al.
Published: (2026)
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
by: Takashiro, Shota, et al.
Published: (2024)
by: Takashiro, Shota, et al.
Published: (2024)
Dynamic Injection of Entity Knowledge into Dense Retrievers
by: Yamada, Ikuya, et al.
Published: (2025)
by: Yamada, Ikuya, et al.
Published: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Neuron-Level Analysis of Cultural Understanding in Large Language Models
by: Yamamoto, Taisei, et al.
Published: (2025)
by: Yamamoto, Taisei, et al.
Published: (2025)
Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
by: Kumon, Ryoma, et al.
Published: (2026)
by: Kumon, Ryoma, et al.
Published: (2026)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
by: Ryu, Koki, et al.
Published: (2026)
by: Ryu, Koki, et al.
Published: (2026)
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
by: Mikami, Yosuke, et al.
Published: (2025)
by: Mikami, Yosuke, et al.
Published: (2025)
Comprehensive Evaluation of Large Language Models for Topic Modeling
by: Doi, Tomoki, et al.
Published: (2024)
by: Doi, Tomoki, et al.
Published: (2024)
Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
by: Huang, Jerry, et al.
Published: (2026)
by: Huang, Jerry, et al.
Published: (2026)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
by: Doi, Tomoki, et al.
Published: (2025)
by: Doi, Tomoki, et al.
Published: (2025)
When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
by: Harada, Keno, et al.
Published: (2025)
by: Harada, Keno, et al.
Published: (2025)
Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models
by: Someya, Taiga, et al.
Published: (2025)
by: Someya, Taiga, et al.
Published: (2025)
Analyzing the Inner Workings of Transformers in Compositional Generalization
by: Kumon, Ryoma, et al.
Published: (2025)
by: Kumon, Ryoma, et al.
Published: (2025)
Enhancing Rating Prediction with Off-the-Shelf LLMs Using In-Context User Reviews
by: Ryu, Koki, et al.
Published: (2025)
by: Ryu, Koki, et al.
Published: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Beyond In-Distribution Success: Scaling Curves of CoT Granularity for Language Model Generalization
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
by: Tang, Tianyi, et al.
Published: (2024)
by: Tang, Tianyi, et al.
Published: (2024)
MKG-Rank: Enhancing Large Language Models with Knowledge Graph for Multilingual Medical Question Answering
by: Li, Feiyang, et al.
Published: (2025)
by: Li, Feiyang, et al.
Published: (2025)
Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
by: Shinozaki, Taiga, et al.
Published: (2025)
by: Shinozaki, Taiga, et al.
Published: (2025)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Evaluating Structural Generalization in Neural Machine Translation
by: Kumon, Ryoma, et al.
Published: (2024)
by: Kumon, Ryoma, et al.
Published: (2024)
LLMs Struggle with NLI for Perfect Aspect: A Cross-Linguistic Study in Chinese and Japanese
by: Lu, Jie, et al.
Published: (2025)
by: Lu, Jie, et al.
Published: (2025)
Implementing a Logical Inference System for Japanese Comparatives
by: Mikami, Yosuke, et al.
Published: (2025)
by: Mikami, Yosuke, et al.
Published: (2025)
Exploring Intra and Inter-language Consistency in Embeddings with ICA
by: Li, Rongzhi, et al.
Published: (2024)
by: Li, Rongzhi, et al.
Published: (2024)
CRANE: Causal Relevance Analysis of Language-Specific Neurons in Multilingual Large Language Models
by: Le, Yifan, et al.
Published: (2026)
by: Le, Yifan, et al.
Published: (2026)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
by: Liu, Yan, et al.
Published: (2024)
by: Liu, Yan, et al.
Published: (2024)
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
by: Ide, Yusuke, et al.
Published: (2026)
by: Ide, Yusuke, et al.
Published: (2026)
Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
by: Yamamoto, Taisei, et al.
Published: (2025)
by: Yamamoto, Taisei, et al.
Published: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
by: Nakata, Wataru, et al.
Published: (2024)
by: Nakata, Wataru, et al.
Published: (2024)
JBBQ: Japanese Bias Benchmark for Analyzing Social Biases in Large Language Models
by: Yanaka, Hitomi, et al.
Published: (2024)
by: Yanaka, Hitomi, et al.
Published: (2024)
From N-grams to Pre-trained Multilingual Models For Language Identification
by: Sindane, Thapelo, et al.
Published: (2024)
by: Sindane, Thapelo, et al.
Published: (2024)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
Similar Items
-
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025) -
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
by: Uchiyama, Fumiya, et al.
Published: (2024) -
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
by: Harada, Keno, et al.
Published: (2025) -
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026) -
NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
by: Li, Rongzhi, et al.
Published: (2026)