Improving Multilingual Language Models by Aligning Representations through Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Mahmoud, Omar, Semage, Buddhika Laknath, Karimpanal, Thommen George, Rana, Santu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
by: Mahmoud, Omar, et al.
Published: (2025)
by: Mahmoud, Omar, et al.
Published: (2025)
Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment
by: Hussonnois, Maxence, et al.
Published: (2025)
by: Hussonnois, Maxence, et al.
Published: (2025)
ECoDe: A Sample-Efficient Method for Co-Design of Robotic Agents
by: Nagiredla, Kishan R., et al.
Published: (2023)
by: Nagiredla, Kishan R., et al.
Published: (2023)
Leveraging Human Feedback for Semantically-Relevant Skill Discovery
by: Hussonnois, Maxence, et al.
Published: (2026)
by: Hussonnois, Maxence, et al.
Published: (2026)
ASPECT:Analogical Semantic Policy Execution via Language Conditioned Transfer
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2026)
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2026)
MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2025)
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2025)
Dynamic Policy Fusion for User Alignment Without Re-Interaction
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2024)
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2024)
Improved Representation Steering for Language Models
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment
by: Bu, Mengyu, et al.
Published: (2025)
by: Bu, Mengyu, et al.
Published: (2025)
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
by: Zhu, Jian-Qiao, et al.
Published: (2025)
by: Zhu, Jian-Qiao, et al.
Published: (2025)
Language Steering for Multilingual In-Context Learning
by: Kirtane, Neeraja, et al.
Published: (2026)
by: Kirtane, Neeraja, et al.
Published: (2026)
Multilingual Political Views of Large Language Models: Identification and Steering
by: Gurgurov, Daniil, et al.
Published: (2025)
by: Gurgurov, Daniil, et al.
Published: (2025)
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
by: Kassem, Aly M., et al.
Published: (2024)
by: Kassem, Aly M., et al.
Published: (2024)
Cross-Lingual Activation Steering for Multilingual Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
The Cylindrical Representation Hypothesis for Language Model Steering
by: Gao, Lang, et al.
Published: (2026)
by: Gao, Lang, et al.
Published: (2026)
Aligning Large Language Models with Human Preferences through Representation Engineering
by: Liu, Wenhao, et al.
Published: (2023)
by: Liu, Wenhao, et al.
Published: (2023)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning Argumentations
by: Ranaldi, Leonardo, et al.
Published: (2025)
by: Ranaldi, Leonardo, et al.
Published: (2025)
Linguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource Languages
by: Fernando, Aloka, et al.
Published: (2025)
by: Fernando, Aloka, et al.
Published: (2025)
CM-Align: Consistency-based Multilingual Alignment for Large Language Models
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
Self-Improving Model Steering
by: Zhu, Rongyi, et al.
Published: (2025)
by: Zhu, Rongyi, et al.
Published: (2025)
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
by: Ghussin, Yusser Al, et al.
Published: (2026)
by: Ghussin, Yusser Al, et al.
Published: (2026)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
by: Knipper, R. Alexander, et al.
Published: (2025)
by: Knipper, R. Alexander, et al.
Published: (2025)
How Language Directions Align with Token Geometry in Multilingual LLMs
by: Kim, JaeSeong, et al.
Published: (2025)
by: Kim, JaeSeong, et al.
Published: (2025)
On Effects of Steering Latent Representation for Large Language Model Unlearning
by: Huu-Tien, Dang, et al.
Published: (2024)
by: Huu-Tien, Dang, et al.
Published: (2024)
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
Aligning Language Models with Demonstrated Feedback
by: Shaikh, Omar, et al.
Published: (2024)
by: Shaikh, Omar, et al.
Published: (2024)
Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
by: Puranegedara, Imalsha, et al.
Published: (2025)
by: Puranegedara, Imalsha, et al.
Published: (2025)
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
Steering off Course: Reliability Challenges in Steering Language Models
by: Da Silva, Patrick Queiroz, et al.
Published: (2025)
by: Da Silva, Patrick Queiroz, et al.
Published: (2025)
Vision-Language Models Align with Human Neural Representations in Concept Processing
by: Bavaresco, Anna, et al.
Published: (2024)
by: Bavaresco, Anna, et al.
Published: (2024)
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning
by: Koleilat, Taha, et al.
Published: (2026)
by: Koleilat, Taha, et al.
Published: (2026)
Improving Multilingual Math Reasoning for African Languages
by: Ogundepo, Odunayo, et al.
Published: (2025)
by: Ogundepo, Odunayo, et al.
Published: (2025)
AlignFreeze: Navigating the Impact of Realignment on the Layers of Multilingual Models Across Diverse Languages
by: Bakos, Steve, et al.
Published: (2025)
by: Bakos, Steve, et al.
Published: (2025)
Redundancy Aware Multi-Reference Based Gainwise Evaluation of Extractive Summarization
by: Akter, Mousumi, et al.
Published: (2023)
by: Akter, Mousumi, et al.
Published: (2023)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
by: Paul, Indraneil, et al.
Published: (2024)
by: Paul, Indraneil, et al.
Published: (2024)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Large Language Models for IT Automation Tasks: Are We There Yet?
by: Hassan, Md Mahadi, et al.
Published: (2025)
by: Hassan, Md Mahadi, et al.
Published: (2025)
Multilingual Large Language Models and Curse of Multilinguality
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
Similar Items
-
The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
by: Mahmoud, Omar, et al.
Published: (2025) -
Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment
by: Hussonnois, Maxence, et al.
Published: (2025) -
ECoDe: A Sample-Efficient Method for Co-Design of Robotic Agents
by: Nagiredla, Kishan R., et al.
Published: (2023) -
Leveraging Human Feedback for Semantically-Relevant Skill Discovery
by: Hussonnois, Maxence, et al.
Published: (2026) -
ASPECT:Analogical Semantic Policy Execution via Language Conditioned Transfer
by: Palattuparambil, Ajsal Shereef, et al.
Published: (2026)