Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Xuan, Richmond Sin Jing, Huseynov, Jalil, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Constrain Alignment with Sparse Autoencoders
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
by: Farid, Sualeha, et al.
Published: (2025)
by: Farid, Sualeha, et al.
Published: (2025)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
by: Zhang, Ruikang, et al.
Published: (2026)
by: Zhang, Ruikang, et al.
Published: (2026)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
by: Liu, Dengcan, et al.
Published: (2025)
by: Liu, Dengcan, et al.
Published: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
Sparse Autoencoders for Hypothesis Generation
by: Movva, Rajiv, et al.
Published: (2025)
by: Movva, Rajiv, et al.
Published: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
by: Jørgensen, Mikkel Godsk, et al.
Published: (2026)
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
by: Nguyen, Xuan-Phi, et al.
Published: (2023)
by: Nguyen, Xuan-Phi, et al.
Published: (2023)
Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
by: Jaipersaud, Brandon, et al.
Published: (2024)
by: Jaipersaud, Brandon, et al.
Published: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
by: Wu, Xinwei, et al.
Published: (2025)
by: Wu, Xinwei, et al.
Published: (2025)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective
by: Yang, Wen, et al.
Published: (2025)
by: Yang, Wen, et al.
Published: (2025)
Popular LLMs Amplify Race and Gender Disparities in Human Mobility
by: Wu, Xinhua, et al.
Published: (2024)
by: Wu, Xinhua, et al.
Published: (2024)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
by: O'Reilly, Cliff, et al.
Published: (2025)
by: O'Reilly, Cliff, et al.
Published: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2024)
by: Chanin, David, et al.
Published: (2024)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
Mapping Clinical Doubt: Locating Linguistic Uncertainty in LLMs
by: Sridhar, Srivarshinee, et al.
Published: (2025)
by: Sridhar, Srivarshinee, et al.
Published: (2025)
Large Linguistic Models: Investigating LLMs' metalinguistic abilities
by: Beguš, Gašper, et al.
Published: (2023)
by: Beguš, Gašper, et al.
Published: (2023)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
by: Shu, Huizhen, et al.
Published: (2025)
by: Shu, Huizhen, et al.
Published: (2025)
Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
by: Joshi, Ananya, et al.
Published: (2025)
by: Joshi, Ananya, et al.
Published: (2025)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
by: Muchane, Mark, et al.
Published: (2025)
by: Muchane, Mark, et al.
Published: (2025)
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations
by: Bhagat, Kirti, et al.
Published: (2024)
by: Bhagat, Kirti, et al.
Published: (2024)
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
by: McCoubrey, Michael, et al.
Published: (2026)
by: McCoubrey, Michael, et al.
Published: (2026)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
by: Yu, Zhuohao, et al.
Published: (2025)
by: Yu, Zhuohao, et al.
Published: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
by: Jing, Yi, et al.
Published: (2026)
by: Jing, Yi, et al.
Published: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2025)
by: Chanin, David, et al.
Published: (2025)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
by: Hwang, Kyomin, et al.
Published: (2025)
by: Hwang, Kyomin, et al.
Published: (2025)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
by: He, Jinwen, et al.
Published: (2023)
by: He, Jinwen, et al.
Published: (2023)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
by: Zhang, Zhuoxuan, et al.
Published: (2025)
by: Zhang, Zhuoxuan, et al.
Published: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
by: Chalnev, Sviatoslav, et al.
Published: (2024)
by: Chalnev, Sviatoslav, et al.
Published: (2024)
Similar Items
-
Constrain Alignment with Sparse Autoencoders
by: Yin, Qingyu, et al.
Published: (2024) -
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
by: Farid, Sualeha, et al.
Published: (2025) -
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
by: Zhang, Ruikang, et al.
Published: (2026) -
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
by: Liu, Dengcan, et al.
Published: (2025) -
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
by: Xiong, Guangzhi, et al.
Published: (2025)