Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
Fuente:
arXiv
Guardado en:
| Autores principales: | O'Reilly, Cliff, Jimenez-Ruiz, Ernesto, Weyde, Tillman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KG-CRAFT: Knowledge Graph-based Contrastive Reasoning with LLMs for Enhancing Automated Fact-checking
por: Lourenço, Vítor N., et al.
Publicado: (2026)
por: Lourenço, Vítor N., et al.
Publicado: (2026)
OWL2Vec4OA: Tailoring Knowledge Graph Embeddings for Ontology Alignment
por: Teymurova, Sevinj, et al.
Publicado: (2024)
por: Teymurova, Sevinj, et al.
Publicado: (2024)
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
por: Liu, Xiangyu, et al.
Publicado: (2025)
por: Liu, Xiangyu, et al.
Publicado: (2025)
Beyond Public Access in LLM Pre-Training Data
por: Rosenblat, Sruly, et al.
Publicado: (2025)
por: Rosenblat, Sruly, et al.
Publicado: (2025)
Constrain Alignment with Sparse Autoencoders
por: Yin, Qingyu, et al.
Publicado: (2024)
por: Yin, Qingyu, et al.
Publicado: (2024)
AI-enhanced semantic feature norms for 786 concepts
por: Suresh, Siddharth, et al.
Publicado: (2025)
por: Suresh, Siddharth, et al.
Publicado: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
por: Fang, Yi, et al.
Publicado: (2026)
por: Fang, Yi, et al.
Publicado: (2026)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
por: Liu, Dengcan, et al.
Publicado: (2025)
por: Liu, Dengcan, et al.
Publicado: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
por: Shu, Huizhen, et al.
Publicado: (2025)
por: Shu, Huizhen, et al.
Publicado: (2025)
Sparse Autoencoders for Hypothesis Generation
por: Movva, Rajiv, et al.
Publicado: (2025)
por: Movva, Rajiv, et al.
Publicado: (2025)
Cross-linguistic disagreement as a conflict of semantic alignment norms in multilingual AI~Linguistic Diversity as a Problem for Philosophy, Cognitive Science, and AI~
por: Mizumoto, Masaharu, et al.
Publicado: (2025)
por: Mizumoto, Masaharu, et al.
Publicado: (2025)
Information for Conversation Generation: Proposals Utilising Knowledge Graphs
por: Clay, Alex, et al.
Publicado: (2024)
por: Clay, Alex, et al.
Publicado: (2024)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
por: Shi, Wei, et al.
Publicado: (2025)
por: Shi, Wei, et al.
Publicado: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
por: Xiong, Guangzhi, et al.
Publicado: (2025)
por: Xiong, Guangzhi, et al.
Publicado: (2025)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
por: Zhang, Ruikang, et al.
Publicado: (2026)
por: Zhang, Ruikang, et al.
Publicado: (2026)
Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
por: Ghilardi, Davide, et al.
Publicado: (2024)
por: Ghilardi, Davide, et al.
Publicado: (2024)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
por: Xuan, Richmond Sin Jing, et al.
Publicado: (2025)
por: Xuan, Richmond Sin Jing, et al.
Publicado: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2024)
por: Chanin, David, et al.
Publicado: (2024)
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025)
por: Gallifant, Jack, et al.
Publicado: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
por: Zhao, Haiyan, et al.
Publicado: (2025)
por: Zhao, Haiyan, et al.
Publicado: (2025)
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
por: Flor, Michael, et al.
Publicado: (2025)
por: Flor, Michael, et al.
Publicado: (2025)
Retrieval-augmented generation in multilingual settings
por: Chirkova, Nadezhda, et al.
Publicado: (2024)
por: Chirkova, Nadezhda, et al.
Publicado: (2024)
Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
por: Joshi, Ananya, et al.
Publicado: (2025)
por: Joshi, Ananya, et al.
Publicado: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
por: Wang, Xu, et al.
Publicado: (2026)
por: Wang, Xu, et al.
Publicado: (2026)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
por: Muchane, Mark, et al.
Publicado: (2025)
por: Muchane, Mark, et al.
Publicado: (2025)
LLM-Supported Natural Language to Bash Translation
por: Westenfelder, Finnian, et al.
Publicado: (2025)
por: Westenfelder, Finnian, et al.
Publicado: (2025)
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
por: Husain, Jaavid Aktar, et al.
Publicado: (2024)
por: Husain, Jaavid Aktar, et al.
Publicado: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
por: Han, Lifeng, et al.
Publicado: (2020)
por: Han, Lifeng, et al.
Publicado: (2020)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
por: Hua, Zhenglin, et al.
Publicado: (2025)
por: Hua, Zhenglin, et al.
Publicado: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
por: Lan, Michael, et al.
Publicado: (2024)
por: Lan, Michael, et al.
Publicado: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
por: Farnik, Lucy, et al.
Publicado: (2025)
por: Farnik, Lucy, et al.
Publicado: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025)
por: Li, Aaron J., et al.
Publicado: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
A multimodal multiplex of the mental lexicon for multilingual individuals
por: Huynh, Maria, et al.
Publicado: (2025)
por: Huynh, Maria, et al.
Publicado: (2025)
Scalable multilingual PII annotation for responsible AI in LLMs
por: Meena, Bharti, et al.
Publicado: (2025)
por: Meena, Bharti, et al.
Publicado: (2025)
The Attribution Crisis in LLM Search Results
por: Strauss, Ilan, et al.
Publicado: (2025)
por: Strauss, Ilan, et al.
Publicado: (2025)
Exploring Content and Social Connections of Fake News with Explainable Text and Graph Learning
por: Lourenço, Vítor N., et al.
Publicado: (2025)
por: Lourenço, Vítor N., et al.
Publicado: (2025)
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication
por: Chung, Philip, et al.
Publicado: (2024)
por: Chung, Philip, et al.
Publicado: (2024)
Ejemplares similares
-
KG-CRAFT: Knowledge Graph-based Contrastive Reasoning with LLMs for Enhancing Automated Fact-checking
por: Lourenço, Vítor N., et al.
Publicado: (2026) -
OWL2Vec4OA: Tailoring Knowledge Graph Embeddings for Ontology Alignment
por: Teymurova, Sevinj, et al.
Publicado: (2024) -
ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
por: Liu, Xiangyu, et al.
Publicado: (2025) -
Beyond Public Access in LLM Pre-Training Data
por: Rosenblat, Sruly, et al.
Publicado: (2025) -
Constrain Alignment with Sparse Autoencoders
por: Yin, Qingyu, et al.
Publicado: (2024)