PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
Fuente:
arXiv
Guardado en:
| Autores principales: | Koromilas, Panagiotis, Demou, Andreas D., Oldfield, James, Panagakis, Yannis, Nicolaou, Mihalis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery
por: Demou, Andreas D., et al.
Publicado: (2026)
por: Demou, Andreas D., et al.
Publicado: (2026)
Neural Collapse by Design: Learning Class Prototypes on the Hypersphere
por: Koromilas, Panagiotis, et al.
Publicado: (2026)
por: Koromilas, Panagiotis, et al.
Publicado: (2026)
Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based Losses
por: Koromilas, Panagiotis, et al.
Publicado: (2024)
por: Koromilas, Panagiotis, et al.
Publicado: (2024)
A Principled Framework for Multi-View Contrastive Learning
por: Koromilas, Panagiotis, et al.
Publicado: (2025)
por: Koromilas, Panagiotis, et al.
Publicado: (2025)
Environment-Aware Satellite Image Generation with Diffusion Models
por: Kostagiolas, Nikos, et al.
Publicado: (2025)
por: Kostagiolas, Nikos, et al.
Publicado: (2025)
Enabling Local Editing in Diffusion Models by Joint and Individual Component Analysis
por: Kouzelis, Theodoros, et al.
Publicado: (2024)
por: Kouzelis, Theodoros, et al.
Publicado: (2024)
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
por: Oldfield, James, et al.
Publicado: (2024)
por: Oldfield, James, et al.
Publicado: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
por: Yang, Minglai, et al.
Publicado: (2025)
por: Yang, Minglai, et al.
Publicado: (2025)
Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
por: Oldfield, James, et al.
Publicado: (2025)
por: Oldfield, James, et al.
Publicado: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection
por: Zhang, Huopu, et al.
Publicado: (2025)
por: Zhang, Huopu, et al.
Publicado: (2025)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
por: Korznikov, Anton, et al.
Publicado: (2025)
por: Korznikov, Anton, et al.
Publicado: (2025)
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
por: Cao, Tue M., et al.
Publicado: (2026)
por: Cao, Tue M., et al.
Publicado: (2026)
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
por: Zhang, Mingxu, et al.
Publicado: (2026)
por: Zhang, Mingxu, et al.
Publicado: (2026)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
por: Dimakopoulos, Dimitris, et al.
Publicado: (2026)
por: Dimakopoulos, Dimitris, et al.
Publicado: (2026)
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025)
por: Gallifant, Jack, et al.
Publicado: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
por: Sainsbury, Chris, et al.
Publicado: (2026)
por: Sainsbury, Chris, et al.
Publicado: (2026)
PrivacyScalpel: Enhancing LLM Privacy via Interpretable Feature Intervention with Sparse Autoencoders
por: Frikha, Ahmed, et al.
Publicado: (2025)
por: Frikha, Ahmed, et al.
Publicado: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
por: Lan, Michael, et al.
Publicado: (2024)
por: Lan, Michael, et al.
Publicado: (2024)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
por: Yan, Xinyuan, et al.
Publicado: (2025)
por: Yan, Xinyuan, et al.
Publicado: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
por: Wang, Xu, et al.
Publicado: (2025)
por: Wang, Xu, et al.
Publicado: (2025)
Dense SAE Latents Are Features, Not Bugs
por: Sun, Xiaoqing, et al.
Publicado: (2025)
por: Sun, Xiaoqing, et al.
Publicado: (2025)
Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search
por: Plitsis, Manos, et al.
Publicado: (2025)
por: Plitsis, Manos, et al.
Publicado: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
por: Muhamed, Aashiq, et al.
Publicado: (2024)
por: Muhamed, Aashiq, et al.
Publicado: (2024)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Training Superior Sparse Autoencoders for Instruct Models
por: Li, Jiaming, et al.
Publicado: (2025)
por: Li, Jiaming, et al.
Publicado: (2025)
Feature Rivalry in Sparse Autoencoder Representations: A Mechanistic Study of Uncertainty-Driven Feature Competition in LLMs
por: Harshavardhan
Publicado: (2026)
por: Harshavardhan
Publicado: (2026)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
Generalization analysis of an unfolding network for analysis-based Compressed Sensing
por: Kouni, Vicky, et al.
Publicado: (2023)
por: Kouni, Vicky, et al.
Publicado: (2023)
Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
por: Kalogeropoulos, Ioannis, et al.
Publicado: (2025)
por: Kalogeropoulos, Ioannis, et al.
Publicado: (2025)
Scale Equivariant Graph Metanetworks
por: Kalogeropoulos, Ioannis, et al.
Publicado: (2024)
por: Kalogeropoulos, Ioannis, et al.
Publicado: (2024)
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
por: Brzozowski, Michał, et al.
Publicado: (2026)
por: Brzozowski, Michał, et al.
Publicado: (2026)
Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits
por: Yap, Jia Qing
Publicado: (2026)
por: Yap, Jia Qing
Publicado: (2026)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
por: Mishra, Anurag
Publicado: (2026)
por: Mishra, Anurag
Publicado: (2026)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
por: Zheng, Carolina, et al.
Publicado: (2025)
por: Zheng, Carolina, et al.
Publicado: (2025)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
por: He, Zhengfu, et al.
Publicado: (2024)
por: He, Zhengfu, et al.
Publicado: (2024)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
por: O'Neill, Charles, et al.
Publicado: (2024)
por: O'Neill, Charles, et al.
Publicado: (2024)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
Ejemplares similares
-
fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery
por: Demou, Andreas D., et al.
Publicado: (2026) -
Neural Collapse by Design: Learning Class Prototypes on the Hypersphere
por: Koromilas, Panagiotis, et al.
Publicado: (2026) -
Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based Losses
por: Koromilas, Panagiotis, et al.
Publicado: (2024) -
A Principled Framework for Multi-View Contrastive Learning
por: Koromilas, Panagiotis, et al.
Publicado: (2025) -
Environment-Aware Satellite Image Generation with Diffusion Models
por: Kostagiolas, Nikos, et al.
Publicado: (2025)