Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
Fuente:
arXiv
Guardado en:
| Autores principales: | Minegishi, Gouki, Furuta, Hiroki, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
por: Furuta, Hiroki, et al.
Publicado: (2024)
por: Furuta, Hiroki, et al.
Publicado: (2024)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
por: Minegishi, Gouki, et al.
Publicado: (2026)
por: Minegishi, Gouki, et al.
Publicado: (2026)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
por: Matsutani, Kohsei, et al.
Publicado: (2026)
por: Matsutani, Kohsei, et al.
Publicado: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
por: Minegishi, Gouki, et al.
Publicado: (2023)
por: Minegishi, Gouki, et al.
Publicado: (2023)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
por: Feng, Jingyuan, et al.
Publicado: (2026)
por: Feng, Jingyuan, et al.
Publicado: (2026)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
por: Gambardella, Andrew, et al.
Publicado: (2024)
por: Gambardella, Andrew, et al.
Publicado: (2024)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
por: Gambardella, Andrew, et al.
Publicado: (2025)
por: Gambardella, Andrew, et al.
Publicado: (2025)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
por: Furuta, Hiroki, et al.
Publicado: (2023)
por: Furuta, Hiroki, et al.
Publicado: (2023)
Mechanism of Task-oriented Information Removal in In-context Learning
por: Cho, Hakaze, et al.
Publicado: (2025)
por: Cho, Hakaze, et al.
Publicado: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
por: Cao, Qi, et al.
Publicado: (2026)
por: Cao, Qi, et al.
Publicado: (2026)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
por: Wang, Ru, et al.
Publicado: (2025)
por: Wang, Ru, et al.
Publicado: (2025)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
por: Gur, Izzeddin, et al.
Publicado: (2023)
por: Gur, Izzeddin, et al.
Publicado: (2023)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025)
por: Li, Aaron J., et al.
Publicado: (2025)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
por: Matsutani, Kohsei, et al.
Publicado: (2025)
por: Matsutani, Kohsei, et al.
Publicado: (2025)
Geometric-Averaged Preference Optimization for Soft Preference Labels
por: Furuta, Hiroki, et al.
Publicado: (2024)
por: Furuta, Hiroki, et al.
Publicado: (2024)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
por: Gu, Xiaojie, et al.
Publicado: (2026)
por: Gu, Xiaojie, et al.
Publicado: (2026)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
por: Iwase, Naoto, et al.
Publicado: (2025)
por: Iwase, Naoto, et al.
Publicado: (2025)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
por: Yang, Bo, et al.
Publicado: (2025)
por: Yang, Bo, et al.
Publicado: (2025)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
por: Oshima, Yuta, et al.
Publicado: (2025)
por: Oshima, Yuta, et al.
Publicado: (2025)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
por: Yang, Xianjun, et al.
Publicado: (2025)
por: Yang, Xianjun, et al.
Publicado: (2025)
Sparse Autoencoder Features for Classifications and Transferability
por: Gallifant, Jack, et al.
Publicado: (2025)
por: Gallifant, Jack, et al.
Publicado: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
por: Sainsbury, Chris, et al.
Publicado: (2026)
por: Sainsbury, Chris, et al.
Publicado: (2026)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
por: Muchane, Mark, et al.
Publicado: (2025)
por: Muchane, Mark, et al.
Publicado: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
por: Farnik, Lucy, et al.
Publicado: (2025)
por: Farnik, Lucy, et al.
Publicado: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
por: Chalnev, Sviatoslav, et al.
Publicado: (2024)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
por: Kim, Bum Jun, et al.
Publicado: (2026)
por: Kim, Bum Jun, et al.
Publicado: (2026)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025)
por: Bhalla, Usha, et al.
Publicado: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
por: Yu, Zhuohao, et al.
Publicado: (2025)
por: Yu, Zhuohao, et al.
Publicado: (2025)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
por: Uchiyama, Fumiya, et al.
Publicado: (2024)
por: Uchiyama, Fumiya, et al.
Publicado: (2024)
Static Word Embeddings for Sentence Semantic Representation
por: Wada, Takashi, et al.
Publicado: (2025)
por: Wada, Takashi, et al.
Publicado: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
por: Shu, Dong, et al.
Publicado: (2025)
por: Shu, Dong, et al.
Publicado: (2025)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
por: Kubo, Kenji, et al.
Publicado: (2026)
por: Kubo, Kenji, et al.
Publicado: (2026)
Ejemplares similares
-
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
por: Furuta, Hiroki, et al.
Publicado: (2024) -
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
por: Minegishi, Gouki, et al.
Publicado: (2025) -
Understanding Emergent Misalignment via Feature Superposition Geometry
por: Minegishi, Gouki, et al.
Publicado: (2026) -
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
por: Minegishi, Gouki, et al.
Publicado: (2025) -
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
por: Matsutani, Kohsei, et al.
Publicado: (2026)