Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Hanqi, Xiang, Yanzheng, Chen, Guangyi, Wang, Yifei, Gui, Lin, He, Yulan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
by: Xiang, Yanzheng, et al.
Published: (2025)
by: Xiang, Yanzheng, et al.
Published: (2025)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models
by: Xiang, Yanzheng, et al.
Published: (2024)
by: Xiang, Yanzheng, et al.
Published: (2024)
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
by: McCoubrey, Michael, et al.
Published: (2026)
by: McCoubrey, Michael, et al.
Published: (2026)
A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models
by: Mamalakis, Michail, et al.
Published: (2026)
by: Mamalakis, Michail, et al.
Published: (2026)
Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning
by: Yan, Hanqi, et al.
Published: (2024)
by: Yan, Hanqi, et al.
Published: (2024)
Measuring and Guiding Monosemanticity
by: Härle, Ruben, et al.
Published: (2025)
by: Härle, Ruben, et al.
Published: (2025)
Learning from Emergence: A Study on Proactively Inhibiting the Monosemantic Neurons of Artificial Neural Networks
by: Wang, Jiachuan, et al.
Published: (2023)
by: Wang, Jiachuan, et al.
Published: (2023)
Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding
by: Xiang, Yanzheng, et al.
Published: (2026)
by: Xiang, Yanzheng, et al.
Published: (2026)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
Monet: Mixture of Monosemantic Experts for Transformers
by: Park, Jungwoo, et al.
Published: (2024)
by: Park, Jungwoo, et al.
Published: (2024)
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
by: Templeton, Adly, et al.
Published: (2026)
by: Templeton, Adly, et al.
Published: (2026)
Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
by: Hu, Zhanghao, et al.
Published: (2026)
by: Hu, Zhanghao, et al.
Published: (2026)
The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis
by: Zhou, Yuxiang, et al.
Published: (2023)
by: Zhou, Yuxiang, et al.
Published: (2023)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025)
by: Hu, Zhanghao, et al.
Published: (2025)
PECAN: LLM-Guided Dynamic Progress Control with Attention-Guided Hierarchical Weighted Graph for Long-Document QA
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Extracting Interaction-Aware Monosemantic Concepts in Recommender Systems
by: Arviv, Dor, et al.
Published: (2025)
by: Arviv, Dor, et al.
Published: (2025)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
by: Qi, Siya, et al.
Published: (2024)
by: Qi, Siya, et al.
Published: (2024)
Revisiting Long-context Modeling from Context Denoising Perspective
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
by: Zhao, Runcong, et al.
Published: (2024)
by: Zhao, Runcong, et al.
Published: (2024)
Constrain Alignment with Sparse Autoencoders
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
by: Li, Loka, et al.
Published: (2024)
by: Li, Loka, et al.
Published: (2024)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
CWTM: Leveraging Contextualized Word Embeddings from BERT for Neural Topic Modeling
by: Fang, Zheng, et al.
Published: (2023)
by: Fang, Zheng, et al.
Published: (2023)
Weak Reward Model Transforms Generative Models into Robust Causal Event Extraction Systems
by: da Silva, Italo Luis, et al.
Published: (2024)
by: da Silva, Italo Luis, et al.
Published: (2024)
Explainable Recommender with Geometric Information Bottleneck
by: Yan, Hanqi, et al.
Published: (2023)
by: Yan, Hanqi, et al.
Published: (2023)
Rehearse With User: Personalized Opinion Summarization via Role-Playing based on Large Language Models
by: Zhang, Yanyue, et al.
Published: (2025)
by: Zhang, Yanyue, et al.
Published: (2025)
Deep Dreams Are Made of This: Visualizing Monosemantic Features in Diffusion Models
by: Szokalski, Adam, et al.
Published: (2026)
by: Szokalski, Adam, et al.
Published: (2026)
Sparse Activation Editing for Reliable Instruction Following in Narratives
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
Revisiting Self-Consistency from Dynamic Distributional Alignment Perspective on Answer Aggregation
by: Li, Yiwei, et al.
Published: (2025)
by: Li, Yiwei, et al.
Published: (2025)
GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
by: da Silva, Italo Luis, et al.
Published: (2025)
by: da Silva, Italo Luis, et al.
Published: (2025)
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
by: Li, Tianlong, et al.
Published: (2024)
by: Li, Tianlong, et al.
Published: (2024)
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
by: Zhu, Qinglin, et al.
Published: (2025)
by: Zhu, Qinglin, et al.
Published: (2025)
Cascading Large Language Models for Salient Event Graph Generation
by: Tan, Xingwei, et al.
Published: (2024)
by: Tan, Xingwei, et al.
Published: (2024)
ExDDI: Explaining Drug-Drug Interaction Predictions with Natural Language
by: Sun, Zhaoyue, et al.
Published: (2024)
by: Sun, Zhaoyue, et al.
Published: (2024)
A Statistical and Multi-Perspective Revisiting of the Membership Inference Attack in Large Language Models
by: Chen, Bowen, et al.
Published: (2024)
by: Chen, Bowen, et al.
Published: (2024)
Hán Dān Xué Bù (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
by: Zheng, Yingming, et al.
Published: (2025)
by: Zheng, Yingming, et al.
Published: (2025)
Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement
by: Zhu, Zipeng, et al.
Published: (2026)
by: Zhu, Zipeng, et al.
Published: (2026)
Similar Items
-
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
by: Xiang, Yanzheng, et al.
Published: (2025) -
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
by: Zhang, Qi, et al.
Published: (2024) -
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
by: Ye, Charles, et al.
Published: (2026) -
Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models
by: Xiang, Yanzheng, et al.
Published: (2024) -
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
by: McCoubrey, Michael, et al.
Published: (2026)