Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
Fuente:
arXiv
Guardado en:
| Autores principales: | Ma, Qingsen, Wang, Dianyun, Lyu, Jiaming, Wang, Yaoye, Ning, Lechen, Zhu, Sujie, Xu, Zhenbo, Xiang, Liuyu, Li, Huining, Wu, Huijia, He, Zhaofeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
por: Ma, Qingsen, et al.
Publicado: (2026)
por: Ma, Qingsen, et al.
Publicado: (2026)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
por: Wang, Dianyun, et al.
Publicado: (2025)
por: Wang, Dianyun, et al.
Publicado: (2025)
Beyond Darkness: Thermal-Supervised 3D Gaussian Splatting for Low-Light Novel View Synthesis
por: Ma, Qingsen, et al.
Publicado: (2025)
por: Ma, Qingsen, et al.
Publicado: (2025)
CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
por: Ma, Qingsen, et al.
Publicado: (2025)
por: Ma, Qingsen, et al.
Publicado: (2025)
Rethinking Class-Incremental Learning from a Dynamic Imbalanced Learning Perspective
por: Wang, Leyuan, et al.
Publicado: (2024)
por: Wang, Leyuan, et al.
Publicado: (2024)
Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
por: Liu, Shuodi, et al.
Publicado: (2025)
por: Liu, Shuodi, et al.
Publicado: (2025)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
por: Ye, Mengyu, et al.
Publicado: (2025)
por: Ye, Mengyu, et al.
Publicado: (2025)
Dynamic Generation of Personalities with Large Language Models
por: Liu, Jianzhi, et al.
Publicado: (2024)
por: Liu, Jianzhi, et al.
Publicado: (2024)
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation
por: Zhang, Hongxuan, et al.
Publicado: (2024)
por: Zhang, Hongxuan, et al.
Publicado: (2024)
MedSAE: Dissecting MedCLIP Representations with Sparse Autoencoders
por: Renzulli, Riccardo, et al.
Publicado: (2025)
por: Renzulli, Riccardo, et al.
Publicado: (2025)
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
por: Kong, Zicheng, et al.
Publicado: (2026)
por: Kong, Zicheng, et al.
Publicado: (2026)
CLIP model is an Efficient Online Lifelong Learner
por: Wang, Leyuan, et al.
Publicado: (2024)
por: Wang, Leyuan, et al.
Publicado: (2024)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
por: Muchane, Mark, et al.
Publicado: (2025)
por: Muchane, Mark, et al.
Publicado: (2025)
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
por: Shi, Wei, et al.
Publicado: (2025)
por: Shi, Wei, et al.
Publicado: (2025)
Interpretable Reward Model via Sparse Autoencoder
por: Zhang, Shuyi, et al.
Publicado: (2025)
por: Zhang, Shuyi, et al.
Publicado: (2025)
Two-Timescale Optimization Framework for Sparse-Feedback Linear-Quadratic Optimal Control
por: Feng, Lechen, et al.
Publicado: (2024)
por: Feng, Lechen, et al.
Publicado: (2024)
Training Superior Sparse Autoencoders for Instruct Models
por: Li, Jiaming, et al.
Publicado: (2025)
por: Li, Jiaming, et al.
Publicado: (2025)
Route Sparse Autoencoder to Interpret Large Language Models
por: Shi, Wei, et al.
Publicado: (2025)
por: Shi, Wei, et al.
Publicado: (2025)
Constrain Alignment with Sparse Autoencoders
por: Yin, Qingyu, et al.
Publicado: (2024)
por: Yin, Qingyu, et al.
Publicado: (2024)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
por: Mishra, Anurag
Publicado: (2026)
por: Mishra, Anurag
Publicado: (2026)
Improving Sparse Autoencoder with Dynamic Attention
por: Wang, Dongsheng, et al.
Publicado: (2026)
por: Wang, Dongsheng, et al.
Publicado: (2026)
Nonconvex Optimization Framework for Group-Sparse Feedback Linear-Quadratic Optimal Control: Penalty Approach
por: Feng, Lechen, et al.
Publicado: (2025)
por: Feng, Lechen, et al.
Publicado: (2025)
Ensembling Sparse Autoencoders
por: Gadgil, Soham, et al.
Publicado: (2025)
por: Gadgil, Soham, et al.
Publicado: (2025)
Sparse Autoencoders, Again?
por: Lu, Yin, et al.
Publicado: (2025)
por: Lu, Yin, et al.
Publicado: (2025)
Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
por: Ro, Yusung, et al.
Publicado: (2026)
por: Ro, Yusung, et al.
Publicado: (2026)
Nonconvex Optimization Framework for Group-Sparse Feedback Linear-Quadratic Optimal Control: Non-Penalty Approach
por: Feng, Lechen, et al.
Publicado: (2025)
por: Feng, Lechen, et al.
Publicado: (2025)
Simulation-Free PSRO: Removing Game Simulation from Policy Space Response Oracles
por: Liu, Yingzhuo, et al.
Publicado: (2025)
por: Liu, Yingzhuo, et al.
Publicado: (2025)
Generative Iris Prior Embedded Transformer for Iris Restoration
por: Huang, Yubo, et al.
Publicado: (2024)
por: Huang, Yubo, et al.
Publicado: (2024)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
por: Yang, Xuan, et al.
Publicado: (2026)
por: Yang, Xuan, et al.
Publicado: (2026)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
por: Zhang, Ruikang, et al.
Publicado: (2026)
por: Zhang, Ruikang, et al.
Publicado: (2026)
CSRv2: Unlocking Ultra-Sparse Embeddings
por: Guo, Lixuan, et al.
Publicado: (2026)
por: Guo, Lixuan, et al.
Publicado: (2026)
Multilingual Safety Alignment Via Sparse Weight Editing
por: Liang, Jiaming, et al.
Publicado: (2026)
por: Liang, Jiaming, et al.
Publicado: (2026)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
por: Cao, Tue M., et al.
Publicado: (2026)
por: Cao, Tue M., et al.
Publicado: (2026)
Sparse Autoencoders for Hypothesis Generation
por: Movva, Rajiv, et al.
Publicado: (2025)
por: Movva, Rajiv, et al.
Publicado: (2025)
Sparse Autoencoders are Topic Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Toward Identifiable Sparse Autoencoders
por: Nelson, Walter, et al.
Publicado: (2026)
por: Nelson, Walter, et al.
Publicado: (2026)
Analysis of Variational Sparse Autoencoders
por: Baker, Zachary, et al.
Publicado: (2025)
por: Baker, Zachary, et al.
Publicado: (2025)
Are Sparse Autoencoder Benchmarks Reliable?
por: Chanin, David
Publicado: (2026)
por: Chanin, David
Publicado: (2026)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
por: Kantamneni, Subhash, et al.
Publicado: (2025)
por: Kantamneni, Subhash, et al.
Publicado: (2025)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
por: Liu, Dengcan, et al.
Publicado: (2025)
por: Liu, Dengcan, et al.
Publicado: (2025)
Ejemplares similares
-
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
por: Ma, Qingsen, et al.
Publicado: (2026) -
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
por: Wang, Dianyun, et al.
Publicado: (2025) -
Beyond Darkness: Thermal-Supervised 3D Gaussian Splatting for Low-Light Novel View Synthesis
por: Ma, Qingsen, et al.
Publicado: (2025) -
CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
por: Ma, Qingsen, et al.
Publicado: (2025) -
Rethinking Class-Incremental Learning from a Dynamic Imbalanced Learning Perspective
por: Wang, Leyuan, et al.
Publicado: (2024)