Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Daniel, Shankarampeta, Abhilash, Hu, Lanxiang, Rosing, Tajana, Zhang, Hao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
por: Hu, Lanxiang, et al.
Publicado: (2024)
por: Hu, Lanxiang, et al.
Publicado: (2024)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
por: Hu, Lanxiang, et al.
Publicado: (2025)
por: Hu, Lanxiang, et al.
Publicado: (2025)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
por: Zhao, Yujie, et al.
Publicado: (2026)
por: Zhao, Yujie, et al.
Publicado: (2026)
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
por: Ponzina, Flavio, et al.
Publicado: (2024)
por: Ponzina, Flavio, et al.
Publicado: (2024)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
por: Parsan, Nithin, et al.
Publicado: (2025)
por: Parsan, Nithin, et al.
Publicado: (2025)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
por: Kong, Jason, et al.
Publicado: (2026)
por: Kong, Jason, et al.
Publicado: (2026)
Evidence-Guided Schema Normalization for Temporal Tabular Reasoning
por: Thanga, Ashish, et al.
Publicado: (2025)
por: Thanga, Ashish, et al.
Publicado: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
por: Jiang, Nick, et al.
Publicado: (2025)
por: Jiang, Nick, et al.
Publicado: (2025)
Divide and Learn: Multi-Objective Combinatorial Optimization at Scale
por: Singh, Esha, et al.
Publicado: (2026)
por: Singh, Esha, et al.
Publicado: (2026)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
por: Tolooshams, Bahareh, et al.
Publicado: (2025)
por: Tolooshams, Bahareh, et al.
Publicado: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
por: Cywiński, Bartosz, et al.
Publicado: (2025)
por: Cywiński, Bartosz, et al.
Publicado: (2025)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
por: Li, Yuhan, et al.
Publicado: (2026)
por: Li, Yuhan, et al.
Publicado: (2026)
Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
por: Lin, Fudong, et al.
Publicado: (2025)
por: Lin, Fudong, et al.
Publicado: (2025)
Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
por: Cao, Tue M., et al.
Publicado: (2026)
por: Cao, Tue M., et al.
Publicado: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
por: Zaigrajew, Vladimir, et al.
Publicado: (2025)
por: Zaigrajew, Vladimir, et al.
Publicado: (2025)
Sparse Autoencoders, Again?
por: Lu, Yin, et al.
Publicado: (2025)
por: Lu, Yin, et al.
Publicado: (2025)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
por: Shu, Dong, et al.
Publicado: (2025)
por: Shu, Dong, et al.
Publicado: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
por: Balagansky, Nikita, et al.
Publicado: (2025)
por: Balagansky, Nikita, et al.
Publicado: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025)
por: Bhalla, Usha, et al.
Publicado: (2025)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
por: Klenitskiy, Anton, et al.
Publicado: (2025)
por: Klenitskiy, Anton, et al.
Publicado: (2025)
Improving Sparse Autoencoder with Dynamic Attention
por: Wang, Dongsheng, et al.
Publicado: (2026)
por: Wang, Dongsheng, et al.
Publicado: (2026)
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
por: Qian, Yu-Yang, et al.
Publicado: (2026)
por: Qian, Yu-Yang, et al.
Publicado: (2026)
Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
por: Jha, Gopi Krishna, et al.
Publicado: (2023)
por: Jha, Gopi Krishna, et al.
Publicado: (2023)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
por: He, Zirui, et al.
Publicado: (2025)
por: He, Zirui, et al.
Publicado: (2025)
COT Flow: Learning Optimal-Transport Image Sampling and Editing by Contrastive Pairs
por: Zu, Xinrui, et al.
Publicado: (2024)
por: Zu, Xinrui, et al.
Publicado: (2024)
Are Sparse Autoencoder Benchmarks Reliable?
por: Chanin, David
Publicado: (2026)
por: Chanin, David
Publicado: (2026)
Do Sparse Autoencoders Capture Concept Manifolds?
por: Bhalla, Usha, et al.
Publicado: (2026)
por: Bhalla, Usha, et al.
Publicado: (2026)
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
por: Weng, Jiaqi, et al.
Publicado: (2025)
por: Weng, Jiaqi, et al.
Publicado: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
por: Jing, Yi, et al.
Publicado: (2026)
por: Jing, Yi, et al.
Publicado: (2026)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
por: Yeung, Calvin, et al.
Publicado: (2026)
por: Yeung, Calvin, et al.
Publicado: (2026)
Towards Understanding the Robustness of Sparse Autoencoders
por: Saiyed, Ahson, et al.
Publicado: (2026)
por: Saiyed, Ahson, et al.
Publicado: (2026)
BatchTopK Sparse Autoencoders
por: Bussmann, Bart, et al.
Publicado: (2024)
por: Bussmann, Bart, et al.
Publicado: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
por: Kantamneni, Subhash, et al.
Publicado: (2025)
por: Kantamneni, Subhash, et al.
Publicado: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
por: Wang, Xu, et al.
Publicado: (2026)
por: Wang, Xu, et al.
Publicado: (2026)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
por: Zhao, Yilong, et al.
Publicado: (2025)
por: Zhao, Yilong, et al.
Publicado: (2025)
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
por: Ayonrinde, Kola
Publicado: (2024)
por: Ayonrinde, Kola
Publicado: (2024)
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
por: Peter, Hans, et al.
Publicado: (2025)
por: Peter, Hans, et al.
Publicado: (2025)
On the transferability of Sparse Autoencoders for interpreting compressed models
por: Gupte, Suchit, et al.
Publicado: (2025)
por: Gupte, Suchit, et al.
Publicado: (2025)
Data Whitening Improves Sparse Autoencoder Learning
por: Saraswatula, Ashwin, et al.
Publicado: (2025)
por: Saraswatula, Ashwin, et al.
Publicado: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
por: Rajamanoharan, Senthooran, et al.
Publicado: (2024)
por: Rajamanoharan, Senthooran, et al.
Publicado: (2024)
Ejemplares similares
-
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
por: Hu, Lanxiang, et al.
Publicado: (2024) -
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
por: Hu, Lanxiang, et al.
Publicado: (2025) -
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
por: Zhao, Yujie, et al.
Publicado: (2026) -
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
por: Ponzina, Flavio, et al.
Publicado: (2024) -
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
por: Parsan, Nithin, et al.
Publicado: (2025)