Finding Belief Geometries with Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Author: | Levinson, Matthew |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents
by: Levinson, Matthew
Published: (2026)
by: Levinson, Matthew
Published: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
Understanding Variational Autoencoders with Intrinsic Dimension and Information Imbalance
by: Camboulin, Charles, et al.
Published: (2024)
by: Camboulin, Charles, et al.
Published: (2024)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
Symbolic Graph Networks for Robust PDE Discovery from Noisy Sparse Data
by: Chen, Xingyu, et al.
Published: (2026)
by: Chen, Xingyu, et al.
Published: (2026)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
by: Anderson, Samuel Cyrenius
Published: (2026)
by: Anderson, Samuel Cyrenius
Published: (2026)
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
by: Haklidir, Mehmet
Published: (2026)
by: Haklidir, Mehmet
Published: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026)
by: DeLeeuw, Caleb
Published: (2026)
SLAY: Geometry-Aware Spherical Linearized Attention with Yat-Kernel
by: Luna, Jose Miguel, et al.
Published: (2026)
by: Luna, Jose Miguel, et al.
Published: (2026)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
The Lattice Geometry of Neural Network Quantization -- A Short Equivalence Proof of GPTQ and Babai's Algorithm
by: Birnick, Johann
Published: (2025)
by: Birnick, Johann
Published: (2025)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
by: Kumar, Sachin
Published: (2026)
by: Kumar, Sachin
Published: (2026)
Reconstructing 12-Lead ECG from 3-Lead ECG using Variational Autoencoder to Improve Cardiac Disease Detection of Wearable ECG Devices
by: Guan, Xinyan, et al.
Published: (2025)
by: Guan, Xinyan, et al.
Published: (2025)
Evaluating Model Robustness Using Adaptive Sparse L0 Regularization
by: Liu, Weiyou, et al.
Published: (2024)
by: Liu, Weiyou, et al.
Published: (2024)
Car Sensors Health Monitoring by Verification Based on Autoencoder and Random Forest Regression
by: Torkhesari, Sahar, et al.
Published: (2025)
by: Torkhesari, Sahar, et al.
Published: (2025)
Distribution Consistency based Self-Training for Graph Neural Networks with Sparse Labels
by: Wang, Fali, et al.
Published: (2024)
by: Wang, Fali, et al.
Published: (2024)
HierCVAE: Hierarchical Attention-Driven Conditional Variational Autoencoders for Multi-Scale Temporal Modeling
by: Wu, Yao
Published: (2025)
by: Wu, Yao
Published: (2025)
UWM-JEPA: Predictive World Models That Imagine in Belief Space
by: Radha, Santosh Kumar, et al.
Published: (2026)
by: Radha, Santosh Kumar, et al.
Published: (2026)
SPACeR: Self-Play Anchoring with Centralized Reference Models
by: Chang, Wei-Jer, et al.
Published: (2025)
by: Chang, Wei-Jer, et al.
Published: (2025)
Learning to Construct Knowledge through Sparse Reference Selection with Reinforcement Learning
by: Yin, Shao-An
Published: (2025)
by: Yin, Shao-An
Published: (2025)
CortexCompile: Harnessing Cortical-Inspired Architectures for Enhanced Multi-Agent NLP Code Synthesis
by: Ramachandran, Gautham, et al.
Published: (2024)
by: Ramachandran, Gautham, et al.
Published: (2024)
DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming
by: Malarkkan, Arun Vignesh, et al.
Published: (2025)
by: Malarkkan, Arun Vignesh, et al.
Published: (2025)
Solve it with EASE
by: Viktorin, Adam, et al.
Published: (2025)
by: Viktorin, Adam, et al.
Published: (2025)
Faster by Design: Interactive Aerodynamics via Neural Surrogates Trained on Expert-Validated CFD
by: Thumiger, Nicholas, et al.
Published: (2026)
by: Thumiger, Nicholas, et al.
Published: (2026)
TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
by: Stripelis, Dimitris, et al.
Published: (2024)
by: Stripelis, Dimitris, et al.
Published: (2024)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
by: Wang, Zhixiang
Published: (2025)
by: Wang, Zhixiang
Published: (2025)
DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer
by: Simon, Sona Elza, et al.
Published: (2025)
by: Simon, Sona Elza, et al.
Published: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
by: Zhang, Xuzhi, et al.
Published: (2025)
by: Zhang, Xuzhi, et al.
Published: (2025)
Brittlebench: Quantifying LLM robustness via prompt sensitivity
by: Romanou, Angelika, et al.
Published: (2026)
by: Romanou, Angelika, et al.
Published: (2026)
From Numbers to Prompts: A Cognitive Symbolic Transition Mechanism for Lightweight Time-Series Forecasting
by: Yoon, Namkyung, et al.
Published: (2026)
by: Yoon, Namkyung, et al.
Published: (2026)
CoATA: Effective Co-Augmentation of Topology and Attribute for Graph Neural Networks
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Data Valuation by Fusing Global and Local Statistical Information
by: Zhou, Xiaoling, et al.
Published: (2024)
by: Zhou, Xiaoling, et al.
Published: (2024)
CroSel: Cross Selection of Confident Pseudo Labels for Partial-Label Learning
by: Tian, Shiyu, et al.
Published: (2023)
by: Tian, Shiyu, et al.
Published: (2023)
DSF-GAN: DownStream Feedback Generative Adversarial Network
by: Perets, Oriel, et al.
Published: (2024)
by: Perets, Oriel, et al.
Published: (2024)
Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic
by: Ji, Tianying, et al.
Published: (2023)
by: Ji, Tianying, et al.
Published: (2023)
Similar Items
-
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents
by: Levinson, Matthew
Published: (2026) -
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025) -
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026) -
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
by: Chen, Siyu, et al.
Published: (2025) -
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)