Salvato in:
| Autore principale: | Poupart, Yoann |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2406.04028 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TDHook: A Lightweight Framework for Interpretability
di: Poupart, Yoann
Pubblicazione: (2025)
di: Poupart, Yoann
Pubblicazione: (2025)
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
di: Poupart, Yoann, et al.
Pubblicazione: (2025)
di: Poupart, Yoann, et al.
Pubblicazione: (2025)
Iterative Inference in a Chess-Playing Neural Network
di: Sandmann, Elias, et al.
Pubblicazione: (2025)
di: Sandmann, Elias, et al.
Pubblicazione: (2025)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
di: Jenner, Erik, et al.
Pubblicazione: (2024)
di: Jenner, Erik, et al.
Pubblicazione: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
di: Tolooshams, Bahareh, et al.
Pubblicazione: (2025)
Causal Interpretation of Sparse Autoencoder Features in Vision
di: Han, Sangyu, et al.
Pubblicazione: (2025)
di: Han, Sangyu, et al.
Pubblicazione: (2025)
Neural Network-based Information Set Weighting for Playing Reconnaissance Blind Chess
di: Bertram, Timo, et al.
Pubblicazione: (2024)
di: Bertram, Timo, et al.
Pubblicazione: (2024)
Mixture of Masters: Sparse Chess Language Models with Player Routing
di: Frisoni, Giacomo, et al.
Pubblicazione: (2026)
di: Frisoni, Giacomo, et al.
Pubblicazione: (2026)
Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
di: Jeong, Jihwan, et al.
Pubblicazione: (2025)
di: Jeong, Jihwan, et al.
Pubblicazione: (2025)
UniMaia: Steering Chess Policies with Language for Human-like Play
di: Siu, Sherman, et al.
Pubblicazione: (2026)
di: Siu, Sherman, et al.
Pubblicazione: (2026)
Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
di: Parsan, Nithin, et al.
Pubblicazione: (2025)
di: Parsan, Nithin, et al.
Pubblicazione: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
di: Jiang, Nick, et al.
Pubblicazione: (2025)
di: Jiang, Nick, et al.
Pubblicazione: (2025)
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
di: Klenitskiy, Anton, et al.
Pubblicazione: (2025)
di: Klenitskiy, Anton, et al.
Pubblicazione: (2025)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
di: Yeung, Calvin, et al.
Pubblicazione: (2026)
di: Yeung, Calvin, et al.
Pubblicazione: (2026)
ChessQA: Evaluating Large Language Models for Chess Understanding
di: Wen, Qianfeng, et al.
Pubblicazione: (2025)
di: Wen, Qianfeng, et al.
Pubblicazione: (2025)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
di: Ruoss, Anian, et al.
Pubblicazione: (2024)
di: Ruoss, Anian, et al.
Pubblicazione: (2024)
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
LouvreSAE: Sparse Autoencoders for Interpretable and Controllable Style Transfer
di: Panda, Raina, et al.
Pubblicazione: (2025)
di: Panda, Raina, et al.
Pubblicazione: (2025)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
di: Zhao, Daniel, et al.
Pubblicazione: (2025)
di: Zhao, Daniel, et al.
Pubblicazione: (2025)
Why Online Reinforcement Learning is Causal
di: Schulte, Oliver, et al.
Pubblicazione: (2024)
di: Schulte, Oliver, et al.
Pubblicazione: (2024)
Sparse Autoencoders, Again?
di: Lu, Yin, et al.
Pubblicazione: (2025)
di: Lu, Yin, et al.
Pubblicazione: (2025)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
di: Li, Yuhan, et al.
Pubblicazione: (2026)
di: Li, Yuhan, et al.
Pubblicazione: (2026)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
di: Balagansky, Nikita, et al.
Pubblicazione: (2025)
di: Balagansky, Nikita, et al.
Pubblicazione: (2025)
SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations
di: Kim, Taehan, et al.
Pubblicazione: (2025)
di: Kim, Taehan, et al.
Pubblicazione: (2025)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
di: Dokme, Atahan, et al.
Pubblicazione: (2026)
di: Dokme, Atahan, et al.
Pubblicazione: (2026)
SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability
di: Nasiri-Sarvi, Ali, et al.
Pubblicazione: (2025)
di: Nasiri-Sarvi, Ali, et al.
Pubblicazione: (2025)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
di: Shu, Dong, et al.
Pubblicazione: (2025)
di: Shu, Dong, et al.
Pubblicazione: (2025)
Complete Chess Games Enable LLM Become A Chess Master
di: Zhang, Yinqi, et al.
Pubblicazione: (2025)
di: Zhang, Yinqi, et al.
Pubblicazione: (2025)
Measuring Sparse Autoencoder Feature Sensitivity
di: Tian, Claire, et al.
Pubblicazione: (2025)
di: Tian, Claire, et al.
Pubblicazione: (2025)
Generating Creative Chess Puzzles
di: Feng, Xidong, et al.
Pubblicazione: (2025)
di: Feng, Xidong, et al.
Pubblicazione: (2025)
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
di: Wang, Sai, et al.
Pubblicazione: (2025)
di: Wang, Sai, et al.
Pubblicazione: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
di: Chen, Xi, et al.
Pubblicazione: (2025)
di: Chen, Xi, et al.
Pubblicazione: (2025)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
di: Liu, Jincheng, et al.
Pubblicazione: (2025)
di: Liu, Jincheng, et al.
Pubblicazione: (2025)
Constrain Alignment with Sparse Autoencoders
di: Yin, Qingyu, et al.
Pubblicazione: (2024)
di: Yin, Qingyu, et al.
Pubblicazione: (2024)
Are Sparse Autoencoder Benchmarks Reliable?
di: Chanin, David
Pubblicazione: (2026)
di: Chanin, David
Pubblicazione: (2026)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
di: Huang, Victor Shea-Jay, et al.
Pubblicazione: (2025)
di: Huang, Victor Shea-Jay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TDHook: A Lightweight Framework for Interpretability
di: Poupart, Yoann
Pubblicazione: (2025) -
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
di: Poupart, Yoann, et al.
Pubblicazione: (2025) -
Iterative Inference in a Chess-Playing Neural Network
di: Sandmann, Elias, et al.
Pubblicazione: (2025) -
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
di: Jenner, Erik, et al.
Pubblicazione: (2024) -
Interpreting CLIP with Hierarchical Sparse Autoencoders
di: Zaigrajew, Vladimir, et al.
Pubblicazione: (2025)