Investigating task-specific prompts and sparse autoencoders for activation monitoring
Fuente:
arXiv
Salvato in:
| Autori principali: | Tillman, Henk, Mossing, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling and evaluating sparse autoencoders
di: Gao, Leo, et al.
Pubblicazione: (2024)
di: Gao, Leo, et al.
Pubblicazione: (2024)
Weight-sparse transformers have interpretable circuits
di: Gao, Leo, et al.
Pubblicazione: (2025)
di: Gao, Leo, et al.
Pubblicazione: (2025)
Understanding sparse autoencoder scaling in the presence of feature manifolds
di: Michaud, Eric J., et al.
Pubblicazione: (2025)
di: Michaud, Eric J., et al.
Pubblicazione: (2025)
Decomposing multimodal embedding spaces with group-sparse autoencoders
di: Kaushik, Chiraag, et al.
Pubblicazione: (2026)
di: Kaushik, Chiraag, et al.
Pubblicazione: (2026)
Applying sparse autoencoders to unlearn knowledge in language models
di: Farrell, Eoin, et al.
Pubblicazione: (2024)
di: Farrell, Eoin, et al.
Pubblicazione: (2024)
Steering CLIP's vision transformer with sparse autoencoders
di: Joseph, Sonia, et al.
Pubblicazione: (2025)
di: Joseph, Sonia, et al.
Pubblicazione: (2025)
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
di: Bouzid, Kenza, et al.
Pubblicazione: (2025)
di: Bouzid, Kenza, et al.
Pubblicazione: (2025)
Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes
di: Riegler, Michael A., et al.
Pubblicazione: (2026)
di: Riegler, Michael A., et al.
Pubblicazione: (2026)
Can sparse autoencoders make sense of gene expression latent variable models?
di: Schuster, Viktoria
Pubblicazione: (2024)
di: Schuster, Viktoria
Pubblicazione: (2024)
Can sparse autoencoders be used to decompose and interpret steering vectors?
di: Mayne, Harry, et al.
Pubblicazione: (2024)
di: Mayne, Harry, et al.
Pubblicazione: (2024)
Transformer autoencoder with local attention for sparse and irregular time series with application on risk estimation
di: Rodis, Panteleimon
Pubblicazione: (2026)
di: Rodis, Panteleimon
Pubblicazione: (2026)
Theoretically informed selection of latent activation in autoencoder based recommender systems
di: Susman, Aviad
Pubblicazione: (2024)
di: Susman, Aviad
Pubblicazione: (2024)
Do different prompting methods yield a common task representation in language models?
di: Davidson, Guy, et al.
Pubblicazione: (2025)
di: Davidson, Guy, et al.
Pubblicazione: (2025)
Learning task-specific predictive models for scientific computing
di: Yin, Jianyuan, et al.
Pubblicazione: (2025)
di: Yin, Jianyuan, et al.
Pubblicazione: (2025)
P2DT: Mitigating Forgetting in task-incremental Learning with progressive prompt Decision Transformer
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Wang, Zhiyuan, et al.
Pubblicazione: (2024)
Why should autoencoders work?
di: Kvalheim, Matthew D., et al.
Pubblicazione: (2023)
di: Kvalheim, Matthew D., et al.
Pubblicazione: (2023)
Transmuting prompts into weights
di: Mazzawi, Hanna, et al.
Pubblicazione: (2025)
di: Mazzawi, Hanna, et al.
Pubblicazione: (2025)
Bilinear autoencoders find interpretable manifolds
di: Dooms, Thomas, et al.
Pubblicazione: (2026)
di: Dooms, Thomas, et al.
Pubblicazione: (2026)
Matching aggregate posteriors in the variational autoencoder
di: Saha, Surojit, et al.
Pubblicazione: (2023)
di: Saha, Surojit, et al.
Pubblicazione: (2023)
SeisT: A foundational deep learning model for earthquake monitoring tasks
di: Li, Sen, et al.
Pubblicazione: (2023)
di: Li, Sen, et al.
Pubblicazione: (2023)
Why is prompting hard? Understanding prompts on binary sequence predictors
di: Wenliang, Li Kevin, et al.
Pubblicazione: (2025)
di: Wenliang, Li Kevin, et al.
Pubblicazione: (2025)
Wavelet-Filtering of Symbolic Music Representations for Folk Tune Segmentation and Classification
di: Velarde, Gissel, et al.
Pubblicazione: (2025)
di: Velarde, Gissel, et al.
Pubblicazione: (2025)
An approach to melodic segmentation and classification based on filtering with the Haar-wavelet
di: Velarde, Gissel, et al.
Pubblicazione: (2025)
di: Velarde, Gissel, et al.
Pubblicazione: (2025)
Complex variational autoencoders admit Kähler structure
di: Gracyk, Andrew
Pubblicazione: (2025)
di: Gracyk, Andrew
Pubblicazione: (2025)
Adaptive sampling using variational autoencoder and reinforcement learning
di: Rasheed, Adil, et al.
Pubblicazione: (2025)
di: Rasheed, Adil, et al.
Pubblicazione: (2025)
DANAE: a denoising autoencoder for underwater attitude estimation
di: Russo, Paolo, et al.
Pubblicazione: (2020)
di: Russo, Paolo, et al.
Pubblicazione: (2020)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
di: Béchard, Patrice, et al.
Pubblicazione: (2025)
di: Béchard, Patrice, et al.
Pubblicazione: (2025)
Large language models surpass domain-specific architectures for antepartum electronic fetal monitoring analysis
di: Wong, Sheng, et al.
Pubblicazione: (2025)
di: Wong, Sheng, et al.
Pubblicazione: (2025)
Dynamic Embeddings with Task-Oriented prompting
di: Balloccu, Allmin, et al.
Pubblicazione: (2024)
di: Balloccu, Allmin, et al.
Pubblicazione: (2024)
Ransomware detection using stacked autoencoder for feature selection
di: Nkongolo, Mike, et al.
Pubblicazione: (2024)
di: Nkongolo, Mike, et al.
Pubblicazione: (2024)
Variational autoencoder-based neural network model compression
di: Cheng, Liang, et al.
Pubblicazione: (2024)
di: Cheng, Liang, et al.
Pubblicazione: (2024)
An autoencoder for compressing angle-resolved photoemission spectroscopy data
di: Agustsson, Steinn Ymir, et al.
Pubblicazione: (2024)
di: Agustsson, Steinn Ymir, et al.
Pubblicazione: (2024)
Data-driven identification of nonlinear dynamical systems with LSTM autoencoders and Normalizing Flows
di: Rostamijavanani, Abdolvahhab, et al.
Pubblicazione: (2025)
di: Rostamijavanani, Abdolvahhab, et al.
Pubblicazione: (2025)
A tutorial on multi-view autoencoders using the multi-view-AE library
di: Aguila, Ana Lawry, et al.
Pubblicazione: (2024)
di: Aguila, Ana Lawry, et al.
Pubblicazione: (2024)
Ask, and it shall be given: On the Turing completeness of prompting
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
Variational autoencoders understand knot topology
di: Braghetto, Anna, et al.
Pubblicazione: (2025)
di: Braghetto, Anna, et al.
Pubblicazione: (2025)
Derivative-based regularization for regression
di: Lopedoto, Enrico, et al.
Pubblicazione: (2024)
di: Lopedoto, Enrico, et al.
Pubblicazione: (2024)
When to Accept Automated Predictions and When to Defer to Human Judgment?
di: Sikar, Daniel, et al.
Pubblicazione: (2024)
di: Sikar, Daniel, et al.
Pubblicazione: (2024)
Predicted-occupancy grids for vehicle safety applications based on autoencoders and the Random Forest algorithm
di: Nadarajan, Parthasarathy, et al.
Pubblicazione: (2025)
di: Nadarajan, Parthasarathy, et al.
Pubblicazione: (2025)
Multi-fidelity aerodynamic data fusion by autoencoder transfer learning
di: Nieto-Centenero, Javier, et al.
Pubblicazione: (2025)
di: Nieto-Centenero, Javier, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Scaling and evaluating sparse autoencoders
di: Gao, Leo, et al.
Pubblicazione: (2024) -
Weight-sparse transformers have interpretable circuits
di: Gao, Leo, et al.
Pubblicazione: (2025) -
Understanding sparse autoencoder scaling in the presence of feature manifolds
di: Michaud, Eric J., et al.
Pubblicazione: (2025) -
Decomposing multimodal embedding spaces with group-sparse autoencoders
di: Kaushik, Chiraag, et al.
Pubblicazione: (2026) -
Applying sparse autoencoders to unlearn knowledge in language models
di: Farrell, Eoin, et al.
Pubblicazione: (2024)