Investigating task-specific prompts and sparse autoencoders for activation monitoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tillman, Henk, Mossing, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling and evaluating sparse autoencoders
von: Gao, Leo, et al.
Veröffentlicht: (2024)
von: Gao, Leo, et al.
Veröffentlicht: (2024)
Weight-sparse transformers have interpretable circuits
von: Gao, Leo, et al.
Veröffentlicht: (2025)
von: Gao, Leo, et al.
Veröffentlicht: (2025)
Understanding sparse autoencoder scaling in the presence of feature manifolds
von: Michaud, Eric J., et al.
Veröffentlicht: (2025)
von: Michaud, Eric J., et al.
Veröffentlicht: (2025)
Decomposing multimodal embedding spaces with group-sparse autoencoders
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2026)
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2026)
Applying sparse autoencoders to unlearn knowledge in language models
von: Farrell, Eoin, et al.
Veröffentlicht: (2024)
von: Farrell, Eoin, et al.
Veröffentlicht: (2024)
Steering CLIP's vision transformer with sparse autoencoders
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
von: Joseph, Sonia, et al.
Veröffentlicht: (2025)
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
von: Bouzid, Kenza, et al.
Veröffentlicht: (2025)
von: Bouzid, Kenza, et al.
Veröffentlicht: (2025)
Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes
von: Riegler, Michael A., et al.
Veröffentlicht: (2026)
von: Riegler, Michael A., et al.
Veröffentlicht: (2026)
Can sparse autoencoders make sense of gene expression latent variable models?
von: Schuster, Viktoria
Veröffentlicht: (2024)
von: Schuster, Viktoria
Veröffentlicht: (2024)
Can sparse autoencoders be used to decompose and interpret steering vectors?
von: Mayne, Harry, et al.
Veröffentlicht: (2024)
von: Mayne, Harry, et al.
Veröffentlicht: (2024)
Transformer autoencoder with local attention for sparse and irregular time series with application on risk estimation
von: Rodis, Panteleimon
Veröffentlicht: (2026)
von: Rodis, Panteleimon
Veröffentlicht: (2026)
Theoretically informed selection of latent activation in autoencoder based recommender systems
von: Susman, Aviad
Veröffentlicht: (2024)
von: Susman, Aviad
Veröffentlicht: (2024)
Do different prompting methods yield a common task representation in language models?
von: Davidson, Guy, et al.
Veröffentlicht: (2025)
von: Davidson, Guy, et al.
Veröffentlicht: (2025)
Learning task-specific predictive models for scientific computing
von: Yin, Jianyuan, et al.
Veröffentlicht: (2025)
von: Yin, Jianyuan, et al.
Veröffentlicht: (2025)
P2DT: Mitigating Forgetting in task-incremental Learning with progressive prompt Decision Transformer
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
Why should autoencoders work?
von: Kvalheim, Matthew D., et al.
Veröffentlicht: (2023)
von: Kvalheim, Matthew D., et al.
Veröffentlicht: (2023)
Transmuting prompts into weights
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025)
Bilinear autoencoders find interpretable manifolds
von: Dooms, Thomas, et al.
Veröffentlicht: (2026)
von: Dooms, Thomas, et al.
Veröffentlicht: (2026)
Matching aggregate posteriors in the variational autoencoder
von: Saha, Surojit, et al.
Veröffentlicht: (2023)
von: Saha, Surojit, et al.
Veröffentlicht: (2023)
SeisT: A foundational deep learning model for earthquake monitoring tasks
von: Li, Sen, et al.
Veröffentlicht: (2023)
von: Li, Sen, et al.
Veröffentlicht: (2023)
Why is prompting hard? Understanding prompts on binary sequence predictors
von: Wenliang, Li Kevin, et al.
Veröffentlicht: (2025)
von: Wenliang, Li Kevin, et al.
Veröffentlicht: (2025)
Wavelet-Filtering of Symbolic Music Representations for Folk Tune Segmentation and Classification
von: Velarde, Gissel, et al.
Veröffentlicht: (2025)
von: Velarde, Gissel, et al.
Veröffentlicht: (2025)
An approach to melodic segmentation and classification based on filtering with the Haar-wavelet
von: Velarde, Gissel, et al.
Veröffentlicht: (2025)
von: Velarde, Gissel, et al.
Veröffentlicht: (2025)
Complex variational autoencoders admit Kähler structure
von: Gracyk, Andrew
Veröffentlicht: (2025)
von: Gracyk, Andrew
Veröffentlicht: (2025)
Adaptive sampling using variational autoencoder and reinforcement learning
von: Rasheed, Adil, et al.
Veröffentlicht: (2025)
von: Rasheed, Adil, et al.
Veröffentlicht: (2025)
DANAE: a denoising autoencoder for underwater attitude estimation
von: Russo, Paolo, et al.
Veröffentlicht: (2020)
von: Russo, Paolo, et al.
Veröffentlicht: (2020)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
Large language models surpass domain-specific architectures for antepartum electronic fetal monitoring analysis
von: Wong, Sheng, et al.
Veröffentlicht: (2025)
von: Wong, Sheng, et al.
Veröffentlicht: (2025)
Dynamic Embeddings with Task-Oriented prompting
von: Balloccu, Allmin, et al.
Veröffentlicht: (2024)
von: Balloccu, Allmin, et al.
Veröffentlicht: (2024)
Ransomware detection using stacked autoencoder for feature selection
von: Nkongolo, Mike, et al.
Veröffentlicht: (2024)
von: Nkongolo, Mike, et al.
Veröffentlicht: (2024)
Variational autoencoder-based neural network model compression
von: Cheng, Liang, et al.
Veröffentlicht: (2024)
von: Cheng, Liang, et al.
Veröffentlicht: (2024)
An autoencoder for compressing angle-resolved photoemission spectroscopy data
von: Agustsson, Steinn Ymir, et al.
Veröffentlicht: (2024)
von: Agustsson, Steinn Ymir, et al.
Veröffentlicht: (2024)
Data-driven identification of nonlinear dynamical systems with LSTM autoencoders and Normalizing Flows
von: Rostamijavanani, Abdolvahhab, et al.
Veröffentlicht: (2025)
von: Rostamijavanani, Abdolvahhab, et al.
Veröffentlicht: (2025)
A tutorial on multi-view autoencoders using the multi-view-AE library
von: Aguila, Ana Lawry, et al.
Veröffentlicht: (2024)
von: Aguila, Ana Lawry, et al.
Veröffentlicht: (2024)
Ask, and it shall be given: On the Turing completeness of prompting
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
Variational autoencoders understand knot topology
von: Braghetto, Anna, et al.
Veröffentlicht: (2025)
von: Braghetto, Anna, et al.
Veröffentlicht: (2025)
Derivative-based regularization for regression
von: Lopedoto, Enrico, et al.
Veröffentlicht: (2024)
von: Lopedoto, Enrico, et al.
Veröffentlicht: (2024)
When to Accept Automated Predictions and When to Defer to Human Judgment?
von: Sikar, Daniel, et al.
Veröffentlicht: (2024)
von: Sikar, Daniel, et al.
Veröffentlicht: (2024)
Predicted-occupancy grids for vehicle safety applications based on autoencoders and the Random Forest algorithm
von: Nadarajan, Parthasarathy, et al.
Veröffentlicht: (2025)
von: Nadarajan, Parthasarathy, et al.
Veröffentlicht: (2025)
Multi-fidelity aerodynamic data fusion by autoencoder transfer learning
von: Nieto-Centenero, Javier, et al.
Veröffentlicht: (2025)
von: Nieto-Centenero, Javier, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling and evaluating sparse autoencoders
von: Gao, Leo, et al.
Veröffentlicht: (2024) -
Weight-sparse transformers have interpretable circuits
von: Gao, Leo, et al.
Veröffentlicht: (2025) -
Understanding sparse autoencoder scaling in the presence of feature manifolds
von: Michaud, Eric J., et al.
Veröffentlicht: (2025) -
Decomposing multimodal embedding spaces with group-sparse autoencoders
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2026) -
Applying sparse autoencoders to unlearn knowledge in language models
von: Farrell, Eoin, et al.
Veröffentlicht: (2024)