Self-Ablating Transformers: More Interpretability, Less Sparsity
Fuente:
arXiv
Guardado en:
| Autores principales: | Ferrao, Jeremias, Mikaelson, Luhan, Pepper, Keenan, Antolin, Natalia Perez-Campanero |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
World Model Agents with Change-Based Intrinsic Motivation
por: Ferrao, Jeremias, et al.
Publicado: (2025)
por: Ferrao, Jeremias, et al.
Publicado: (2025)
More is Less: Inducing Sparsity via Overparameterization
por: Chou, Hung-Hsu, et al.
Publicado: (2021)
por: Chou, Hung-Hsu, et al.
Publicado: (2021)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
por: Lermen, Simon, et al.
Publicado: (2025)
por: Lermen, Simon, et al.
Publicado: (2025)
Transformer Multivariate Forecasting: Less is More?
por: Xu, Jingjing, et al.
Publicado: (2023)
por: Xu, Jingjing, et al.
Publicado: (2023)
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
por: Pepper, Keenan, et al.
Publicado: (2026)
por: Pepper, Keenan, et al.
Publicado: (2026)
Beyond Mimicry: Preference Coherence in LLMs
por: Mikaelson, Luhan, et al.
Publicado: (2025)
por: Mikaelson, Luhan, et al.
Publicado: (2025)
Evaluating Uncertainty in Deep Gaussian Processes
por: van der Lende, Matthijs, et al.
Publicado: (2025)
por: van der Lende, Matthijs, et al.
Publicado: (2025)
Latent Adversarial Training Improves the Representation of Refusal
por: Abbas, Alexandra, et al.
Publicado: (2025)
por: Abbas, Alexandra, et al.
Publicado: (2025)
(Sometimes) Less is More: Mitigating the Complexity of Rule-based Representation for Interpretable Classification
por: Bergamin, Luca, et al.
Publicado: (2025)
por: Bergamin, Luca, et al.
Publicado: (2025)
Less is More: on the Over-Globalizing Problem in Graph Transformers
por: Xing, Yujie, et al.
Publicado: (2024)
por: Xing, Yujie, et al.
Publicado: (2024)
Supernova: Achieving More with Less in Transformer Architectures
por: Tanase, Andrei-Valentin, et al.
Publicado: (2025)
por: Tanase, Andrei-Valentin, et al.
Publicado: (2025)
Leaner Transformers: More Heads, Less Depth
por: Saratchandran, Hemanth, et al.
Publicado: (2025)
por: Saratchandran, Hemanth, et al.
Publicado: (2025)
The Steganographic Potentials of Language Models
por: Karpov, Artem, et al.
Publicado: (2025)
por: Karpov, Artem, et al.
Publicado: (2025)
Less is More: Fewer Interpretable Region via Submodular Subset Selection
por: Chen, Ruoyu, et al.
Publicado: (2024)
por: Chen, Ruoyu, et al.
Publicado: (2024)
Transformers for Green Semantic Communication: Less Energy, More Semantics
por: Mukherjee, Shubhabrata, et al.
Publicado: (2023)
por: Mukherjee, Shubhabrata, et al.
Publicado: (2023)
Less Is More -- On the Importance of Sparsification for Transformers and Graph Neural Networks for TSP
por: Lischka, Attila, et al.
Publicado: (2024)
por: Lischka, Attila, et al.
Publicado: (2024)
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
por: Brzozowski, Michał, et al.
Publicado: (2026)
por: Brzozowski, Michał, et al.
Publicado: (2026)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
por: Lou, Chao, et al.
Publicado: (2024)
por: Lou, Chao, et al.
Publicado: (2024)
Sense Less, Infer More: Agentic Multimodal Transformers for Edge Medical Intelligence
por: Zhou, Chengwei, et al.
Publicado: (2026)
por: Zhou, Chengwei, et al.
Publicado: (2026)
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
por: Chen, Ruoyu, et al.
Publicado: (2025)
por: Chen, Ruoyu, et al.
Publicado: (2025)
Train Less, Infer Faster: Efficient Model Finetuning and Compression via Structured Sparsity
por: Svirsky, Jonathan, et al.
Publicado: (2026)
por: Svirsky, Jonathan, et al.
Publicado: (2026)
Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections
por: Peng, William, et al.
Publicado: (2026)
por: Peng, William, et al.
Publicado: (2026)
Less is More: Adaptive Coverage for Synthetic Training Data
por: Tavakkol, Sasan, et al.
Publicado: (2025)
por: Tavakkol, Sasan, et al.
Publicado: (2025)
Quantize What Counts: More for Keys, Less for Values
por: Hariri, Mohsen, et al.
Publicado: (2025)
por: Hariri, Mohsen, et al.
Publicado: (2025)
Less is More: Towards Simple Graph Contrastive Learning
por: Zhao, Yanan, et al.
Publicado: (2025)
por: Zhao, Yanan, et al.
Publicado: (2025)
Homeostasis and Sparsity in Transformer
por: Kotyuzanskiy, Leonid, et al.
Publicado: (2024)
por: Kotyuzanskiy, Leonid, et al.
Publicado: (2024)
Out-of-Distribution Detection & Applications With Ablated Learned Temperature Energy
por: LeVine, Will, et al.
Publicado: (2024)
por: LeVine, Will, et al.
Publicado: (2024)
Less is More: Recursive Reasoning with Tiny Networks
por: Jolicoeur-Martineau, Alexia
Publicado: (2025)
por: Jolicoeur-Martineau, Alexia
Publicado: (2025)
No More, No Less: Least-Privilege Language Models
por: Rauba, Paulius, et al.
Publicado: (2026)
por: Rauba, Paulius, et al.
Publicado: (2026)
Rethinking Tokenization for Clinical Time Series: When Less is More
por: Attrach, Rafi Al, et al.
Publicado: (2025)
por: Attrach, Rafi Al, et al.
Publicado: (2025)
Why Less is More (Sometimes): A Theory of Data Curation
por: Dohmatob, Elvis, et al.
Publicado: (2025)
por: Dohmatob, Elvis, et al.
Publicado: (2025)
When Less is More: The LLM Scaling Paradox in Context Compression
por: Guo, Ruishan, et al.
Publicado: (2026)
por: Guo, Ruishan, et al.
Publicado: (2026)
Less is More: Efficient Model Merging with Binary Task Switch
por: Qi, Biqing, et al.
Publicado: (2024)
por: Qi, Biqing, et al.
Publicado: (2024)
Less is More: Clustered Cross-Covariance Control for Offline RL
por: Qiao, Nan, et al.
Publicado: (2026)
por: Qiao, Nan, et al.
Publicado: (2026)
Learning Interpretable PDE Representations for Generative Reconstructions with Structured Sparsity
por: Tsao, Valerie, et al.
Publicado: (2026)
por: Tsao, Valerie, et al.
Publicado: (2026)
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025)
por: Li, Xuefeng, et al.
Publicado: (2025)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
por: Wu, Shaojin, et al.
Publicado: (2025)
por: Wu, Shaojin, et al.
Publicado: (2025)
Learning More with Less: A Generalizable, Self-Supervised Framework for Privacy-Preserving Capacity Estimation with EV Charging Data
por: Arunan, Anushiya, et al.
Publicado: (2025)
por: Arunan, Anushiya, et al.
Publicado: (2025)
Distributionally Robust Self Paced Curriculum Reinforcement Learning
por: Satheesh, Anirudh, et al.
Publicado: (2025)
por: Satheesh, Anirudh, et al.
Publicado: (2025)
Spark Transformer: Reactivating Sparsity in FFN and Attention
por: You, Chong, et al.
Publicado: (2025)
por: You, Chong, et al.
Publicado: (2025)
Ejemplares similares
-
World Model Agents with Change-Based Intrinsic Motivation
por: Ferrao, Jeremias, et al.
Publicado: (2025) -
More is Less: Inducing Sparsity via Overparameterization
por: Chou, Hung-Hsu, et al.
Publicado: (2021) -
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
por: Lermen, Simon, et al.
Publicado: (2025) -
Transformer Multivariate Forecasting: Less is More?
por: Xu, Jingjing, et al.
Publicado: (2023) -
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
por: Pepper, Keenan, et al.
Publicado: (2026)