Transformer Normalisation Layers and the Independence of Semantic Subspaces
Fuente:
arXiv
Salvato in:
| Autori principali: | Menary, Stephen, Kaski, Samuel, Freitas, Andre |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
In-Context Black-Box Optimization with Unreliable Feedback
di: Blumer, Nicolas Samuel, et al.
Pubblicazione: (2026)
di: Blumer, Nicolas Samuel, et al.
Pubblicazione: (2026)
More Than Irrational: Modeling Belief-Biased Agents
di: Zhu, Yifan, et al.
Pubblicazione: (2025)
di: Zhu, Yifan, et al.
Pubblicazione: (2025)
In-Context Multi-Objective Optimization
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
From Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal Transport
di: Bouniot, Quentin, et al.
Pubblicazione: (2023)
di: Bouniot, Quentin, et al.
Pubblicazione: (2023)
Gradient Regularized Natural Gradients
di: Dash, Satya Prakash, et al.
Pubblicazione: (2026)
di: Dash, Satya Prakash, et al.
Pubblicazione: (2026)
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
di: Huo, Yingxiao, et al.
Pubblicazione: (2026)
di: Huo, Yingxiao, et al.
Pubblicazione: (2026)
Gated Subspace Inference for Transformer Acceleration
di: Thomas, Stephen J.
Pubblicazione: (2026)
di: Thomas, Stephen J.
Pubblicazione: (2026)
Cooperative Bayesian Optimization for Imperfect Agents
di: Khoshvishkaie, Ali, et al.
Pubblicazione: (2024)
di: Khoshvishkaie, Ali, et al.
Pubblicazione: (2024)
Lifted Model Construction without Normalisation: A Vectorised Approach to Exploit Symmetries in Factor Graphs
di: Luttermann, Malte, et al.
Pubblicazione: (2024)
di: Luttermann, Malte, et al.
Pubblicazione: (2024)
Towards modeling evolving longitudinal health trajectories with a transformer-based deep learning model
di: Moen, Hans, et al.
Pubblicazione: (2024)
di: Moen, Hans, et al.
Pubblicazione: (2024)
Interpretable-by-Design Transformers via Architectural Stream Independence
di: Kerce, Clayton, et al.
Pubblicazione: (2026)
di: Kerce, Clayton, et al.
Pubblicazione: (2026)
Model Merging in the Essential Subspace
di: Li, Longhua, et al.
Pubblicazione: (2026)
di: Li, Longhua, et al.
Pubblicazione: (2026)
Dynamic Layer Tying for Parameter-Efficient Transformers
di: Hay, Tamir David, et al.
Pubblicazione: (2024)
di: Hay, Tamir David, et al.
Pubblicazione: (2024)
Layer Specialization Underlying Compositional Reasoning in Transformers
di: Liu, Jing
Pubblicazione: (2025)
di: Liu, Jing
Pubblicazione: (2025)
Mixture-of-Subspaces in Low-Rank Adaptation
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
Uncoupled Learning of Differential Stackelberg Equilibria with Commitments
di: Loftin, Robert, et al.
Pubblicazione: (2023)
di: Loftin, Robert, et al.
Pubblicazione: (2023)
Universal Approximation Theorem for a Single-Layer Transformer
di: Gumaan, Esmail
Pubblicazione: (2025)
di: Gumaan, Esmail
Pubblicazione: (2025)
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
di: Hsu, Hsin-Ling, et al.
Pubblicazione: (2026)
di: Hsu, Hsin-Ling, et al.
Pubblicazione: (2026)
Test-time Adaptation for Regression by Subspace Alignment
di: Adachi, Kazuki, et al.
Pubblicazione: (2024)
di: Adachi, Kazuki, et al.
Pubblicazione: (2024)
Conjunction Subspaces Test for Conformal and Selective Classification
di: He, Zengyou, et al.
Pubblicazione: (2024)
di: He, Zengyou, et al.
Pubblicazione: (2024)
Machine Unlearning in Low-Dimensional Feature Subspace
di: Fang, Kun, et al.
Pubblicazione: (2026)
di: Fang, Kun, et al.
Pubblicazione: (2026)
Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs
di: Xia, Wei
Pubblicazione: (2025)
di: Xia, Wei
Pubblicazione: (2025)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Engineering Verifiable Modularity in Transformers via Per-Layer Supervision
di: Kerce, J. Clayton
Pubblicazione: (2026)
di: Kerce, J. Clayton
Pubblicazione: (2026)
Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
di: Xu, Jin, et al.
Pubblicazione: (2026)
di: Xu, Jin, et al.
Pubblicazione: (2026)
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
di: Sun, Haoyuan, et al.
Pubblicazione: (2025)
di: Sun, Haoyuan, et al.
Pubblicazione: (2025)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
di: Yu, Ziming, et al.
Pubblicazione: (2024)
di: Yu, Ziming, et al.
Pubblicazione: (2024)
Rethinking Inter-LoRA Orthogonality in Adapter Merging: Insights from Orthogonal Monte Carlo Dropout
di: Zhang, Andi, et al.
Pubblicazione: (2025)
di: Zhang, Andi, et al.
Pubblicazione: (2025)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
di: Wu, Qitian, et al.
Pubblicazione: (2024)
di: Wu, Qitian, et al.
Pubblicazione: (2024)
Uncovering Layer-Dependent Activation Sparsity Patterns in ReLU Transformers
di: Wild, Cody, et al.
Pubblicazione: (2024)
di: Wild, Cody, et al.
Pubblicazione: (2024)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
di: Gavito, Andrea Treviño, et al.
Pubblicazione: (2023)
di: Gavito, Andrea Treviño, et al.
Pubblicazione: (2023)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
di: Poduval, Prathyush, et al.
Pubblicazione: (2026)
di: Poduval, Prathyush, et al.
Pubblicazione: (2026)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
di: Cutler, Dylan, et al.
Pubblicazione: (2025)
di: Cutler, Dylan, et al.
Pubblicazione: (2025)
Transformer Is Inherently a Causal Learner
di: Wang, Xinyue, et al.
Pubblicazione: (2026)
di: Wang, Xinyue, et al.
Pubblicazione: (2026)
InjectTST: A Transformer Method of Injecting Global Information into Independent Channels for Long Time Series Forecasting
di: Chi, Ce, et al.
Pubblicazione: (2024)
di: Chi, Ce, et al.
Pubblicazione: (2024)
Subspace-based Approximate Hessian Method for Zeroth-Order Optimization
di: Kim, Dongyoon, et al.
Pubblicazione: (2025)
di: Kim, Dongyoon, et al.
Pubblicazione: (2025)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
di: Freytes, Luis Rosario
Pubblicazione: (2026)
di: Freytes, Luis Rosario
Pubblicazione: (2026)
Hierarchical Transformers are Efficient Meta-Reinforcement Learners
di: Shala, Gresa, et al.
Pubblicazione: (2024)
di: Shala, Gresa, et al.
Pubblicazione: (2024)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
di: Arturi, Daniel Aarao Reis, et al.
Pubblicazione: (2025)
di: Arturi, Daniel Aarao Reis, et al.
Pubblicazione: (2025)
Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification
di: Duan, Zenghao, et al.
Pubblicazione: (2026)
di: Duan, Zenghao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
In-Context Black-Box Optimization with Unreliable Feedback
di: Blumer, Nicolas Samuel, et al.
Pubblicazione: (2026) -
More Than Irrational: Modeling Belief-Biased Agents
di: Zhu, Yifan, et al.
Pubblicazione: (2025) -
In-Context Multi-Objective Optimization
di: Zhang, Xinyu, et al.
Pubblicazione: (2025) -
From Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal Transport
di: Bouniot, Quentin, et al.
Pubblicazione: (2023) -
Gradient Regularized Natural Gradients
di: Dash, Satya Prakash, et al.
Pubblicazione: (2026)