Transformers without Normalization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Jiachen, Chen, Xinlei, He, Kaiming, LeCun, Yann, Liu, Zhuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stronger Normalization-Free Transformers
by: Chen, Mingzhi, et al.
Published: (2025)
by: Chen, Mingzhi, et al.
Published: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023)
by: Zhu, Jiachen, et al.
Published: (2023)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
Revisiting Feature Prediction for Learning Visual Representations from Video
by: Bardes, Adrien, et al.
Published: (2024)
by: Bardes, Adrien, et al.
Published: (2024)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
by: Zhai, Yuexiang, et al.
Published: (2024)
by: Zhai, Yuexiang, et al.
Published: (2024)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
by: Yun, Zeyu, et al.
Published: (2023)
by: Yun, Zeyu, et al.
Published: (2023)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
A hierarchical loss and its problems when classifying non-hierarchically
by: Wu, Cinna, et al.
Published: (2017)
by: Wu, Cinna, et al.
Published: (2017)
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
by: Rudman, William, et al.
Published: (2025)
by: Rudman, William, et al.
Published: (2025)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity
by: Jiang, Jiachen, et al.
Published: (2024)
by: Jiang, Jiachen, et al.
Published: (2024)
Deconstructing Denoising Diffusion Models for Self-Supervised Learning
by: Chen, Xinlei, et al.
Published: (2024)
by: Chen, Xinlei, et al.
Published: (2024)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
Diffuse and Disperse: Image Generation with Representation Regularization
by: Wang, Runqian, et al.
Published: (2025)
by: Wang, Runqian, et al.
Published: (2025)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
by: Denton, Remi, et al.
Published: (2014)
by: Denton, Remi, et al.
Published: (2014)
CHASM: Unveiling Covert Advertisements on Chinese Social Media
by: Zheng, Jingyi, et al.
Published: (2026)
by: Zheng, Jingyi, et al.
Published: (2026)
Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models
by: Fu, Shuai, et al.
Published: (2024)
by: Fu, Shuai, et al.
Published: (2024)
Physical Property Understanding from Language-Embedded Feature Fields
by: Zhai, Albert J., et al.
Published: (2024)
by: Zhai, Albert J., et al.
Published: (2024)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
A General and Efficient Training for Transformer via Token Expansion
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
by: Agashe, Saaket, et al.
Published: (2025)
by: Agashe, Saaket, et al.
Published: (2025)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
PooDLe: Pooled and dense self-supervised learning from naturalistic videos
by: Wang, Alex N., et al.
Published: (2024)
by: Wang, Alex N., et al.
Published: (2024)
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
by: Wang, Lirui, et al.
Published: (2024)
by: Wang, Lirui, et al.
Published: (2024)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
by: Javadi, Farnoosh, et al.
Published: (2023)
by: Javadi, Farnoosh, et al.
Published: (2023)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
A Decade's Battle on Dataset Bias: Are We There Yet?
by: Liu, Zhuang, et al.
Published: (2024)
by: Liu, Zhuang, et al.
Published: (2024)
REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations
by: Sushko, Peter, et al.
Published: (2025)
by: Sushko, Peter, et al.
Published: (2025)
Similar Items
-
Stronger Normalization-Free Transformers
by: Chen, Mingzhi, et al.
Published: (2025) -
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025) -
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024) -
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023) -
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025)