Taming Transformer Without Using Learning Rate Warmup
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Xianbiao, He, Yelin, Ye, Jiaquan, Li, Chun-Guang, Zi, Bojia, Dai, Xili, Zou, Qin, Xiao, Rong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimpleGPT: Improving GPT via A Simple Normalization Strategy
by: Chen, Marco, et al.
Published: (2026)
by: Chen, Marco, et al.
Published: (2026)
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
Refaçade: Editing Object with Given Reference Texture
by: Huang, Youze, et al.
Published: (2025)
by: Huang, Youze, et al.
Published: (2025)
Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
by: He, Wei, et al.
Published: (2026)
by: He, Wei, et al.
Published: (2026)
Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation
by: Ruan, Penghui, et al.
Published: (2026)
by: Ruan, Penghui, et al.
Published: (2026)
Delving into Muon and Beyond: Deep Analysis and Extensions
by: Qi, Xianbiao, et al.
Published: (2026)
by: Qi, Xianbiao, et al.
Published: (2026)
Exploring a Principled Framework for Deep Subspace Clustering
by: Meng, Xianghan, et al.
Published: (2025)
by: Meng, Xianghan, et al.
Published: (2025)
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
by: Huang, Yaxuan, et al.
Published: (2025)
by: Huang, Yaxuan, et al.
Published: (2025)
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
by: Lin, Jiahao, et al.
Published: (2025)
by: Lin, Jiahao, et al.
Published: (2025)
CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
by: Zi, Bojia, et al.
Published: (2024)
by: Zi, Bojia, et al.
Published: (2024)
Double Helix Diffusion for Cross-Domain Anomaly Image Generation
by: Wu, Linchun, et al.
Published: (2025)
by: Wu, Linchun, et al.
Published: (2025)
Anti-Collapse Loss for Deep Metric Learning Based on Coding Rate Metric
by: Jiang, Xiruo, et al.
Published: (2024)
by: Jiang, Xiruo, et al.
Published: (2024)
Elucidating the design space of language models for image generation
by: Liu, Xuantong, et al.
Published: (2024)
by: Liu, Xuantong, et al.
Published: (2024)
DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
by: Duan, Zheng-Peng, et al.
Published: (2025)
by: Duan, Zheng-Peng, et al.
Published: (2025)
MIA-Mind: A Multidimensional Interactive Attention Mechanism Based on MindSpore
by: Qin, Zhenkai, et al.
Published: (2025)
by: Qin, Zhenkai, et al.
Published: (2025)
TeleStyle: Content-Preserving Style Transfer in Images and Videos
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
by: Chu, Tianzhe, et al.
Published: (2023)
by: Chu, Tianzhe, et al.
Published: (2023)
WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection
by: Zi, Bojia, et al.
Published: (2021)
by: Zi, Bojia, et al.
Published: (2021)
ConsistEdit: Highly Consistent and Precise Training-free Visual Editing
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
CrashChat: A Multimodal Large Language Model for Multitask Traffic Crash Video Analysis
by: Liang, Kaidi, et al.
Published: (2025)
by: Liang, Kaidi, et al.
Published: (2025)
DPFormer: Dynamic Prompt Transformer for Continual Learning
by: Huang, Sheng-Kai, et al.
Published: (2025)
by: Huang, Sheng-Kai, et al.
Published: (2025)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
by: Wei, Guoting, et al.
Published: (2026)
by: Wei, Guoting, et al.
Published: (2026)
Multi-label Scene Classification for Autonomous Vehicles: Acquiring and Accumulating Knowledge from Diverse Datasets
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
HMPDM: A Diffusion Model for Driving Video Prediction with Historical Motion Priors
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Temporal Rate Reduction Clustering for Human Motion Segmentation
by: Meng, Xianghan, et al.
Published: (2025)
by: Meng, Xianghan, et al.
Published: (2025)
LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
Exploring Color Invariance through Image-Level Ensemble Learning
by: Gong, Yunpeng, et al.
Published: (2024)
by: Gong, Yunpeng, et al.
Published: (2024)
Taming Transformer for Emotion-Controllable Talking Face Generation
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
BAFNet: Bilateral Attention Fusion Network for Lightweight Semantic Segmentation of Urban Remote Sensing Images
by: Wang, Wentao, et al.
Published: (2024)
by: Wang, Wentao, et al.
Published: (2024)
Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression Perspective
by: Wen, Qishuai, et al.
Published: (2024)
by: Wen, Qishuai, et al.
Published: (2024)
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024)
by: Luo, Lin, et al.
Published: (2024)
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
by: Guo, Qin, et al.
Published: (2025)
by: Guo, Qin, et al.
Published: (2025)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Graph Cut-guided Maximal Coding Rate Reduction for Learning Image Embedding and Clustering
by: He, W., et al.
Published: (2024)
by: He, W., et al.
Published: (2024)
Winner Team Mia at TextVQA Challenge 2021: Vision-and-Language Representation Learning with Pre-trained Sequence-to-Sequence Model
by: Qiao, Yixuan, et al.
Published: (2021)
by: Qiao, Yixuan, et al.
Published: (2021)
Similar Items
-
SimpleGPT: Improving GPT via A Simple Normalization Strategy
by: Chen, Marco, et al.
Published: (2026) -
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
by: Qi, Xianbiao, et al.
Published: (2025) -
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025) -
Refaçade: Editing Object with Given Reference Texture
by: Huang, Youze, et al.
Published: (2025) -
Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
by: He, Wei, et al.
Published: (2026)