Oscillation-Reduced MXFP4 Training for Vision Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yuxiang, Xi, Haocheng, Zhu, Jun, Chen, Jianfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024)
FrameBridge: Improving Image-to-Video Generation with Bridge Models
von: Wang, Yuji, et al.
Veröffentlicht: (2024)
von: Wang, Yuji, et al.
Veröffentlicht: (2024)
Visual Generation Without Guidance
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2025)
When Training-Free NAS Meets Vision Transformer: A Neural Tangent Kernel Perspective
von: Zhou, Qiqi, et al.
Veröffentlicht: (2024)
von: Zhou, Qiqi, et al.
Veröffentlicht: (2024)
The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design
von: Liu, Anjie, et al.
Veröffentlicht: (2026)
von: Liu, Anjie, et al.
Veröffentlicht: (2026)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
von: Du, Zilin, et al.
Veröffentlicht: (2024)
von: Du, Zilin, et al.
Veröffentlicht: (2024)
PointSSC: A Cooperative Vehicle-Infrastructure Point Cloud Benchmark for Semantic Scene Completion
von: Yan, Yuxiang, et al.
Veröffentlicht: (2023)
von: Yan, Yuxiang, et al.
Veröffentlicht: (2023)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets
von: Rangwani, Harsh, et al.
Veröffentlicht: (2024)
von: Rangwani, Harsh, et al.
Veröffentlicht: (2024)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Fairness-aware Vision Transformer via Debiased Self-Attention
von: Qiang, Yao, et al.
Veröffentlicht: (2023)
von: Qiang, Yao, et al.
Veröffentlicht: (2023)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
HGTDP-DTA: Hybrid Graph-Transformer with Dynamic Prompt for Drug-Target Binding Affinity Prediction
von: Xiao, Xi, et al.
Veröffentlicht: (2024)
von: Xiao, Xi, et al.
Veröffentlicht: (2024)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
Block-Recurrent Dynamics in Vision Transformers
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
Improving Interpretation Faithfulness for Vision Transformers
von: Hu, Lijie, et al.
Veröffentlicht: (2023)
von: Hu, Lijie, et al.
Veröffentlicht: (2023)
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
On the Surprising Effectiveness of Attention Transfer for Vision Transformers
von: Li, Alexander C., et al.
Veröffentlicht: (2024)
von: Li, Alexander C., et al.
Veröffentlicht: (2024)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
von: Mou, Linzhan, et al.
Veröffentlicht: (2024)
von: Mou, Linzhan, et al.
Veröffentlicht: (2024)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
Discovering Influential Neuron Path in Vision Transformers
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Accelerating Vision Transformers with Adaptive Patch Sizes
von: Choudhury, Rohan, et al.
Veröffentlicht: (2025)
von: Choudhury, Rohan, et al.
Veröffentlicht: (2025)
Continual Adaptation of Vision Transformers for Federated Learning
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023)
DiffiT: Diffusion Vision Transformers for Image Generation
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
ADAPT to Robustify Prompt Tuning Vision Transformers
von: Eskandar, Masih, et al.
Veröffentlicht: (2024)
von: Eskandar, Masih, et al.
Veröffentlicht: (2024)
Class-Discriminative Attention Maps for Vision Transformers
von: Brocki, Lennart, et al.
Veröffentlicht: (2023)
von: Brocki, Lennart, et al.
Veröffentlicht: (2023)
Fast Training of Diffusion Models with Masked Transformers
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
A Survey on Graph Neural Networks and Graph Transformers in Computer Vision: A Task-Oriented Perspective
von: Chen, Chaoqi, et al.
Veröffentlicht: (2022)
von: Chen, Chaoqi, et al.
Veröffentlicht: (2022)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Enhancing Vision Transformer Explainability Using Artificial Astrocytes
von: Echevarrieta-Catalan, Nicolas, et al.
Veröffentlicht: (2025)
von: Echevarrieta-Catalan, Nicolas, et al.
Veröffentlicht: (2025)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
FasterViT: Fast Vision Transformers with Hierarchical Attention
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
von: Xi, Yuanhao, et al.
Veröffentlicht: (2025)
von: Xi, Yuanhao, et al.
Veröffentlicht: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
von: Kim, Bum Jun, et al.
Veröffentlicht: (2024) -
FrameBridge: Improving Image-to-Video Generation with Bridge Models
von: Wang, Yuji, et al.
Veröffentlicht: (2024) -
Visual Generation Without Guidance
von: Chen, Huayu, et al.
Veröffentlicht: (2025)