BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Chaodong, Zhang, Zhengqiang, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Polyline Path Masked Attention for Vision Transformer
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025)
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025)
Spatial-Mamba: Effective Visual State Space Models via Structure-aware State Fusion
von: Xiao, Chaodong, et al.
Veröffentlicht: (2024)
von: Xiao, Chaodong, et al.
Veröffentlicht: (2024)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
Vision Transformers with Hierarchical Attention
von: Liu, Yun, et al.
Veröffentlicht: (2021)
von: Liu, Yun, et al.
Veröffentlicht: (2021)
Spiking Vision Transformer with Saccadic Attention
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Representative Attention For Vision Transformers
von: Li, Yuntong, et al.
Veröffentlicht: (2026)
von: Li, Yuntong, et al.
Veröffentlicht: (2026)
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
von: Sun, Lingchen, et al.
Veröffentlicht: (2026)
von: Sun, Lingchen, et al.
Veröffentlicht: (2026)
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
von: Sun, Yujing, et al.
Veröffentlicht: (2025)
von: Sun, Yujing, et al.
Veröffentlicht: (2025)
HAViT: Historical Attention Vision Transformer
von: Banik, Swarnendu, et al.
Veröffentlicht: (2026)
von: Banik, Swarnendu, et al.
Veröffentlicht: (2026)
Structured Initialization for Attention in Vision Transformers
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2024)
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2024)
Vision Transformers are Circulant Attention Learners
von: Han, Dongchen, et al.
Veröffentlicht: (2025)
von: Han, Dongchen, et al.
Veröffentlicht: (2025)
Multi-manifold Attention for Vision Transformers
von: Konstantinidis, Dimitrios, et al.
Veröffentlicht: (2022)
von: Konstantinidis, Dimitrios, et al.
Veröffentlicht: (2022)
FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
von: Lee, Sanghyeok, et al.
Veröffentlicht: (2024)
von: Lee, Sanghyeok, et al.
Veröffentlicht: (2024)
Attention Retention for Continual Learning with Vision Transformers
von: Lu, Yue, et al.
Veröffentlicht: (2026)
von: Lu, Yue, et al.
Veröffentlicht: (2026)
One Attention, One Scale: Phase-Aligned Rotary Positional Embeddings for Mixed-Resolution Diffusion Transformer
von: Wu, Haoyu, et al.
Veröffentlicht: (2025)
von: Wu, Haoyu, et al.
Veröffentlicht: (2025)
Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
von: Gao, Xinle, et al.
Veröffentlicht: (2025)
von: Gao, Xinle, et al.
Veröffentlicht: (2025)
ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention
von: He, Chenhang, et al.
Veröffentlicht: (2024)
von: He, Chenhang, et al.
Veröffentlicht: (2024)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
Analysis of Attention in Video Diffusion Transformers
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
von: Ding, Hangliang, et al.
Veröffentlicht: (2025)
von: Ding, Hangliang, et al.
Veröffentlicht: (2025)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
von: Liang, Wenhao, et al.
Veröffentlicht: (2025)
von: Liang, Wenhao, et al.
Veröffentlicht: (2025)
You Only Need Less Attention at Each Stage in Vision Transformers
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
One Last Attention for Your Vision-Language Model
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution
von: Yi, Qiaosi, et al.
Veröffentlicht: (2026)
von: Yi, Qiaosi, et al.
Veröffentlicht: (2026)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
von: You, Haoran, et al.
Veröffentlicht: (2022)
von: You, Haoran, et al.
Veröffentlicht: (2022)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
Decision-Aware Attention Propagation for Vision Transformer Explainability
von: Jo, Sehyeong, et al.
Veröffentlicht: (2026)
von: Jo, Sehyeong, et al.
Veröffentlicht: (2026)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024)
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024)
Elastic Attention Cores for Scalable Vision Transformers
von: Song, Alan Z., et al.
Veröffentlicht: (2026)
von: Song, Alan Z., et al.
Veröffentlicht: (2026)
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
von: Shmilovich, Dor, et al.
Veröffentlicht: (2025)
von: Shmilovich, Dor, et al.
Veröffentlicht: (2025)
Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers
von: Zhang, Chunyang, et al.
Veröffentlicht: (2025)
von: Zhang, Chunyang, et al.
Veröffentlicht: (2025)
Attention Sinks in Diffusion Transformers: A Causal Analysis
von: Wu, Fangzheng, et al.
Veröffentlicht: (2026)
von: Wu, Fangzheng, et al.
Veröffentlicht: (2026)
SSL: A Self-similarity Loss for Improving Generative Image Super-resolution
von: Chen, Du, et al.
Veröffentlicht: (2024)
von: Chen, Du, et al.
Veröffentlicht: (2024)
GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2025)
The Linear Attention Resurrection in Vision Transformer
von: Zheng, Chuanyang
Veröffentlicht: (2025)
von: Zheng, Chuanyang
Veröffentlicht: (2025)
VOSR: A Vision-Only Generative Model for Image Super-Resolution
von: Wu, Rongyuan, et al.
Veröffentlicht: (2026)
von: Wu, Rongyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Polyline Path Masked Attention for Vision Transformer
von: Zhao, Zhongchen, et al.
Veröffentlicht: (2025) -
Spatial-Mamba: Effective Visual State Space Models via Structure-aware State Fusion
von: Xiao, Chaodong, et al.
Veröffentlicht: (2024) -
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
von: Luo, Jiayi, et al.
Veröffentlicht: (2026) -
Vision Transformers with Hierarchical Attention
von: Liu, Yun, et al.
Veröffentlicht: (2021) -
Spiking Vision Transformer with Saccadic Attention
von: Wang, Shuai, et al.
Veröffentlicht: (2025)