PolaFormer: Polarity-aware Linear Attention for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Weikang, Luo, Yadan, Li, Xin, Jiang, Dongmei, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026)
by: Feng, Zhe, et al.
Published: (2026)
ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025)
by: Li, Wenxi, et al.
Published: (2025)
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
GRAD-Former: Gated Robust Attention-based Differential Transformer for Change Detection
by: Ameta, Durgesh, et al.
Published: (2026)
by: Ameta, Durgesh, et al.
Published: (2026)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Attention Retention for Continual Learning with Vision Transformers
by: Lu, Yue, et al.
Published: (2026)
by: Lu, Yue, et al.
Published: (2026)
LaCViT: A Label-aware Contrastive Fine-tuning Framework for Vision Transformers
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
Transferable Adversarial Face Attack with Text Controlled Attribute
by: Li, Wenyun, et al.
Published: (2024)
by: Li, Wenyun, et al.
Published: (2024)
Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
by: Hu, Yuanwei, et al.
Published: (2026)
by: Hu, Yuanwei, et al.
Published: (2026)
MetaFormer Baselines for Vision
by: Yu, Weihao, et al.
Published: (2022)
by: Yu, Weihao, et al.
Published: (2022)
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
by: Tang, Xiaoya, et al.
Published: (2024)
by: Tang, Xiaoya, et al.
Published: (2024)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
by: Zheng, Naichuan, et al.
Published: (2026)
by: Zheng, Naichuan, et al.
Published: (2026)
PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
by: Aniraj, Ananthu, et al.
Published: (2024)
by: Aniraj, Ananthu, et al.
Published: (2024)
ParkFormer: A Transformer-Based Parking Policy with Goal Embedding and Pedestrian-Aware Control
by: Fu, Jun, et al.
Published: (2025)
by: Fu, Jun, et al.
Published: (2025)
Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments
by: Etchegaray, Djamahl, et al.
Published: (2024)
by: Etchegaray, Djamahl, et al.
Published: (2024)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
SentiFormer: Metadata Enhanced Transformer for Image Sentiment Analysis
by: Feng, Bin, et al.
Published: (2025)
by: Feng, Bin, et al.
Published: (2025)
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
by: Li, Wenyun, et al.
Published: (2026)
by: Li, Wenyun, et al.
Published: (2026)
iFormer: Integrating ConvNet and Transformer for Mobile Application
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
by: Leem, Saebom, et al.
Published: (2024)
by: Leem, Saebom, et al.
Published: (2024)
WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation
by: Shu, Zishan, et al.
Published: (2026)
by: Shu, Zishan, et al.
Published: (2026)
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
by: Zhang, Zhizhen, et al.
Published: (2025)
by: Zhang, Zhizhen, et al.
Published: (2025)
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
by: Liao, Yi, et al.
Published: (2025)
by: Liao, Yi, et al.
Published: (2025)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
by: Zou, Shihao, et al.
Published: (2025)
by: Zou, Shihao, et al.
Published: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
by: Zhang, Junyang, et al.
Published: (2024)
by: Zhang, Junyang, et al.
Published: (2024)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
by: Dong, Xinpeng, et al.
Published: (2026)
by: Dong, Xinpeng, et al.
Published: (2026)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
SERNet-Former: Semantic Segmentation by Efficient Residual Network with Attention-Boosting Gates and Attention-Fusion Networks
by: Erisen, Serdar
Published: (2024)
by: Erisen, Serdar
Published: (2024)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
by: Vani, Ankit, et al.
Published: (2024)
by: Vani, Ankit, et al.
Published: (2024)
Dynamic Pondering Sparsity-aware Mixture-of-Experts Transformer for Event Stream based Visual Object Tracking
by: Wang, Shiao, et al.
Published: (2026)
by: Wang, Shiao, et al.
Published: (2026)
FactoFormer: Factorized Hyperspectral Transformers with Self-Supervised Pretraining
by: Mohamed, Shaheer, et al.
Published: (2023)
by: Mohamed, Shaheer, et al.
Published: (2023)
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
by: Kang, Beoungwoo, et al.
Published: (2024)
by: Kang, Beoungwoo, et al.
Published: (2024)
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
by: Chen, Zhuoxiao, et al.
Published: (2024)
by: Chen, Zhuoxiao, et al.
Published: (2024)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
Similar Items
-
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025) -
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026) -
ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
by: Li, Yan, et al.
Published: (2025) -
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025) -
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)