Linear-Time Global Visual Modeling without Explicit Attention
Fuente:
arXiv
Saved in:
| Main Authors: | He, Ruize, Han, Dongchen, Huang, Gao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
Demystify Mamba in Vision: A Linear Attention Perspective
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023)
by: Xia, Zhuofan, et al.
Published: (2023)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking
by: Liang, Haiji, et al.
Published: (2024)
by: Liang, Haiji, et al.
Published: (2024)
From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera Calibration
by: Qian, Zekun, et al.
Published: (2022)
by: Qian, Zekun, et al.
Published: (2022)
Breaking the Low-Rank Dilemma of Linear Attention
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024)
by: Pu, Yifan, et al.
Published: (2024)
Rectifying Magnitude Neglect in Linear Attention
by: Fan, Qihang, et al.
Published: (2025)
by: Fan, Qihang, et al.
Published: (2025)
ViT$^3$: Unlocking Test-Time Training in Vision
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
by: Liu, Yuhe, et al.
Published: (2026)
by: Liu, Yuhe, et al.
Published: (2026)
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
A Unified Perspective on Adversarial Membership Manipulation in Vision Models
by: Gao, Ruize, et al.
Published: (2026)
by: Gao, Ruize, et al.
Published: (2026)
Constructing a High Temporal Resolution Global Lakes Dataset via Swin-Unet with Applications to Area Prediction
by: Han, Yutian, et al.
Published: (2024)
by: Han, Yutian, et al.
Published: (2024)
Step by Step Network
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
by: Qian, Zekun, et al.
Published: (2026)
by: Qian, Zekun, et al.
Published: (2026)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
Explicit Visual Prompts for Visual Object Tracking
by: Shi, Liangtao, et al.
Published: (2024)
by: Shi, Liangtao, et al.
Published: (2024)
OffsetOPT: Explicit Surface Reconstruction without Normals
by: Lei, Huan
Published: (2025)
by: Lei, Huan
Published: (2025)
Attention Distillation: A Unified Approach to Visual Characteristics Transfer
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
by: Han, Lawrence
Published: (2026)
by: Han, Lawrence
Published: (2026)
BEVANet: Bilateral Efficient Visual Attention Network for Real-Time Semantic Segmentation
by: Huang, Ping-Mao, et al.
Published: (2025)
by: Huang, Ping-Mao, et al.
Published: (2025)
Morpho-Aware Global Attention for Image Matting
by: Yang, Jingru, et al.
Published: (2024)
by: Yang, Jingru, et al.
Published: (2024)
AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Models
by: Wang, Teng, et al.
Published: (2026)
by: Wang, Teng, et al.
Published: (2026)
CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
by: Wang, Likai, et al.
Published: (2025)
by: Wang, Likai, et al.
Published: (2025)
From Indoor To Outdoor: Unsupervised Domain Adaptive Gait Recognition
by: Wang, Likai, et al.
Published: (2022)
by: Wang, Likai, et al.
Published: (2022)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
by: Peng, Qihang, et al.
Published: (2025)
by: Peng, Qihang, et al.
Published: (2025)
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention
by: Ai, Yuang, et al.
Published: (2025)
by: Ai, Yuang, et al.
Published: (2025)
DeepLatent: Think with Images via Parallel Latent Visual Reasoning
by: Lu, Dongchen, et al.
Published: (2026)
by: Lu, Dongchen, et al.
Published: (2026)
Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation
by: Zheng, Hongbo, et al.
Published: (2026)
by: Zheng, Hongbo, et al.
Published: (2026)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
by: Li, Kai, et al.
Published: (2025)
by: Li, Kai, et al.
Published: (2025)
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
by: Li, Junzhou, et al.
Published: (2026)
by: Li, Junzhou, et al.
Published: (2026)
Explicit Context Reasoning with Supervision for Visual Tracking
by: Zeng, Fansheng, et al.
Published: (2025)
by: Zeng, Fansheng, et al.
Published: (2025)
Linear Attention Modeling for Learned Image Compression
by: Feng, Donghui, et al.
Published: (2025)
by: Feng, Donghui, et al.
Published: (2025)
RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation
by: Xu, Guoan, et al.
Published: (2026)
by: Xu, Guoan, et al.
Published: (2026)
BiomechGPT: Towards a Biomechanically Fluent Multimodal Foundation Model for Clinically Relevant Motion Tasks
by: Yang, Ruize, et al.
Published: (2025)
by: Yang, Ruize, et al.
Published: (2025)
LoFLAT: Local Feature Matching using Focused Linear Attention Transformer
by: Cao, Naijian, et al.
Published: (2024)
by: Cao, Naijian, et al.
Published: (2024)
Similar Items
-
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023) -
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025) -
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024) -
Demystify Mamba in Vision: A Linear Attention Perspective
by: Han, Dongchen, et al.
Published: (2024) -
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023)