ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Bencheng, Wang, Xinggang, Zhu, Lianghui, Zhang, Qian, Huang, Chang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
by: Zuo, Lin, et al.
Published: (2024)
by: Zuo, Lin, et al.
Published: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
by: Mane, Atharv Mahesh, et al.
Published: (2025)
by: Mane, Atharv Mahesh, et al.
Published: (2025)
ViG-Bias: Visually Grounded Bias Discovery and Mitigation
by: Marani, Badr-Eddine, et al.
Published: (2024)
by: Marani, Badr-Eddine, et al.
Published: (2024)
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
by: Li, Zongming, et al.
Published: (2025)
by: Li, Zongming, et al.
Published: (2025)
WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
by: Wei, Guoyizhe, et al.
Published: (2025)
by: Wei, Guoyizhe, et al.
Published: (2025)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
by: Liao, Bencheng, et al.
Published: (2025)
by: Liao, Bencheng, et al.
Published: (2025)
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
by: Zhu, Lianghui, et al.
Published: (2025)
by: Zhu, Lianghui, et al.
Published: (2025)
H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
by: Li, Xueyang, et al.
Published: (2025)
by: Li, Xueyang, et al.
Published: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026)
by: Feng, Zhe, et al.
Published: (2026)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
by: Wu, Fengyi, et al.
Published: (2025)
by: Wu, Fengyi, et al.
Published: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025)
by: Li, Yingyue, et al.
Published: (2025)
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
by: Zhang, Kewei, et al.
Published: (2026)
by: Zhang, Kewei, et al.
Published: (2026)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
VSViG: Real-time Video-based Seizure Detection via Skeleton-based Spatiotemporal ViG
by: Xu, Yankun, et al.
Published: (2023)
by: Xu, Yankun, et al.
Published: (2023)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
CNN-ViT Fusion with Adaptive Attention Gate for Brain Tumor MRI Classification: A Hybrid Deep Learning Model
by: Hasnain, Syed Ibad, et al.
Published: (2026)
by: Hasnain, Syed Ibad, et al.
Published: (2026)
ViPO: Visual Preference Optimization at Scale
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
by: Picard, David, et al.
Published: (2026)
by: Picard, David, et al.
Published: (2026)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
by: Tian, Yufeng, et al.
Published: (2026)
by: Tian, Yufeng, et al.
Published: (2026)
Interpretable and Sparse Linear Attention with Decoupled Membership-Subspace Modeling via MCR2 Objective
by: Liu, Tianyuan, et al.
Published: (2026)
by: Liu, Tianyuan, et al.
Published: (2026)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
by: Wang, Hui, et al.
Published: (2026)
by: Wang, Hui, et al.
Published: (2026)
PEANO-ViT: Power-Efficient Approximations of Non-Linearities in Vision Transformers
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
by: Gedik, Hakan Emre, et al.
Published: (2025)
by: Gedik, Hakan Emre, et al.
Published: (2025)
Linear Attention Based Deep Nonlocal Means Filtering for Multiplicative Noise Removal
by: Siyao, Xiao, et al.
Published: (2024)
by: Siyao, Xiao, et al.
Published: (2024)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification
by: Xiang, Suncheng, et al.
Published: (2025)
by: Xiang, Suncheng, et al.
Published: (2025)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
by: Li, Kailing, et al.
Published: (2025)
by: Li, Kailing, et al.
Published: (2025)
ViRED: Prediction of Visual Relations in Engineering Drawings
by: Gu, Chao, et al.
Published: (2024)
by: Gu, Chao, et al.
Published: (2024)
Similar Items
-
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
by: Zhu, Lianghui, et al.
Published: (2024) -
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025) -
Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
by: Zuo, Lin, et al.
Published: (2024) -
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024) -
Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
by: Mane, Atharv Mahesh, et al.
Published: (2025)