Vim-F: Visual State Space Model Benefiting from Learning in the Frequency Domain
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Juntao, Liu, Shaogeng, Zhou, Jun, Bian, Kun, Zhou, You, Liu, Jianning, Zhang, Pei, Liu, Bingyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Separable Self-attention Inspired by the State Space Model for Computer Vision
von: Zhang, Juntao, et al.
Veröffentlicht: (2025)
von: Zhang, Juntao, et al.
Veröffentlicht: (2025)
BadVim: Unveiling Backdoor Threats in Visual State Space Model
von: Lee, Cheng-Yi, et al.
Veröffentlicht: (2024)
von: Lee, Cheng-Yi, et al.
Veröffentlicht: (2024)
DefMamba: Deformable Visual State Space Model
von: Liu, Leiye, et al.
Veröffentlicht: (2025)
von: Liu, Leiye, et al.
Veröffentlicht: (2025)
ECMamba: Consolidating Selective State Space Model with Retinex Guidance for Efficient Multiple Exposure Correction
von: Dong, Wei, et al.
Veröffentlicht: (2024)
von: Dong, Wei, et al.
Veröffentlicht: (2024)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
VMamba: Visual State Space Model
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
SpectralMamba-UNet: Frequency-Disentangled State Space Modeling for Texture-Structure Consistent Medical Image Segmentation
von: Zhang, Fuhao, et al.
Veröffentlicht: (2026)
von: Zhang, Fuhao, et al.
Veröffentlicht: (2026)
Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
von: Shen, Yunzhe, et al.
Veröffentlicht: (2025)
von: Shen, Yunzhe, et al.
Veröffentlicht: (2025)
CFTrack: Enhancing Lightweight Visual Tracking through Contrastive Learning and Feature Matching
von: Liang, Juntao, et al.
Veröffentlicht: (2025)
von: Liang, Juntao, et al.
Veröffentlicht: (2025)
Continual Learning in the Frequency Domain
von: Liu, Ruiqi, et al.
Veröffentlicht: (2024)
von: Liu, Ruiqi, et al.
Veröffentlicht: (2024)
VFGS-Net: Frequency-Guided State-Space Learning for Topology-Preserving Retinal Vessel Segmentation
von: Song, Ruiqi, et al.
Veröffentlicht: (2026)
von: Song, Ruiqi, et al.
Veröffentlicht: (2026)
LocalMamba: Visual State Space Model with Windowed Selective Scan
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
von: Hou, Wenjin, et al.
Veröffentlicht: (2024)
von: Hou, Wenjin, et al.
Veröffentlicht: (2024)
Selective Structured State Space for Multispectral-fused Small Target Detection
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking
von: Liu, Xinqi, et al.
Veröffentlicht: (2024)
von: Liu, Xinqi, et al.
Veröffentlicht: (2024)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning
von: Zhou, Fei, et al.
Veröffentlicht: (2024)
von: Zhou, Fei, et al.
Veröffentlicht: (2024)
Multimodal Instruction Tuning with Hybrid State Space Models
von: Zhou, Jianing, et al.
Veröffentlicht: (2024)
von: Zhou, Jianing, et al.
Veröffentlicht: (2024)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
von: Zhou, Guanglin, et al.
Veröffentlicht: (2024)
von: Zhou, Guanglin, et al.
Veröffentlicht: (2024)
CMViM: Contrastive Masked Vim Autoencoder for 3D Multi-modal Representation Learning for AD classification
von: Yang, Guangqian, et al.
Veröffentlicht: (2024)
von: Yang, Guangqian, et al.
Veröffentlicht: (2024)
AtrousMamaba: An Atrous-Window Scanning Visual State Space Model for Remote Sensing Change Detection
von: Wang, Tao, et al.
Veröffentlicht: (2025)
von: Wang, Tao, et al.
Veröffentlicht: (2025)
FrequencyCT: Frequency Domain Self-supervised Low-dose CT Denoising
von: Wei, Guoquan, et al.
Veröffentlicht: (2026)
von: Wei, Guoquan, et al.
Veröffentlicht: (2026)
DifAttack++: Query-Efficient Black-Box Adversarial Attack via Hierarchical Disentangled Feature Space in Cross-Domain
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
DGMamba: Domain Generalization via Generalized State Space Model
von: Long, Shaocong, et al.
Veröffentlicht: (2024)
von: Long, Shaocong, et al.
Veröffentlicht: (2024)
Visual Foundation Models Boost Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
von: Liu, Yanan, et al.
Veröffentlicht: (2025)
von: Liu, Yanan, et al.
Veröffentlicht: (2025)
DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection
von: Li, Haochen, et al.
Veröffentlicht: (2026)
von: Li, Haochen, et al.
Veröffentlicht: (2026)
Learning View-Dependent Splatting Kernels
von: Ding, Huakeng, et al.
Veröffentlicht: (2026)
von: Ding, Huakeng, et al.
Veröffentlicht: (2026)
SFDFusion: An Efficient Spatial-Frequency Domain Fusion Network for Infrared and Visible Image Fusion
von: Hu, Kun, et al.
Veröffentlicht: (2024)
von: Hu, Kun, et al.
Veröffentlicht: (2024)
Language as an Anchor: Preserving Relative Visual Geometry for Domain Incremental Learning
von: Geng, Shuyi, et al.
Veröffentlicht: (2025)
von: Geng, Shuyi, et al.
Veröffentlicht: (2025)
SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
von: Li, Jiajun, et al.
Veröffentlicht: (2025)
von: Li, Jiajun, et al.
Veröffentlicht: (2025)
SPJFNet: Self-Mining Prior-Guided Joint Frequency Enhancement for Ultra-Efficient Dark Image Restoration
von: Zhang, Tongshun, et al.
Veröffentlicht: (2025)
von: Zhang, Tongshun, et al.
Veröffentlicht: (2025)
DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
von: Feng, Hao, et al.
Veröffentlicht: (2023)
von: Feng, Hao, et al.
Veröffentlicht: (2023)
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model
von: Xie, Fei, et al.
Veröffentlicht: (2024)
von: Xie, Fei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Separable Self-attention Inspired by the State Space Model for Computer Vision
von: Zhang, Juntao, et al.
Veröffentlicht: (2025) -
BadVim: Unveiling Backdoor Threats in Visual State Space Model
von: Lee, Cheng-Yi, et al.
Veröffentlicht: (2024) -
DefMamba: Deformable Visual State Space Model
von: Liu, Leiye, et al.
Veröffentlicht: (2025) -
ECMamba: Consolidating Selective State Space Model with Retinex Guidance for Efficient Multiple Exposure Correction
von: Dong, Wei, et al.
Veröffentlicht: (2024) -
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)