Scalable Visual State Space Model with Fractal Scanning
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Lv, Xiao, HaoKe, Jiang, Peng-Tao, Zhang, Hao, Chen, Jinwei, Li, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
by: Tang, Lv, et al.
Published: (2023)
by: Tang, Lv, et al.
Published: (2023)
Towards Training-free Open-world Segmentation via Image Prompt Foundation Models
by: Tang, Lv, et al.
Published: (2023)
by: Tang, Lv, et al.
Published: (2023)
FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
Improving Consistency in Diffusion Models for Image Super-Resolution
by: Gu, Junhao, et al.
Published: (2024)
by: Gu, Junhao, et al.
Published: (2024)
Empowering Segmentation Ability to Multi-modal Large Language Models
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
Multi-Task Dense Prediction via Mixture of Low-Rank Experts
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
by: Wan, Yuhao, et al.
Published: (2024)
by: Wan, Yuhao, et al.
Published: (2024)
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity
by: Yu, Qian, et al.
Published: (2024)
by: Yu, Qian, et al.
Published: (2024)
LocalMamba: Visual State Space Model with Windowed Selective Scan
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
VMambaCC: A Visual State Space Model for Crowd Counting
by: Ma, Hao-Yuan, et al.
Published: (2024)
by: Ma, Hao-Yuan, et al.
Published: (2024)
AtrousMamaba: An Atrous-Window Scanning Visual State Space Model for Remote Sensing Change Detection
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
SDMatte: Grafting Diffusion Models for Interactive Matting
by: Huang, Longfei, et al.
Published: (2025)
by: Huang, Longfei, et al.
Published: (2025)
Improving Adversarial Energy-Based Model via Diffusion Process
by: Geng, Cong, et al.
Published: (2024)
by: Geng, Cong, et al.
Published: (2024)
Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework
by: Shi, Linxiao, et al.
Published: (2026)
by: Shi, Linxiao, et al.
Published: (2026)
BadScan: An Architectural Backdoor Attack on Visual State Space Models
by: Deshmukh, Om Suhas, et al.
Published: (2024)
by: Deshmukh, Om Suhas, et al.
Published: (2024)
QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model
by: Xie, Fei, et al.
Published: (2024)
by: Xie, Fei, et al.
Published: (2024)
CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
by: Wang, Yuelei, et al.
Published: (2024)
by: Wang, Yuelei, et al.
Published: (2024)
DAMamba: Vision State Space Model with Dynamic Adaptive Scan
by: Li, Tanzhe, et al.
Published: (2025)
by: Li, Tanzhe, et al.
Published: (2025)
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
by: Wu, Chen, et al.
Published: (2026)
by: Wu, Chen, et al.
Published: (2026)
EDCSSM: Edge Detection with Convolutional State Space Model
by: Hong, Qinghui, et al.
Published: (2024)
by: Hong, Qinghui, et al.
Published: (2024)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
VMambaMorph: a Multi-Modality Deformable Image Registration Framework based on Visual State Space Model with Cross-Scan Module
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
ASAM: Boosting Segment Anything Model with Adversarial Tuning
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic Segmentation
by: Xia, Ruihao, et al.
Published: (2024)
by: Xia, Ruihao, et al.
Published: (2024)
Partial Ring Scan: Revisiting Scan Order in Vision State Space Models
by: Hsieh, Yi-Kuan, et al.
Published: (2026)
by: Hsieh, Yi-Kuan, et al.
Published: (2026)
MFil-Mamba: Multi-Filter Scanning for Spatial Redundancy-Aware Visual State Space Models
by: Khadka, Puskal, et al.
Published: (2026)
by: Khadka, Puskal, et al.
Published: (2026)
Learning Adaptive Lighting via Channel-Aware Guidance
by: Yang, Qirui, et al.
Published: (2024)
by: Yang, Qirui, et al.
Published: (2024)
Spatial-Mamba: Effective Visual State Space Models via Structure-aware State Fusion
by: Xiao, Chaodong, et al.
Published: (2024)
by: Xiao, Chaodong, et al.
Published: (2024)
Beyond Global Scanning: Adaptive Visual State Space Modeling for Salient Object Detection in Optical Remote Sensing Images
by: Ren, Mengyu, et al.
Published: (2025)
by: Ren, Mengyu, et al.
Published: (2025)
MambaVSR: Content-Aware Scanning State Space Model for Video Super-Resolution
by: He, Linfeng, et al.
Published: (2025)
by: He, Linfeng, et al.
Published: (2025)
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Mamba-Adaptor: State Space Model Adaptor for Visual Recognition
by: Xie, Fei, et al.
Published: (2025)
by: Xie, Fei, et al.
Published: (2025)
Dual Prototype-Conditioned Diffusion Model for Scalable Multi-Class Unsupervised Anomaly Detection in Large Category Spaces
by: Feng, Yaoxuan, et al.
Published: (2026)
by: Feng, Yaoxuan, et al.
Published: (2026)
Efficient Visual State Space Model for Image Deblurring
by: Kong, Lingshun, et al.
Published: (2024)
by: Kong, Lingshun, et al.
Published: (2024)
Selecting and Pruning: A Differentiable Causal Sequentialized State-Space Model for Two-View Correspondence Learning
by: Fang, Xiang, et al.
Published: (2025)
by: Fang, Xiang, et al.
Published: (2025)
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
by: Liu, Yanan, et al.
Published: (2025)
by: Liu, Yanan, et al.
Published: (2025)
Learning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
WaterMamba: Visual State Space Model for Underwater Image Enhancement
by: Guan, Meisheng, et al.
Published: (2024)
by: Guan, Meisheng, et al.
Published: (2024)
RED: Robust Event-Guided Motion Deblurring with Modality-Specific Disentanglement
by: Leng, Yihong, et al.
Published: (2025)
by: Leng, Yihong, et al.
Published: (2025)
Similar Items
-
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
by: Tang, Lv, et al.
Published: (2023) -
Towards Training-free Open-world Segmentation via Image Prompt Foundation Models
by: Tang, Lv, et al.
Published: (2023) -
FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry
by: Li, Bo, et al.
Published: (2025) -
Improving Consistency in Diffusion Models for Image Super-Resolution
by: Gu, Junhao, et al.
Published: (2024) -
Empowering Segmentation Ability to Multi-modal Large Language Models
by: Yang, Yuqi, et al.
Published: (2024)