Parameter Aware Mamba Model for Multi-task Dense Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Xinzhuo, Zhuge, Yunzhi, Gong, Sitong, Zhang, Lu, Zhang, Pingping, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
by: Zhuge, Yunzhi, et al.
Published: (2025)
by: Zhuge, Yunzhi, et al.
Published: (2025)
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
by: Feng, Xiyan, et al.
Published: (2026)
by: Feng, Xiyan, et al.
Published: (2026)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
by: Ye, Chengyang, et al.
Published: (2024)
by: Ye, Chengyang, et al.
Published: (2024)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification
by: Yu, Chenyang, et al.
Published: (2025)
by: Yu, Chenyang, et al.
Published: (2025)
Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
Learning Universal Features for Generalizable Image Forgery Localization
by: Zhao, Hengrun, et al.
Published: (2025)
by: Zhao, Hengrun, et al.
Published: (2025)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
by: Gao, Shixuan, et al.
Published: (2024)
by: Gao, Shixuan, et al.
Published: (2024)
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
by: Zhu, Yixin, et al.
Published: (2026)
by: Zhu, Yixin, et al.
Published: (2026)
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification
by: Zhang, Pingping, et al.
Published: (2025)
by: Zhang, Pingping, et al.
Published: (2025)
StableIdentity: Inserting Anybody into Anywhere at First Sight
by: Wang, Qinghe, et al.
Published: (2024)
by: Wang, Qinghe, et al.
Published: (2024)
DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting
by: Yang, Yicheng, et al.
Published: (2024)
by: Yang, Yicheng, et al.
Published: (2024)
ComPtr: Towards Diverse Bi-source Dense Prediction Tasks via A Simple yet General Complementary Transformer
by: Pang, Youwei, et al.
Published: (2023)
by: Pang, Youwei, et al.
Published: (2023)
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
by: Wang, Yingquan, et al.
Published: (2024)
by: Wang, Yingquan, et al.
Published: (2024)
Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
by: Liu, Qing'an, et al.
Published: (2026)
by: Liu, Qing'an, et al.
Published: (2026)
What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
by: Wang, Yingquan, et al.
Published: (2025)
by: Wang, Yingquan, et al.
Published: (2025)
SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
DefMamba: Deformable Visual State Space Model
by: Liu, Leiye, et al.
Published: (2025)
by: Liu, Leiye, et al.
Published: (2025)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
High-Performance Few-Shot Segmentation with Foundation Models: An Empirical Study
by: Chang, Shijie, et al.
Published: (2024)
by: Chang, Shijie, et al.
Published: (2024)
Unity is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Regularizing Subspace Redundancy of Low-Rank Adaptation
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
by: Li, Baolu, et al.
Published: (2025)
by: Li, Baolu, et al.
Published: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Multi-view Aggregation Network for Dichotomous Image Segmentation
by: Yu, Qian, et al.
Published: (2024)
by: Yu, Qian, et al.
Published: (2024)
MambaVT: Spatio-Temporal Contextual Modeling for robust RGB-T Tracking
by: Lai, Simiao, et al.
Published: (2024)
by: Lai, Simiao, et al.
Published: (2024)
MAS-SAM: Segment Any Marine Animal with Aggregated Features
by: Yan, Tianyu, et al.
Published: (2024)
by: Yan, Tianyu, et al.
Published: (2024)
Similar Items
-
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025) -
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025) -
Reinforcing Video Reasoning Segmentation to Think Before It Segments
by: Gong, Sitong, et al.
Published: (2025) -
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
by: Gong, Sitong, et al.
Published: (2025) -
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)