MTMamba++: Enhancing Multi-Task Dense Scene Understanding via Mamba-Based Decoders
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Baijiong, Jiang, Weisen, Chen, Pengguang, Liu, Shu, Chen, Ying-Cong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
by: Lin, Baijiong, et al.
Published: (2024)
by: Lin, Baijiong, et al.
Published: (2024)
BYOM: Building Your Own Multi-Task Model For Free
by: Jiang, Weisen, et al.
Published: (2023)
by: Jiang, Weisen, et al.
Published: (2023)
Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
by: Cao, Mang, et al.
Published: (2025)
by: Cao, Mang, et al.
Published: (2025)
Dual-Balancing for Multi-Task Learning
by: Lin, Baijiong, et al.
Published: (2023)
by: Lin, Baijiong, et al.
Published: (2023)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
by: Wang, Xiaoye, et al.
Published: (2025)
by: Wang, Xiaoye, et al.
Published: (2025)
PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
by: Lin, Baijiong, et al.
Published: (2025)
by: Lin, Baijiong, et al.
Published: (2025)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
by: Li, Chenjun
Published: (2026)
by: Li, Chenjun
Published: (2026)
DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes
by: Chen, Jingru, et al.
Published: (2026)
by: Chen, Jingru, et al.
Published: (2026)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
by: Cheng, Dingxin, et al.
Published: (2024)
by: Cheng, Dingxin, et al.
Published: (2024)
Frequency-Dynamic Attention Modulation for Dense Prediction
by: Chen, Linwei, et al.
Published: (2025)
by: Chen, Linwei, et al.
Published: (2025)
TextMamba: Scene Text Detector with Mamba
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning
by: Mei, Xiaodong, et al.
Published: (2025)
by: Mei, Xiaodong, et al.
Published: (2025)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors
by: Zhang, Jingdong, et al.
Published: (2026)
by: Zhang, Jingdong, et al.
Published: (2026)
HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation
by: Chen, Zini, et al.
Published: (2026)
by: Chen, Zini, et al.
Published: (2026)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
by: Liu, Hanqing, et al.
Published: (2026)
by: Liu, Hanqing, et al.
Published: (2026)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
Frequency Dynamic Convolution for Dense Image Prediction
by: Chen, Linwei, et al.
Published: (2025)
by: Chen, Linwei, et al.
Published: (2025)
FSCM: Frequency-Enhanced Spatial-Spectral Coupled Mamba for Infrared Hyperspectral Image Colorization
by: Liu, Tingting, et al.
Published: (2026)
by: Liu, Tingting, et al.
Published: (2026)
HexPlane Representation for 3D Semantic Scene Understanding
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
GFE-Mamba: Mamba-based AD Multi-modal Progression Assessment via Generative Feature Extraction from MCI
by: Fang, Zhaojie, et al.
Published: (2024)
by: Fang, Zhaojie, et al.
Published: (2024)
Multi-Task Dense Prediction via Mixture of Low-Rank Experts
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
by: Zeng, Nianbo, et al.
Published: (2025)
by: Zeng, Nianbo, et al.
Published: (2025)
Enhancing Human-Centered Dynamic Scene Understanding via Multiple LLMs Collaborated Reasoning
by: Zhang, Hang, et al.
Published: (2024)
by: Zhang, Hang, et al.
Published: (2024)
GridMask Data Augmentation
by: Chen, Pengguang, et al.
Published: (2020)
by: Chen, Pengguang, et al.
Published: (2020)
TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation
by: Li, Jingyao, et al.
Published: (2023)
by: Li, Jingyao, et al.
Published: (2023)
MOODv2: Masked Image Modeling for Out-of-Distribution Detection
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
Multi-dimensional Visual Prompt Enhanced Image Restoration via Mamba-Transformer Aggregation
by: Jiang, Aiwen, et al.
Published: (2024)
by: Jiang, Aiwen, et al.
Published: (2024)
Hypergraph Mamba for Efficient Whole Slide Image Understanding
by: Lu, Jiaxuan, et al.
Published: (2025)
by: Lu, Jiaxuan, et al.
Published: (2025)
Frequency-aware Feature Fusion for Dense Image Prediction
by: Chen, Linwei, et al.
Published: (2024)
by: Chen, Linwei, et al.
Published: (2024)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
Optimizing Dense Visual Predictions Through Multi-Task Coherence and Prioritization
by: Fontana, Maxime, et al.
Published: (2024)
by: Fontana, Maxime, et al.
Published: (2024)
DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models
by: Shi, Xinrui, et al.
Published: (2026)
by: Shi, Xinrui, et al.
Published: (2026)
MROSS: Multi-Round Region-based Optimization for Scene Sketching
by: Liang, Yiqi, et al.
Published: (2024)
by: Liang, Yiqi, et al.
Published: (2024)
Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
by: Fu, Rao, et al.
Published: (2024)
by: Fu, Rao, et al.
Published: (2024)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
by: Nguyen, Vinh
Published: (2024)
by: Nguyen, Vinh
Published: (2024)
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
by: Gavrikov, Paul, et al.
Published: (2025)
by: Gavrikov, Paul, et al.
Published: (2025)
CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization
by: Liang, Yue, et al.
Published: (2026)
by: Liang, Yue, et al.
Published: (2026)
Similar Items
-
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
by: Lin, Baijiong, et al.
Published: (2024) -
BYOM: Building Your Own Multi-Task Model For Free
by: Jiang, Weisen, et al.
Published: (2023) -
Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
by: Cao, Mang, et al.
Published: (2025) -
Dual-Balancing for Multi-Task Learning
by: Lin, Baijiong, et al.
Published: (2023) -
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
by: Wang, Xiaoye, et al.
Published: (2025)