Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jingchao, Wu, Zhijian, Huang, Dingjiang, Zheng, Yefeng, Wang, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
VPTracker: Global Vision-Language Tracking via Visual Prompt
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Learning Spectral Diffusion Prior for Hyperspectral Image Reconstruction
by: Yu, Mingyang, et al.
Published: (2025)
by: Yu, Mingyang, et al.
Published: (2025)
A Simple and Better Baseline for Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
by: Zuo, Rui, et al.
Published: (2025)
by: Zuo, Rui, et al.
Published: (2025)
Refer to Any Segmentation Mask Group With Vision-Language Prompts
by: Cao, Shengcao, et al.
Published: (2025)
by: Cao, Shengcao, et al.
Published: (2025)
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
A Refreshed Similarity-based Upsampler for Direct High-Ratio Feature Upsampling
by: Zhou, Minghao, et al.
Published: (2024)
by: Zhou, Minghao, et al.
Published: (2024)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
by: Shen, Li-Cheng, et al.
Published: (2025)
by: Shen, Li-Cheng, et al.
Published: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
by: Chen, Tongfei, et al.
Published: (2026)
by: Chen, Tongfei, et al.
Published: (2026)
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
by: Chen, Yanhui, et al.
Published: (2026)
by: Chen, Yanhui, et al.
Published: (2026)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
X-SAM: From Segment Anything to Any Segmentation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
by: Zhu, Armando, et al.
Published: (2024)
by: Zhu, Armando, et al.
Published: (2024)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization
by: Chen, Yuqi, et al.
Published: (2026)
by: Chen, Yuqi, et al.
Published: (2026)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
Single Domain Generalization for Multimodal Cross-Cancer Prognosis via Dirac Rebalancer and Distribution Entanglement
by: Jiang, Jia-Xuan, et al.
Published: (2025)
by: Jiang, Jia-Xuan, et al.
Published: (2025)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
by: Yao, Yuan, et al.
Published: (2026)
by: Yao, Yuan, et al.
Published: (2026)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
by: Kim, Dong-Hee, et al.
Published: (2025)
by: Kim, Dong-Hee, et al.
Published: (2025)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
MatchSeg: Towards Better Segmentation via Reference Image Matching
by: Huo, Jiayu, et al.
Published: (2024)
by: Huo, Jiayu, et al.
Published: (2024)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025)
by: Liu, Xianjie, et al.
Published: (2025)
ConformalSAM: Unlocking the Potential of Foundational Segmentation Models in Semi-Supervised Semantic Segmentation with Conformal Prediction
by: Chen, Danhui, et al.
Published: (2025)
by: Chen, Danhui, et al.
Published: (2025)
TopoTTA: Topology-Enhanced Test-Time Adaptation for Tubular Structure Segmentation
by: Zhou, Jiale, et al.
Published: (2025)
by: Zhou, Jiale, et al.
Published: (2025)
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
by: Wei, Meng, et al.
Published: (2026)
by: Wei, Meng, et al.
Published: (2026)
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
by: Jiang, Lifan, et al.
Published: (2026)
by: Jiang, Lifan, et al.
Published: (2026)
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
by: Jiang, Yankai, et al.
Published: (2024)
by: Jiang, Yankai, et al.
Published: (2024)
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025)
by: Goto, Kanoko, et al.
Published: (2025)
Similar Items
-
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025) -
VPTracker: Global Vision-Language Tracking via Visual Prompt
by: Wang, Jingchao, et al.
Published: (2025) -
Learning Spectral Diffusion Prior for Hyperspectral Image Reconstruction
by: Yu, Mingyang, et al.
Published: (2025) -
A Simple and Better Baseline for Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025) -
CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024)