Improving Generalized Visual Grounding with Instance-aware Joint Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Ming, Cheng, Wenxuan, Liu, Jiang-Jiang, Yang, Lingfeng, Feng, Zhenhua, Yang, Wankou, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
by: Dai, Ming, et al.
Published: (2024)
by: Dai, Ming, et al.
Published: (2024)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
by: Feng, Ze, et al.
Published: (2025)
by: Feng, Ze, et al.
Published: (2025)
MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
by: Feng, Ze, et al.
Published: (2025)
by: Feng, Ze, et al.
Published: (2025)
Drone Referring Localization: An Efficient Heterogeneous Spatial Feature Interaction Method For UAV Self-Localization
by: Dai, Ming, et al.
Published: (2022)
by: Dai, Ming, et al.
Published: (2022)
Improving underwater semantic segmentation with underwater image quality attention and muti-scale aggregation attention
by: Zuo, Xin, et al.
Published: (2025)
by: Zuo, Xin, et al.
Published: (2025)
Precise GPS-Denied UAV Self-Positioning via Context-Enhanced Cross-View Geo-Localization
by: Xu, Yuanze, et al.
Published: (2025)
by: Xu, Yuanze, et al.
Published: (2025)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
by: Wu, Ruiqi, et al.
Published: (2024)
by: Wu, Ruiqi, et al.
Published: (2024)
Object-level Geometric Structure Preserving for Natural Image Stitching
by: Cai, Wenxiao, et al.
Published: (2024)
by: Cai, Wenxiao, et al.
Published: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
GAInS: Gradient Anomaly-aware Biomedical Instance Segmentation
by: Liu, Runsheng, et al.
Published: (2024)
by: Liu, Runsheng, et al.
Published: (2024)
Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model
by: Yin, Junhui, et al.
Published: (2026)
by: Yin, Junhui, et al.
Published: (2026)
StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
by: Li, Wen, et al.
Published: (2024)
by: Li, Wen, et al.
Published: (2024)
Sharpness-aware Dynamic Anchor Selection for Generalized Category Discovery
by: Peng, Zhimao, et al.
Published: (2025)
by: Peng, Zhimao, et al.
Published: (2025)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Lighting-aware Unified Model for Instance Segmentation
by: Liu, Qisai, et al.
Published: (2026)
by: Liu, Qisai, et al.
Published: (2026)
LSSInst: Improving Geometric Modeling in LSS-Based BEV Perception with Instance Representation
by: Ma, Weijie, et al.
Published: (2024)
by: Ma, Weijie, et al.
Published: (2024)
Exploring Effective Factors for Improving Visual In-Context Learning
by: Sun, Yanpeng, et al.
Published: (2023)
by: Sun, Yanpeng, et al.
Published: (2023)
Visual Instance-aware Prompt Tuning
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Visual Grounding for Object-Level Generalization in Reinforcement Learning
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring
by: Sun, Rui, et al.
Published: (2026)
by: Sun, Rui, et al.
Published: (2026)
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
by: Cai, Chen, et al.
Published: (2025)
by: Cai, Chen, et al.
Published: (2025)
TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation
by: Zheng, Rongkun, et al.
Published: (2023)
by: Zheng, Rongkun, et al.
Published: (2023)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Guiding Visual Autoregressive Models through Spectrum Weakening
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
by: Li, Honglin, et al.
Published: (2024)
by: Li, Honglin, et al.
Published: (2024)
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
by: Zheng, Minghang, et al.
Published: (2024)
by: Zheng, Minghang, et al.
Published: (2024)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation
by: Ren, Biaoyu, et al.
Published: (2026)
by: Ren, Biaoyu, et al.
Published: (2026)
Add-SD: Rational Generation without Manual Reference
by: Yang, Lingfeng, et al.
Published: (2024)
by: Yang, Lingfeng, et al.
Published: (2024)
LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
by: Shi, Xianglong, et al.
Published: (2025)
by: Shi, Xianglong, et al.
Published: (2025)
EIMC: Efficient Instance-aware Multi-modal Collaborative Perception
by: Yang, Kang, et al.
Published: (2026)
by: Yang, Kang, et al.
Published: (2026)
Director: Instance-aware Gaussian Splatting for Dynamic Scene Modeling and Understanding
by: Jiang, Yuheng, et al.
Published: (2026)
by: Jiang, Yuheng, et al.
Published: (2026)
Low-Biased General Annotated Dataset Generation
by: Jiang, Dengyang, et al.
Published: (2024)
by: Jiang, Dengyang, et al.
Published: (2024)
Similar Items
-
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
by: Dai, Ming, et al.
Published: (2024) -
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
by: Dai, Ming, et al.
Published: (2025) -
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
by: Feng, Ze, et al.
Published: (2025) -
MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
by: Dai, Ming, et al.
Published: (2025) -
DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
by: Dai, Ming, et al.
Published: (2025)