New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Xuzheng, Liu, Junzhuo, Wang, Peng, Wang, Guoqing, Yang, Yang, Shen, Heng Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024)
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024)
GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
von: Liu, Bing, et al.
Veröffentlicht: (2025)
von: Liu, Bing, et al.
Veröffentlicht: (2025)
Region-aware Distribution Contrast: A Novel Approach to Multi-Task Partially Supervised Learning
von: Li, Meixuan, et al.
Veröffentlicht: (2024)
von: Li, Meixuan, et al.
Veröffentlicht: (2024)
Fine-Grained Representation for Lane Topology Reasoning
von: Xu, Guoqing, et al.
Veröffentlicht: (2025)
von: Xu, Guoqing, et al.
Veröffentlicht: (2025)
CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
Implicit Counterfactual Learning for Audio-Visual Segmentation
von: Zha, Mingfeng, et al.
Veröffentlicht: (2025)
von: Zha, Mingfeng, et al.
Veröffentlicht: (2025)
WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
von: Zhang, Shilan, et al.
Veröffentlicht: (2025)
von: Zhang, Shilan, et al.
Veröffentlicht: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
von: Qiu, Shulei, et al.
Veröffentlicht: (2024)
von: Qiu, Shulei, et al.
Veröffentlicht: (2024)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
von: Lei, Yinjie, et al.
Veröffentlicht: (2023)
von: Lei, Yinjie, et al.
Veröffentlicht: (2023)
JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image Enhancement
von: Wu, Yuhui, et al.
Veröffentlicht: (2023)
von: Wu, Yuhui, et al.
Veröffentlicht: (2023)
EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery
von: Wang, Gui, et al.
Veröffentlicht: (2025)
von: Wang, Gui, et al.
Veröffentlicht: (2025)
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
von: Zheng, Lihao, et al.
Veröffentlicht: (2026)
von: Zheng, Lihao, et al.
Veröffentlicht: (2026)
FineXtrol: Controllable Motion Generation via Fine-Grained Text
von: Shen, Keming, et al.
Veröffentlicht: (2025)
von: Shen, Keming, et al.
Veröffentlicht: (2025)
UAV as Urban Construction Change Monitor: A New Benchmark and Change Captioning Model
von: Gao, Yupeng, et al.
Veröffentlicht: (2026)
von: Gao, Yupeng, et al.
Veröffentlicht: (2026)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
von: Ji, Fengxian, et al.
Veröffentlicht: (2025)
von: Ji, Fengxian, et al.
Veröffentlicht: (2025)
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
von: Tao, Chaofan, et al.
Veröffentlicht: (2024)
MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
von: Qiao, Yang, et al.
Veröffentlicht: (2025)
von: Qiao, Yang, et al.
Veröffentlicht: (2025)
Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
von: Tang, Zhenchen, et al.
Veröffentlicht: (2025)
von: Tang, Zhenchen, et al.
Veröffentlicht: (2025)
Hierarchical Consistency Learning for Test-time Adaptation in Camouflage Perception
von: Zha, Mingfeng, et al.
Veröffentlicht: (2026)
von: Zha, Mingfeng, et al.
Veröffentlicht: (2026)
GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
von: Dou, Weijia, et al.
Veröffentlicht: (2025)
von: Dou, Weijia, et al.
Veröffentlicht: (2025)
Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
von: Yang, Yiting, et al.
Veröffentlicht: (2025)
von: Yang, Yiting, et al.
Veröffentlicht: (2025)
Micro-Expression Recognition via Fine-Grained Dynamic Perception
von: Shao, Zhiwen, et al.
Veröffentlicht: (2025)
von: Shao, Zhiwen, et al.
Veröffentlicht: (2025)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
von: Zhong, Liangyu, et al.
Veröffentlicht: (2025)
von: Zhong, Liangyu, et al.
Veröffentlicht: (2025)
Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
von: Huang, Longfei, et al.
Veröffentlicht: (2024)
von: Huang, Longfei, et al.
Veröffentlicht: (2024)
Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
von: Yang, Xiaobo, et al.
Veröffentlicht: (2025)
von: Yang, Xiaobo, et al.
Veröffentlicht: (2025)
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
von: Han, Mingfei, et al.
Veröffentlicht: (2023)
von: Han, Mingfei, et al.
Veröffentlicht: (2023)
Balancing Multi-Target Semi-Supervised Medical Image Segmentation with Collaborative Generalist and Specialists
von: Wang, You, et al.
Veröffentlicht: (2025)
von: Wang, You, et al.
Veröffentlicht: (2025)
Fast SAM2 with Text-Driven Token Pruning
von: Mandal, Avilasha, et al.
Veröffentlicht: (2025)
von: Mandal, Avilasha, et al.
Veröffentlicht: (2025)
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
von: Chen, Jierun, et al.
Veröffentlicht: (2024)
von: Chen, Jierun, et al.
Veröffentlicht: (2024)
Hierarchical Feature Learning for Medical Point Clouds via State Space Model
von: Zhang, Guoqing, et al.
Veröffentlicht: (2025)
von: Zhang, Guoqing, et al.
Veröffentlicht: (2025)
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
von: Yu, Enhui, et al.
Veröffentlicht: (2026)
von: Yu, Enhui, et al.
Veröffentlicht: (2026)
Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
von: Shi, Yuheng, et al.
Veröffentlicht: (2025)
von: Shi, Yuheng, et al.
Veröffentlicht: (2025)
Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning
von: Yan, Xudong, et al.
Veröffentlicht: (2024)
von: Yan, Xudong, et al.
Veröffentlicht: (2024)
Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation
von: Lei, Sen, et al.
Veröffentlicht: (2024)
von: Lei, Sen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024) -
GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
von: Liu, Bing, et al.
Veröffentlicht: (2025) -
Region-aware Distribution Contrast: A Novel Approach to Multi-Task Partially Supervised Learning
von: Li, Meixuan, et al.
Veröffentlicht: (2024) -
Fine-Grained Representation for Lane Topology Reasoning
von: Xu, Guoqing, et al.
Veröffentlicht: (2025) -
CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection
von: Li, Yiheng, et al.
Veröffentlicht: (2026)