Beyond Referring Expressions: Scenario Comprehension Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | He, Ruozhen, Shah, Nisarg A., Dong, Qihua, Xiao, Zilin, Koo, Jaywon, Ordonez, Vicente |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
by: Koo, Jaywon, et al.
Published: (2026)
by: Koo, Jaywon, et al.
Published: (2026)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
by: Xiao, Zilin, et al.
Published: (2025)
by: Xiao, Zilin, et al.
Published: (2025)
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024)
by: Koo, Jaywon, et al.
Published: (2024)
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024)
by: He, Ruozhen, et al.
Published: (2024)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)
by: Koo, Jaywon, et al.
Published: (2025)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
GViT: Representing Images as Gaussians for Visual Recognition
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
NoiseShift: Resolution-Aware Noise Recalibration for Better Low-Resolution Image Generation
by: He, Ruozhen, et al.
Published: (2025)
by: He, Ruozhen, et al.
Published: (2025)
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
by: Wang, Yaxian, et al.
Published: (2025)
by: Wang, Yaxian, et al.
Published: (2025)
RGBT-Ground Benchmark: Visual Grounding Beyond RGB in Complex Real-World Scenarios
by: Zhao, Tianyi, et al.
Published: (2025)
by: Zhao, Tianyi, et al.
Published: (2025)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
by: Willemsen, Bram, et al.
Published: (2024)
by: Willemsen, Bram, et al.
Published: (2024)
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
by: Li, Ling, et al.
Published: (2026)
by: Li, Ling, et al.
Published: (2026)
Comprehension of Multilingual Expressions Referring to Target Objects in Visual Inputs
by: Nogueira, Francisco, et al.
Published: (2025)
by: Nogueira, Francisco, et al.
Published: (2025)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Fairness and Bias Mitigation in Computer Vision: A Survey
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
by: Ding, Henghui, et al.
Published: (2026)
by: Ding, Henghui, et al.
Published: (2026)
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2026)
by: Wu, Changli, et al.
Published: (2026)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
ScanFormer: Referring Expression Comprehension by Iteratively Scanning
by: Su, Wei, et al.
Published: (2024)
by: Su, Wei, et al.
Published: (2024)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
by: Guo, Hao, et al.
Published: (2025)
by: Guo, Hao, et al.
Published: (2025)
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025)
by: Goto, Kanoko, et al.
Published: (2025)
Generative Visual Instruction Tuning
by: Hernandez, Jefferson, et al.
Published: (2024)
by: Hernandez, Jefferson, et al.
Published: (2024)
WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
by: Gao, Tianyi, et al.
Published: (2025)
by: Gao, Tianyi, et al.
Published: (2025)
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
by: Chen, Jierun, et al.
Published: (2024)
by: Chen, Jierun, et al.
Published: (2024)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Multimodal Reference Visual Grounding
by: Lu, Yangxiao, et al.
Published: (2025)
by: Lu, Yangxiao, et al.
Published: (2025)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
by: Sun, Zhichao, et al.
Published: (2025)
by: Sun, Zhichao, et al.
Published: (2025)
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
by: Koo, Juil, et al.
Published: (2025)
by: Koo, Juil, et al.
Published: (2025)
LOCORE: Image Re-ranking with Long-Context Sequence Modeling
by: Xiao, Zilin, et al.
Published: (2025)
by: Xiao, Zilin, et al.
Published: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
StepAL: Step-aware Active Learning for Cataract Surgical Videos
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
Exploring Spatial Language Grounding Through Referring Expressions
by: Tumu, Akshar, et al.
Published: (2025)
by: Tumu, Akshar, et al.
Published: (2025)
Similar Items
-
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
by: Koo, Jaywon, et al.
Published: (2026) -
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
by: Xiao, Zilin, et al.
Published: (2025) -
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024) -
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024) -
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)