Multimodal Reference Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yangxiao, Li, Ruosen, Jing, Liqiang, Wang, Jikai, Du, Xinya, Guo, Yunhui, Ruozzi, Nicholas, Xiang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
by: Lu, Yangxiao, et al.
Published: (2024)
by: Lu, Yangxiao, et al.
Published: (2024)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
by: Jing, Liqiang, et al.
Published: (2023)
by: Jing, Liqiang, et al.
Published: (2023)
Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
by: Mao, Ruiyu, et al.
Published: (2026)
by: Mao, Ruiyu, et al.
Published: (2026)
Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning
by: Zhang, Qifan, et al.
Published: (2024)
by: Zhang, Qifan, et al.
Published: (2024)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
by: Jing, Liqiang, et al.
Published: (2024)
by: Jing, Liqiang, et al.
Published: (2024)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
by: Zhang, Qifan, et al.
Published: (2026)
by: Zhang, Qifan, et al.
Published: (2026)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
SCENEREPLICA: Benchmarking Real-World Robot Manipulation by Creating Replicable Scenes
by: Khargonkar, Ninad, et al.
Published: (2023)
by: Khargonkar, Ninad, et al.
Published: (2023)
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
by: Du, Zilin, et al.
Published: (2024)
by: Du, Zilin, et al.
Published: (2024)
Learning Semantic Proxies from Visual Prompts for Parameter-Efficient Fine-Tuning in Deep Metric Learning
by: Ren, Li, et al.
Published: (2024)
by: Ren, Li, et al.
Published: (2024)
Multimodal Machine Translation with Visual Scene Graph Pruning
by: Lu, Chenyu, et al.
Published: (2025)
by: Lu, Chenyu, et al.
Published: (2025)
TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding
by: Guo, Wenxuan, et al.
Published: (2025)
by: Guo, Wenxuan, et al.
Published: (2025)
H2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution Detection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Inconsistency-Based Data-Centric Active Open-Set Annotation
by: Mao, Ruiyu, et al.
Published: (2024)
by: Mao, Ruiyu, et al.
Published: (2024)
PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time Adaptation
by: Maharana, Sarthak Kumar, et al.
Published: (2024)
by: Maharana, Sarthak Kumar, et al.
Published: (2024)
Introducing Visual Perception Token into Multimodal Large Language Model
by: Yu, Runpeng, et al.
Published: (2025)
by: Yu, Runpeng, et al.
Published: (2025)
Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study
by: Torop, Max, et al.
Published: (2025)
by: Torop, Max, et al.
Published: (2025)
Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
Not Just Change the Labels, Learn the Features: Watermarking Deep Neural Networks with Multi-View Data
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
by: Liang, Zichen, et al.
Published: (2025)
by: Liang, Zichen, et al.
Published: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
by: Wang, Wenkai, et al.
Published: (2026)
by: Wang, Wenkai, et al.
Published: (2026)
A Visual-inertial Localization Algorithm using Opportunistic Visual Beacons and Dead-Reckoning for GNSS-Denied Large-scale Applications
by: Zhang, Liqiang, et al.
Published: (2024)
by: Zhang, Liqiang, et al.
Published: (2024)
RaggeDi: Diffusion-based State Estimation of Disordered Rags, Sheets, Towels and Blankets
by: Ye, Jikai, et al.
Published: (2024)
by: Ye, Jikai, et al.
Published: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
Visual Prompting in Multimodal Large Language Models: A Survey
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
On the Role of Discrete Tokenization in Visual Representation Learning
by: Du, Tianqi, et al.
Published: (2024)
by: Du, Tianqi, et al.
Published: (2024)
Neutral-Reference Prompting for Vision-Language Models
by: Tian, Senmao, et al.
Published: (2026)
by: Tian, Senmao, et al.
Published: (2026)
Composition-Grounded Data Synthesis for Visual Reasoning
by: Gu, Xinyi, et al.
Published: (2025)
by: Gu, Xinyi, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics
by: Cao, Junyi, et al.
Published: (2024)
by: Cao, Junyi, et al.
Published: (2024)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
Grounding Descriptions in Images informs Zero-Shot Visual Recognition
by: Halbe, Shaunak, et al.
Published: (2024)
by: Halbe, Shaunak, et al.
Published: (2024)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
by: Safwan, Itbaan, et al.
Published: (2025)
by: Safwan, Itbaan, et al.
Published: (2025)
FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
by: Yan, Bowen, et al.
Published: (2024)
by: Yan, Bowen, et al.
Published: (2024)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
Token Activation Map to Visually Explain Multimodal LLMs
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
Similar Items
-
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
by: Lu, Yangxiao, et al.
Published: (2024) -
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
by: Jing, Liqiang, et al.
Published: (2023) -
Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
by: Mao, Ruiyu, et al.
Published: (2026) -
Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning
by: Zhang, Qifan, et al.
Published: (2024) -
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
by: Jing, Liqiang, et al.
Published: (2024)