Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Shengcao, Gui, Liang-Yan, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
von: Kang, Weitai, et al.
Veröffentlicht: (2025)
von: Kang, Weitai, et al.
Veröffentlicht: (2025)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
von: Ma, Chuofan, et al.
Veröffentlicht: (2024)
von: Ma, Chuofan, et al.
Veröffentlicht: (2024)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
von: Agarwal, Sakshi, et al.
Veröffentlicht: (2026)
von: Agarwal, Sakshi, et al.
Veröffentlicht: (2026)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
von: Woo, Byeongju, et al.
Veröffentlicht: (2026)
von: Woo, Byeongju, et al.
Veröffentlicht: (2026)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
von: Feng, Chun, et al.
Veröffentlicht: (2024)
von: Feng, Chun, et al.
Veröffentlicht: (2024)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
von: Du, Zilin, et al.
Veröffentlicht: (2024)
von: Du, Zilin, et al.
Veröffentlicht: (2024)
Situational Awareness Matters in 3D Vision Language Reasoning
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
Why are Visually-Grounded Language Models Bad at Image Classification?
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2024)
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025)
von: Kim, Junho, et al.
Veröffentlicht: (2025)
Visual Test-time Scaling for GUI Agent Grounding
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
von: He, Haoyang, et al.
Veröffentlicht: (2025)
von: He, Haoyang, et al.
Veröffentlicht: (2025)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
von: Wang, Wenkai, et al.
Veröffentlicht: (2026)
von: Wang, Wenkai, et al.
Veröffentlicht: (2026)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2024)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2024)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2026)
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2026)
GroCo: Ground Constraint for Metric Self-Supervised Monocular Depth
von: Cecille, Aurélien, et al.
Veröffentlicht: (2024)
von: Cecille, Aurélien, et al.
Veröffentlicht: (2024)
Multimodal Reference Visual Grounding
von: Lu, Yangxiao, et al.
Veröffentlicht: (2025)
von: Lu, Yangxiao, et al.
Veröffentlicht: (2025)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
Cortex-Grounded Diffusion Models for Brain Image Generation
von: Bongratz, Fabian, et al.
Veröffentlicht: (2026)
von: Bongratz, Fabian, et al.
Veröffentlicht: (2026)
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
von: Lin, Juyi, et al.
Veröffentlicht: (2026)
von: Lin, Juyi, et al.
Veröffentlicht: (2026)
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
von: Yu, Tianjiao, et al.
Veröffentlicht: (2026)
von: Yu, Tianjiao, et al.
Veröffentlicht: (2026)
Grounded Object Centric Learning
von: Kori, Avinash, et al.
Veröffentlicht: (2023)
von: Kori, Avinash, et al.
Veröffentlicht: (2023)
Visual Generation Without Guidance
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
von: Li, Chenghao, et al.
Veröffentlicht: (2026)
von: Li, Chenghao, et al.
Veröffentlicht: (2026)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
von: Cao-Dinh, Duc, et al.
Veröffentlicht: (2025)
von: Cao-Dinh, Duc, et al.
Veröffentlicht: (2025)
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
von: Kenthapadi, Krishnaram, et al.
Veröffentlicht: (2024)
von: Kenthapadi, Krishnaram, et al.
Veröffentlicht: (2024)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
von: Miao, Yanting, et al.
Veröffentlicht: (2026)
von: Miao, Yanting, et al.
Veröffentlicht: (2026)
Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
von: Sun, Shenghuan, et al.
Veröffentlicht: (2024)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
von: Liu, Zhongye, et al.
Veröffentlicht: (2024)
von: Liu, Zhongye, et al.
Veröffentlicht: (2024)
DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2025)
von: Alvar, Saeed Ranjbar, et al.
Veröffentlicht: (2025)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
von: Cao, Shengcao, et al.
Veröffentlicht: (2024) -
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025) -
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
von: Kang, Weitai, et al.
Veröffentlicht: (2025) -
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
von: Ma, Chuofan, et al.
Veröffentlicht: (2024) -
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2024)