GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Shurong, Zhu, Yousong, Zhao, Hongyin, Yang, Fan, Zhan, Yufei, Tang, Ming, Wang, Jinqiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
di: Zhan, Yufei, et al.
Pubblicazione: (2025)
di: Zhan, Yufei, et al.
Pubblicazione: (2025)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
di: Zhan, Yufei, et al.
Pubblicazione: (2024)
di: Zhan, Yufei, et al.
Pubblicazione: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
di: Zhan, Yufei, et al.
Pubblicazione: (2025)
di: Zhan, Yufei, et al.
Pubblicazione: (2025)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
di: Zhan, Yufei, et al.
Pubblicazione: (2024)
di: Zhan, Yufei, et al.
Pubblicazione: (2024)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
di: Yang, Fan, et al.
Pubblicazione: (2026)
di: Yang, Fan, et al.
Pubblicazione: (2026)
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
di: Yang, Fan, et al.
Pubblicazione: (2025)
di: Yang, Fan, et al.
Pubblicazione: (2025)
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
di: Zhan, Yufei, et al.
Pubblicazione: (2023)
di: Zhan, Yufei, et al.
Pubblicazione: (2023)
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
di: Yu, Jiachen, et al.
Pubblicazione: (2025)
di: Yu, Jiachen, et al.
Pubblicazione: (2025)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
di: Zhan, Yufei, et al.
Pubblicazione: (2025)
di: Zhan, Yufei, et al.
Pubblicazione: (2025)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
di: Kang, Weitai, et al.
Pubblicazione: (2025)
di: Kang, Weitai, et al.
Pubblicazione: (2025)
From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
di: Yang, Fan, et al.
Pubblicazione: (2025)
di: Yang, Fan, et al.
Pubblicazione: (2025)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
di: Li, Haocheng, et al.
Pubblicazione: (2026)
di: Li, Haocheng, et al.
Pubblicazione: (2026)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
di: Dai, Ming, et al.
Pubblicazione: (2025)
di: Dai, Ming, et al.
Pubblicazione: (2025)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
di: Shi, Liangtao, et al.
Pubblicazione: (2025)
di: Shi, Liangtao, et al.
Pubblicazione: (2025)
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
di: Xiao, Linhui, et al.
Pubblicazione: (2024)
di: Xiao, Linhui, et al.
Pubblicazione: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
di: Ilaslan, Muhammet Furkan, et al.
Pubblicazione: (2024)
di: Ilaslan, Muhammet Furkan, et al.
Pubblicazione: (2024)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
di: Dai, Ming, et al.
Pubblicazione: (2024)
di: Dai, Ming, et al.
Pubblicazione: (2024)
GeM-EA: A Generative and Meta-learning Enhanced Evolutionary Algorithm for Streaming Data-Driven Optimization
di: Wu, Yue, et al.
Pubblicazione: (2026)
di: Wu, Yue, et al.
Pubblicazione: (2026)
Towards Visual Text Grounding of Multimodal Large Language Model
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
LLM4VG: Large Language Models Evaluation for Video Grounding
di: Feng, Wei, et al.
Pubblicazione: (2023)
di: Feng, Wei, et al.
Pubblicazione: (2023)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
di: Xiao, Linhui, et al.
Pubblicazione: (2023)
di: Xiao, Linhui, et al.
Pubblicazione: (2023)
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
di: Li, Ke, et al.
Pubblicazione: (2026)
di: Li, Ke, et al.
Pubblicazione: (2026)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
Challenges and Responses in the Practice of Large Language Models
di: Zhu, Hongyin
Pubblicazione: (2024)
di: Zhu, Hongyin
Pubblicazione: (2024)
Architectural Foundations for the Large Language Model Infrastructures
di: Zhu, Hongyin
Pubblicazione: (2024)
di: Zhu, Hongyin
Pubblicazione: (2024)
PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
di: Zhong, Chunlin, et al.
Pubblicazione: (2025)
di: Zhong, Chunlin, et al.
Pubblicazione: (2025)
VG3T: Visual Geometry Grounded Gaussian Transformer
di: Kim, Junho, et al.
Pubblicazione: (2025)
di: Kim, Junho, et al.
Pubblicazione: (2025)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
di: Liu, Junli, et al.
Pubblicazione: (2025)
di: Liu, Junli, et al.
Pubblicazione: (2025)
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
di: Zheng, Minghang, et al.
Pubblicazione: (2024)
di: Zheng, Minghang, et al.
Pubblicazione: (2024)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
di: Fan, Rong, et al.
Pubblicazione: (2026)
di: Fan, Rong, et al.
Pubblicazione: (2026)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
di: Bai, Sule, et al.
Pubblicazione: (2025)
di: Bai, Sule, et al.
Pubblicazione: (2025)
Efficient Masked Autoencoders with Self-Consistency
di: Li, Zhaowen, et al.
Pubblicazione: (2023)
di: Li, Zhaowen, et al.
Pubblicazione: (2023)
Systematic Outliers in Large Language Models
di: An, Yongqi, et al.
Pubblicazione: (2025)
di: An, Yongqi, et al.
Pubblicazione: (2025)
AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
di: Gu, Zhaopeng, et al.
Pubblicazione: (2025)
di: Gu, Zhaopeng, et al.
Pubblicazione: (2025)
$\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer
di: Zhao, Yibin, et al.
Pubblicazione: (2026)
di: Zhao, Yibin, et al.
Pubblicazione: (2026)
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
di: Kang, Weitai, et al.
Pubblicazione: (2024)
di: Kang, Weitai, et al.
Pubblicazione: (2024)
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
di: Wang, Yuji, et al.
Pubblicazione: (2025)
di: Wang, Yuji, et al.
Pubblicazione: (2025)
Climate Change from Large Language Models
di: Zhu, Hongyin, et al.
Pubblicazione: (2023)
di: Zhu, Hongyin, et al.
Pubblicazione: (2023)
Uma análise multimodal à luz do modelo GeM: homepage do Ministério do Meio Ambiente do Brasil
di: Thiago Brazileiro Vilar Hermont
Pubblicazione: (2016)
di: Thiago Brazileiro Vilar Hermont
Pubblicazione: (2016)
Documenti analoghi
-
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
di: Zhan, Yufei, et al.
Pubblicazione: (2025) -
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
di: Zhan, Yufei, et al.
Pubblicazione: (2024) -
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
di: Zhan, Yufei, et al.
Pubblicazione: (2025) -
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
di: Zhan, Yufei, et al.
Pubblicazione: (2024) -
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
di: Yang, Fan, et al.
Pubblicazione: (2026)