HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Linhui, Yang, Xiaoshan, Peng, Fang, Wang, Yaowei, Xu, Changsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
Towards Visual Grounding: A Survey
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
Pilot: Building the Federated Multimodal Instruction Tuning Framework
von: Xiong, Baochen, et al.
Veröffentlicht: (2025)
von: Xiong, Baochen, et al.
Veröffentlicht: (2025)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
von: Lu, Zhiguang, et al.
Veröffentlicht: (2025)
von: Lu, Zhiguang, et al.
Veröffentlicht: (2025)
HiProto: Hierarchical Prototype Learning for Interpretable Object Detection Under Low-quality Conditions
von: Xiang, Jianlin, et al.
Veröffentlicht: (2026)
von: Xiang, Jianlin, et al.
Veröffentlicht: (2026)
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
PathVG: A New Benchmark and Dataset for Pathology Visual Grounding
von: Zhong, Chunlin, et al.
Veröffentlicht: (2025)
von: Zhong, Chunlin, et al.
Veröffentlicht: (2025)
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
von: Dai, Ming, et al.
Veröffentlicht: (2024)
von: Dai, Ming, et al.
Veröffentlicht: (2024)
$\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer
von: Zhao, Yibin, et al.
Veröffentlicht: (2026)
von: Zhao, Yibin, et al.
Veröffentlicht: (2026)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
von: Li, Haocheng, et al.
Veröffentlicht: (2026)
von: Li, Haocheng, et al.
Veröffentlicht: (2026)
Libra: Building Decoupled Vision System on Large Language Models
von: Xu, Yifan, et al.
Veröffentlicht: (2024)
von: Xu, Yifan, et al.
Veröffentlicht: (2024)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)
von: Bai, Sule, et al.
Veröffentlicht: (2025)
RGBT-Ground Benchmark: Visual Grounding Beyond RGB in Complex Real-World Scenarios
von: Zhao, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhao, Tianyi, et al.
Veröffentlicht: (2025)
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
von: Tao, Ming, et al.
Veröffentlicht: (2024)
von: Tao, Ming, et al.
Veröffentlicht: (2024)
VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
von: Yan, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Yan, Xiaoyang, et al.
Veröffentlicht: (2026)
LLM4VG: Large Language Models Evaluation for Video Grounding
von: Feng, Wei, et al.
Veröffentlicht: (2023)
von: Feng, Wei, et al.
Veröffentlicht: (2023)
A Comprehensive Review of Few-shot Action Recognition
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2024)
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2024)
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
von: Yang, Tianshuo, et al.
Veröffentlicht: (2026)
von: Yang, Tianshuo, et al.
Veröffentlicht: (2026)
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025)
von: Kim, Junho, et al.
Veröffentlicht: (2025)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
von: Kang, Weitai, et al.
Veröffentlicht: (2025)
von: Kang, Weitai, et al.
Veröffentlicht: (2025)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
TrajVG: 3D Trajectory-Coupled Visual Geometry Learning
von: Miao, Xingyu, et al.
Veröffentlicht: (2026)
von: Miao, Xingyu, et al.
Veröffentlicht: (2026)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
von: Liu, Junli, et al.
Veröffentlicht: (2025)
von: Liu, Junli, et al.
Veröffentlicht: (2025)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
von: Shao, Dian, et al.
Veröffentlicht: (2026)
von: Shao, Dian, et al.
Veröffentlicht: (2026)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
VG-SSL: Benchmarking Self-supervised Representation Learning Approaches for Visual Geo-localization
von: Xiao, Jiuhong, et al.
Veröffentlicht: (2023)
von: Xiao, Jiuhong, et al.
Veröffentlicht: (2023)
BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
von: Li, Hongbing, et al.
Veröffentlicht: (2026)
von: Li, Hongbing, et al.
Veröffentlicht: (2026)
UniVG: Towards UNIfied-modal Video Generation
von: Ruan, Ludan, et al.
Veröffentlicht: (2024)
von: Ruan, Ludan, et al.
Veröffentlicht: (2024)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
von: Tao, Ming, et al.
Veröffentlicht: (2024)
von: Tao, Ming, et al.
Veröffentlicht: (2024)
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
von: Xu, Xiaoran, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoran, et al.
Veröffentlicht: (2026)
Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition
von: Zuo, Yukun, et al.
Veröffentlicht: (2024)
von: Zuo, Yukun, et al.
Veröffentlicht: (2024)
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2025)
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2023) -
Towards Visual Grounding: A Survey
von: Xiao, Linhui, et al.
Veröffentlicht: (2024) -
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
von: Xiao, Linhui, et al.
Veröffentlicht: (2024) -
Pilot: Building the Federated Multimodal Instruction Tuning Framework
von: Xiong, Baochen, et al.
Veröffentlicht: (2025) -
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
von: Shi, Liangtao, et al.
Veröffentlicht: (2025)