ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Minghang, Zhang, Jiahua, Chen, Qingchao, Peng, Yuxin, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
by: Zheng, Minghang, et al.
Published: (2024)
by: Zheng, Minghang, et al.
Published: (2024)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
by: Zheng, Peng, et al.
Published: (2024)
by: Zheng, Peng, et al.
Published: (2024)
Perceptual Flow Network for Visually Grounded Reasoning
by: Li, Yangfu, et al.
Published: (2026)
by: Li, Yangfu, et al.
Published: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
by: Deichler, Anna
Published: (2026)
by: Deichler, Anna
Published: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2025)
by: Zheng, Minghang, et al.
Published: (2025)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
by: Elsharkawi, Ismael, et al.
Published: (2026)
by: Elsharkawi, Ismael, et al.
Published: (2026)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2026)
by: Zheng, Minghang, et al.
Published: (2026)
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding
by: Wang, Chiyue, et al.
Published: (2026)
by: Wang, Chiyue, et al.
Published: (2026)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
by: Feng, Yigui, et al.
Published: (2026)
by: Feng, Yigui, et al.
Published: (2026)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
by: Han, Yudong, et al.
Published: (2026)
by: Han, Yudong, et al.
Published: (2026)
Pixel-Level Pavement Distress Assessment Using Instance Segmentation
by: Dewick, Logan, et al.
Published: (2026)
by: Dewick, Logan, et al.
Published: (2026)
Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal Retrieval
by: Liu, Yizhi, et al.
Published: (2025)
by: Liu, Yizhi, et al.
Published: (2025)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
by: Yang, Xiaoyu, et al.
Published: (2024)
by: Yang, Xiaoyu, et al.
Published: (2024)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers
by: Peng, Duo, et al.
Published: (2024)
by: Peng, Duo, et al.
Published: (2024)
A Simple Baseline for Streaming Video Understanding
by: Shen, Yujiao, et al.
Published: (2026)
by: Shen, Yujiao, et al.
Published: (2026)
RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
by: Ge, Junyao, et al.
Published: (2024)
by: Ge, Junyao, et al.
Published: (2024)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
by: Peng, Hongxing, et al.
Published: (2025)
by: Peng, Hongxing, et al.
Published: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
by: Oliveira, Daniel A. P., et al.
Published: (2024)
by: Oliveira, Daniel A. P., et al.
Published: (2024)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
by: Han, Yudong, et al.
Published: (2024)
by: Han, Yudong, et al.
Published: (2024)
NOCTIS: Novel Object Cyclic Threshold based Instance Segmentation
by: Gandyra, Max, et al.
Published: (2025)
by: Gandyra, Max, et al.
Published: (2025)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
by: Deng, Zhongchen, et al.
Published: (2024)
by: Deng, Zhongchen, et al.
Published: (2024)
AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback
by: Tupini, Andrea, et al.
Published: (2026)
by: Tupini, Andrea, et al.
Published: (2026)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
by: Khurdula, Harsha Vardhan, et al.
Published: (2024)
by: Khurdula, Harsha Vardhan, et al.
Published: (2024)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
by: Zhou, Fang, et al.
Published: (2025)
by: Zhou, Fang, et al.
Published: (2025)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
by: Panek, Vojtech, et al.
Published: (2024)
by: Panek, Vojtech, et al.
Published: (2024)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
by: Tran, Duc Dang Trung, et al.
Published: (2024)
by: Tran, Duc Dang Trung, et al.
Published: (2024)
AutoInst: Automatic Instance-Based Segmentation of LiDAR 3D Scans
by: Perauer, Cedric, et al.
Published: (2024)
by: Perauer, Cedric, et al.
Published: (2024)
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
by: Wang, Zishuo, et al.
Published: (2024)
by: Wang, Zishuo, et al.
Published: (2024)
T-HITL Effectively Addresses Problematic Associations in Image Generation and Maintains Overall Visual Quality
by: Epstein, Susan, et al.
Published: (2024)
by: Epstein, Susan, et al.
Published: (2024)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
by: Oliveira, Daniel, et al.
Published: (2026)
by: Oliveira, Daniel, et al.
Published: (2026)
Toward Simple and Robust Contrastive Explanations for Image Classification by Leveraging Instance Similarity and Concept Relevance
by: Kaidashova, Yuliia, et al.
Published: (2025)
by: Kaidashova, Yuliia, et al.
Published: (2025)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
by: Yan, Xiaoyang, et al.
Published: (2026)
by: Yan, Xiaoyang, et al.
Published: (2026)
Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
by: Estepa, Imanol G., et al.
Published: (2026)
by: Estepa, Imanol G., et al.
Published: (2026)
Similar Items
-
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
by: Zheng, Minghang, et al.
Published: (2024) -
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024) -
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
by: Zheng, Peng, et al.
Published: (2024) -
Perceptual Flow Network for Visually Grounded Reasoning
by: Li, Yangfu, et al.
Published: (2026) -
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)