Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yike, Bolucu, Necva, Wan, Stephen, Wang, Dadong, Xia, Jiahao, Zhang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression
by: Ke, Jingcheng, et al.
Published: (2024)
by: Ke, Jingcheng, et al.
Published: (2024)
GAIA: Zero-shot Talking Avatar Generation
by: He, Tianyu, et al.
Published: (2023)
by: He, Tianyu, et al.
Published: (2023)
Modularized Zero-shot VQA with Pre-trained Models
by: Cao, Rui, et al.
Published: (2023)
by: Cao, Rui, et al.
Published: (2023)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2024)
by: Zhu, Jiaqi, et al.
Published: (2024)
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation
by: Zhang, Yingying, et al.
Published: (2024)
by: Zhang, Yingying, et al.
Published: (2024)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
by: Zhang, Zhi, et al.
Published: (2023)
by: Zhang, Zhi, et al.
Published: (2023)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
by: Yang, Danni, et al.
Published: (2024)
by: Yang, Danni, et al.
Published: (2024)
DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes
by: Yan, Hao, et al.
Published: (2024)
by: Yan, Hao, et al.
Published: (2024)
Scene Graph Generation with Role-Playing Large Language Models
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
by: Jiang, Diqiong, et al.
Published: (2026)
by: Jiang, Diqiong, et al.
Published: (2026)
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
by: Li, Yanjun, et al.
Published: (2025)
by: Li, Yanjun, et al.
Published: (2025)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026)
by: Meng, Jiahao, et al.
Published: (2026)
Zero-shot image privacy classification with Vision-Language Models
by: Baia, Alina Elena, et al.
Published: (2025)
by: Baia, Alina Elena, et al.
Published: (2025)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
by: Lan, Xing, et al.
Published: (2024)
by: Lan, Xing, et al.
Published: (2024)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
by: Duan, Yiqun, et al.
Published: (2025)
by: Duan, Yiqun, et al.
Published: (2025)
Selective Vision-Language Subspace Projection for Few-shot CLIP
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
by: An, Wenbin, et al.
Published: (2025)
by: An, Wenbin, et al.
Published: (2025)
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
A Unified Non-Parametric and Interpretable Point Cloud Analysis via t-FCW Graph Representation
by: Lai, Haijian, et al.
Published: (2026)
by: Lai, Haijian, et al.
Published: (2026)
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
GaussianForest: Hierarchical-Hybrid 3D Gaussian Splatting for Compressed Scene Modeling
by: Zhang, Fengyi, et al.
Published: (2024)
by: Zhang, Fengyi, et al.
Published: (2024)
MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues
by: Zhang, Liyun
Published: (2024)
by: Zhang, Liyun
Published: (2024)
Pseudo-triplet Guided Few-shot Composed Image Retrieval
by: Hou, Bohan, et al.
Published: (2024)
by: Hou, Bohan, et al.
Published: (2024)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Referring Flexible Image Restoration
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
Querying Autonomous Vehicle Point Clouds: Enhanced by 3D Object Counting with CounterNet
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Cross-domain Multi-step Thinking: Zero-shot Fine-grained Traffic Sign Recognition in the Wild
by: Gan, Yaozong, et al.
Published: (2024)
by: Gan, Yaozong, et al.
Published: (2024)
Dynamic Resolution Guidance for Facial Expression Recognition
by: Wang, Songpan, et al.
Published: (2024)
by: Wang, Songpan, et al.
Published: (2024)
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
by: Hong, Jiacheng, et al.
Published: (2025)
by: Hong, Jiacheng, et al.
Published: (2025)
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
by: Li, Zeju, et al.
Published: (2024)
by: Li, Zeju, et al.
Published: (2024)
Optimized Learned Image Compression for Facial Expression Recognition
by: Li, Xiumei, et al.
Published: (2025)
by: Li, Xiumei, et al.
Published: (2025)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
by: Liu, Lingyu, et al.
Published: (2026)
by: Liu, Lingyu, et al.
Published: (2026)
AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
by: Chen, Liyang, et al.
Published: (2023)
by: Chen, Liyang, et al.
Published: (2023)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
by: Shan, Ziyu, et al.
Published: (2024)
by: Shan, Ziyu, et al.
Published: (2024)
Similar Items
-
Make Graph-based Referring Expression Comprehension Great Again through Expression-guided Dynamic Gating and Regression
by: Ke, Jingcheng, et al.
Published: (2024) -
GAIA: Zero-shot Talking Avatar Generation
by: He, Tianyu, et al.
Published: (2023) -
Modularized Zero-shot VQA with Pre-trained Models
by: Cao, Rui, et al.
Published: (2023) -
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
by: Liu, Yu, et al.
Published: (2025) -
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2024)