Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yupeng, Ding, Changxing, Sun, Chang, Huang, Shaoli, Xu, Xiangmin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streamlined Open-Vocabulary Human-Object Interaction Detection
by: Sun, Chang, et al.
Published: (2026)
by: Sun, Chang, et al.
Published: (2026)
Disentangled Pre-training for Human-Object Interaction Detection
by: Li, Zhuolong, et al.
Published: (2024)
by: Li, Zhuolong, et al.
Published: (2024)
Guiding Human-Object Interactions with Rich Geometry and Relations
by: Xue, Mengqing, et al.
Published: (2025)
by: Xue, Mengqing, et al.
Published: (2025)
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
by: Xue, Weiying, et al.
Published: (2024)
by: Xue, Weiying, et al.
Published: (2024)
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions
by: Wen, Boran, et al.
Published: (2025)
by: Wen, Boran, et al.
Published: (2025)
Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification
by: Jiang, Jiayu, et al.
Published: (2025)
by: Jiang, Jiayu, et al.
Published: (2025)
ViHOI: Human-Object Interaction Synthesis with Visual Priors
by: Cai, Songjin, et al.
Published: (2026)
by: Cai, Songjin, et al.
Published: (2026)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
by: Wang, Kuo, et al.
Published: (2024)
by: Wang, Kuo, et al.
Published: (2024)
COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection
by: Zhang, Yupeng, et al.
Published: (2026)
by: Zhang, Yupeng, et al.
Published: (2026)
NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
by: Zhang, Yupeng, et al.
Published: (2026)
by: Zhang, Yupeng, et al.
Published: (2026)
LV-OSD: Language-Vision-Complementary Open-Set Object Detection
by: Zhang, Yupeng, et al.
Published: (2026)
by: Zhang, Yupeng, et al.
Published: (2026)
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
by: Zhao, Kai, et al.
Published: (2025)
by: Zhao, Kai, et al.
Published: (2025)
Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy
by: Deng, Zekai, et al.
Published: (2025)
by: Deng, Zekai, et al.
Published: (2025)
From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects
by: Li, Zizhao, et al.
Published: (2024)
by: Li, Zizhao, et al.
Published: (2024)
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
by: Nguyen, Minh Anh, et al.
Published: (2026)
by: Nguyen, Minh Anh, et al.
Published: (2026)
Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
by: Li, Liulei, et al.
Published: (2024)
by: Li, Liulei, et al.
Published: (2024)
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
A Training-Free Guess What Vision Language Model from Snippets to Open-Vocabulary Object Detection
by: Zhu, Guiying, et al.
Published: (2026)
by: Zhu, Guiying, et al.
Published: (2026)
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023)
by: Minderer, Matthias, et al.
Published: (2023)
Democratizing High-Fidelity Co-Speech Gesture Video Generation
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
by: Du, Penghui, et al.
Published: (2024)
by: Du, Penghui, et al.
Published: (2024)
Open-Vocabulary Object Detection via Language Hierarchy
by: Huang, Jiaxing, et al.
Published: (2024)
by: Huang, Jiaxing, et al.
Published: (2024)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
by: Zeng, Ruizhe, et al.
Published: (2024)
by: Zeng, Ruizhe, et al.
Published: (2024)
Learning to Detect and Segment for Open Vocabulary Object Detection
by: Wang, Tao, et al.
Published: (2022)
by: Wang, Tao, et al.
Published: (2022)
Retrieval-Augmented Open-Vocabulary Object Detection
by: Kim, Jooyeon, et al.
Published: (2024)
by: Kim, Jooyeon, et al.
Published: (2024)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
by: Ranasinghe, Yasiru, et al.
Published: (2026)
by: Ranasinghe, Yasiru, et al.
Published: (2026)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
by: Chen, Fangyi, et al.
Published: (2024)
by: Chen, Fangyi, et al.
Published: (2024)
Superpowering Open-Vocabulary Object Detectors for X-ray Vision
by: Garcia-Fernandez, Pablo, et al.
Published: (2025)
by: Garcia-Fernandez, Pablo, et al.
Published: (2025)
Realistic Human Motion Generation with Cross-Diffusion Models
by: Ren, Zeping, et al.
Published: (2023)
by: Ren, Zeping, et al.
Published: (2023)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
On the Potential of Open-Vocabulary Models for Object Detection in Unusual Street Scenes
by: Ilyas, Sadia, et al.
Published: (2024)
by: Ilyas, Sadia, et al.
Published: (2024)
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
by: Yao, Jianhang, et al.
Published: (2025)
by: Yao, Jianhang, et al.
Published: (2025)
Similar Items
-
Streamlined Open-Vocabulary Human-Object Interaction Detection
by: Sun, Chang, et al.
Published: (2026) -
Disentangled Pre-training for Human-Object Interaction Detection
by: Li, Zhuolong, et al.
Published: (2024) -
Guiding Human-Object Interactions with Rich Geometry and Relations
by: Xue, Mengqing, et al.
Published: (2025) -
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
by: Xue, Weiying, et al.
Published: (2024) -
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
by: Zhang, Yupeng, et al.
Published: (2025)