Open-World Human-Object Interaction Detection via Multi-modal Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Jie, Li, Bingliang, Zeng, Ailing, Zhang, Lei, Zhang, Ruimao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
X-Pose: Detecting Any Keypoints
di: Yang, Jie, et al.
Pubblicazione: (2023)
di: Yang, Jie, et al.
Pubblicazione: (2023)
FreeMan: Towards Benchmarking 3D Human Pose Estimation under Real-World Conditions
di: Wang, Jiong, et al.
Pubblicazione: (2023)
di: Wang, Jiong, et al.
Pubblicazione: (2023)
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
di: Chen, Ling-Hao, et al.
Pubblicazione: (2024)
di: Chen, Ling-Hao, et al.
Pubblicazione: (2024)
F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions
di: Yang, Jie, et al.
Pubblicazione: (2024)
di: Yang, Jie, et al.
Pubblicazione: (2024)
Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
di: Lin, Jing, et al.
Pubblicazione: (2023)
di: Lin, Jing, et al.
Pubblicazione: (2023)
Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
di: Cai, Zhaolin, et al.
Pubblicazione: (2025)
di: Cai, Zhaolin, et al.
Pubblicazione: (2025)
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
di: Qin, Yiran, et al.
Pubblicazione: (2023)
di: Qin, Yiran, et al.
Pubblicazione: (2023)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
di: Guo, Pinxue, et al.
Pubblicazione: (2024)
di: Guo, Pinxue, et al.
Pubblicazione: (2024)
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
di: Zhang, Yuhong, et al.
Pubblicazione: (2025)
di: Zhang, Yuhong, et al.
Pubblicazione: (2025)
Semantic-Supervised Spatial-Temporal Fusion for LiDAR-based 3D Object Detection
di: Wang, Chaoqun, et al.
Pubblicazione: (2025)
di: Wang, Chaoqun, et al.
Pubblicazione: (2025)
MR-GDINO: Efficient Open-World Continual Object Detection
di: Dong, Bowen, et al.
Pubblicazione: (2024)
di: Dong, Bowen, et al.
Pubblicazione: (2024)
RGB-T Object Detection via Group Shuffled Multi-receptive Attention and Multi-modal Supervision
di: Wang, Jinzhong, et al.
Pubblicazione: (2024)
di: Wang, Jinzhong, et al.
Pubblicazione: (2024)
Unified-modal Salient Object Detection via Adaptive Prompt Learning
di: Wang, Kunpeng, et al.
Pubblicazione: (2023)
di: Wang, Kunpeng, et al.
Pubblicazione: (2023)
Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and Calibration
di: Wang, Chaoqun, et al.
Pubblicazione: (2024)
di: Wang, Chaoqun, et al.
Pubblicazione: (2024)
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
di: Jiang, Qing, et al.
Pubblicazione: (2024)
di: Jiang, Qing, et al.
Pubblicazione: (2024)
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
di: Zhang, Zhenhao, et al.
Pubblicazione: (2025)
di: Zhang, Zhenhao, et al.
Pubblicazione: (2025)
Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection
di: Jia, Mingda, et al.
Pubblicazione: (2024)
di: Jia, Mingda, et al.
Pubblicazione: (2024)
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
di: Ren, Tianhe, et al.
Pubblicazione: (2024)
di: Ren, Tianhe, et al.
Pubblicazione: (2024)
Streamlined Open-Vocabulary Human-Object Interaction Detection
di: Sun, Chang, et al.
Pubblicazione: (2026)
di: Sun, Chang, et al.
Pubblicazione: (2026)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
di: Wang, Yongqi, et al.
Pubblicazione: (2024)
di: Wang, Yongqi, et al.
Pubblicazione: (2024)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
di: Zhang, Jinlu, et al.
Pubblicazione: (2025)
di: Zhang, Jinlu, et al.
Pubblicazione: (2025)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
di: Liu, Shilong, et al.
Pubblicazione: (2023)
di: Liu, Shilong, et al.
Pubblicazione: (2023)
Prototype Embedding Optimization for Human-Object Interaction Detection in Livestreaming
di: Zhang, Menghui, et al.
Pubblicazione: (2025)
di: Zhang, Menghui, et al.
Pubblicazione: (2025)
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
di: Lei, Ting, et al.
Pubblicazione: (2025)
di: Lei, Ting, et al.
Pubblicazione: (2025)
No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
di: Yang, Bin, et al.
Pubblicazione: (2025)
di: Yang, Bin, et al.
Pubblicazione: (2025)
Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions
di: Wen, Boran, et al.
Pubblicazione: (2025)
di: Wen, Boran, et al.
Pubblicazione: (2025)
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
di: Li, Jiaming, et al.
Pubblicazione: (2024)
di: Li, Jiaming, et al.
Pubblicazione: (2024)
Unlock the Power of Unlabeled Data in Language Driving Model
di: Wang, Chaoqun, et al.
Pubblicazione: (2025)
di: Wang, Chaoqun, et al.
Pubblicazione: (2025)
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
di: Wang, Kunpeng, et al.
Pubblicazione: (2024)
di: Wang, Kunpeng, et al.
Pubblicazione: (2024)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
di: Zhang, Zhenhao, et al.
Pubblicazione: (2025)
di: Zhang, Zhenhao, et al.
Pubblicazione: (2025)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
di: Zeng, Ruizhe, et al.
Pubblicazione: (2024)
di: Zeng, Ruizhe, et al.
Pubblicazione: (2024)
Dual-Modal Prompting for Sketch-Based Image Retrieval
di: Gao, Liying, et al.
Pubblicazione: (2024)
di: Gao, Liying, et al.
Pubblicazione: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
di: Ren, Tianhe, et al.
Pubblicazione: (2024)
di: Ren, Tianhe, et al.
Pubblicazione: (2024)
Open World Object Detection: A Survey
di: Li, Yiming, et al.
Pubblicazione: (2024)
di: Li, Yiming, et al.
Pubblicazione: (2024)
Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
di: Heo, Jun-Woo, et al.
Pubblicazione: (2026)
di: Heo, Jun-Woo, et al.
Pubblicazione: (2026)
Task-Driven Prompt Learning: A Joint Framework for Multi-modal Cloud Removal and Segmentation
di: Zhang, Zaiyan, et al.
Pubblicazione: (2026)
di: Zhang, Zaiyan, et al.
Pubblicazione: (2026)
Beyond Flat Unknown Labels in Open-World Object Detection
di: Zhang, Yuchen, et al.
Pubblicazione: (2025)
di: Zhang, Yuchen, et al.
Pubblicazione: (2025)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
Unified Open-World Segmentation with Multi-Modal Prompts
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Visual Consensus Prompting for Co-Salient Object Detection
di: Wang, Jie, et al.
Pubblicazione: (2025)
di: Wang, Jie, et al.
Pubblicazione: (2025)
Documenti analoghi
-
X-Pose: Detecting Any Keypoints
di: Yang, Jie, et al.
Pubblicazione: (2023) -
FreeMan: Towards Benchmarking 3D Human Pose Estimation under Real-World Conditions
di: Wang, Jiong, et al.
Pubblicazione: (2023) -
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
di: Chen, Ling-Hao, et al.
Pubblicazione: (2024) -
F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions
di: Yang, Jie, et al.
Pubblicazione: (2024) -
Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
di: Lin, Jing, et al.
Pubblicazione: (2023)