Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Xuan, Shiyu, Wang, Dongkai, Li, Zechao, Tang, Jinhui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
por: Xuan, Shiyu, et al.
Publicado: (2025)
por: Xuan, Shiyu, et al.
Publicado: (2025)
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
por: Wang, Dongkai, et al.
Publicado: (2024)
por: Wang, Dongkai, et al.
Publicado: (2024)
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
por: Guo, Yixin, et al.
Publicado: (2024)
por: Guo, Yixin, et al.
Publicado: (2024)
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection
por: Lei, Ting, et al.
Publicado: (2024)
por: Lei, Ting, et al.
Publicado: (2024)
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
por: Yang, Chanhyeong, et al.
Publicado: (2025)
por: Yang, Chanhyeong, et al.
Publicado: (2025)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
por: Liu, Yu, et al.
Publicado: (2025)
por: Liu, Yu, et al.
Publicado: (2025)
Visual Position Prompt for MLLM based Visual Grounding
por: Tang, Wei, et al.
Publicado: (2025)
por: Tang, Wei, et al.
Publicado: (2025)
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
por: Cao, Meiqi, et al.
Publicado: (2024)
por: Cao, Meiqi, et al.
Publicado: (2024)
Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection
por: Sarma, Sandipan, et al.
Publicado: (2025)
por: Sarma, Sandipan, et al.
Publicado: (2025)
AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation
por: Dai, Sisi, et al.
Publicado: (2025)
por: Dai, Sisi, et al.
Publicado: (2025)
A Recover-then-Discriminate Framework for Robust Anomaly Detection
por: Xing, Peng, et al.
Publicado: (2024)
por: Xing, Peng, et al.
Publicado: (2024)
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
por: Lei, Qinqian, et al.
Publicado: (2024)
por: Lei, Qinqian, et al.
Publicado: (2024)
Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model
por: Ge, Mengying, et al.
Publicado: (2024)
por: Ge, Mengying, et al.
Publicado: (2024)
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
por: Zhou, Qihang, et al.
Publicado: (2023)
por: Zhou, Qihang, et al.
Publicado: (2023)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
GlocalCLIP: Object-agnostic Global-Local Prompt Learning for Zero-shot Anomaly Detection
por: Ham, Jiyul, et al.
Publicado: (2024)
por: Ham, Jiyul, et al.
Publicado: (2024)
CycleHOI: Improving Human-Object Interaction Detection with Cycle Consistency of Detection and Generation
por: Wang, Yisen, et al.
Publicado: (2024)
por: Wang, Yisen, et al.
Publicado: (2024)
ContextHOI: Spatial Context Learning for Human-Object Interaction Detection
por: Jia, Mingda, et al.
Publicado: (2024)
por: Jia, Mingda, et al.
Publicado: (2024)
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
por: Lei, Qinqian, et al.
Publicado: (2025)
por: Lei, Qinqian, et al.
Publicado: (2025)
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
por: Li, Lei, et al.
Publicado: (2025)
por: Li, Lei, et al.
Publicado: (2025)
DQEN: Dual Query Enhancement Network for DETR-based HOI Detection
por: Li, Zhehao, et al.
Publicado: (2025)
por: Li, Zhehao, et al.
Publicado: (2025)
UAHOI: Uncertainty-aware Robust Interaction Learning for HOI Detection
por: Chen, Mu, et al.
Publicado: (2024)
por: Chen, Mu, et al.
Publicado: (2024)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
por: Jiang, Xin, et al.
Publicado: (2024)
por: Jiang, Xin, et al.
Publicado: (2024)
TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification
por: Yan, Rui, et al.
Publicado: (2025)
por: Yan, Rui, et al.
Publicado: (2025)
Decoupled Contrastive Learning for Long-Tailed Recognition
por: Xuan, Shiyu, et al.
Publicado: (2024)
por: Xuan, Shiyu, et al.
Publicado: (2024)
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
por: Lei, Ting, et al.
Publicado: (2025)
por: Lei, Ting, et al.
Publicado: (2025)
ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation
por: Zeng, Ling-An, et al.
Publicado: (2025)
por: Zeng, Ling-An, et al.
Publicado: (2025)
See the Text: From Tokenization to Visual Reading
por: Xing, Ling, et al.
Publicado: (2025)
por: Xing, Ling, et al.
Publicado: (2025)
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
por: Rao, Mingxing, et al.
Publicado: (2024)
por: Rao, Mingxing, et al.
Publicado: (2024)
ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
por: Li, Ao, et al.
Publicado: (2025)
por: Li, Ao, et al.
Publicado: (2025)
Zero-shot Compositional Action Recognition with Neural Logic Constraints
por: Ye, Gefan, et al.
Publicado: (2025)
por: Ye, Gefan, et al.
Publicado: (2025)
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
por: Wei, Lai, et al.
Publicado: (2024)
por: Wei, Lai, et al.
Publicado: (2024)
Contextual Interaction via Primitive-based Adversarial Training For Compositional Zero-shot Learning
por: Li, Suyi, et al.
Publicado: (2024)
por: Li, Suyi, et al.
Publicado: (2024)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
por: Dong, Jihao, et al.
Publicado: (2024)
por: Dong, Jihao, et al.
Publicado: (2024)
PartHOI: Part-based Hand-Object Interaction Transfer via Generalized Cylinders
por: Wang, Qiaochu, et al.
Publicado: (2025)
por: Wang, Qiaochu, et al.
Publicado: (2025)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
por: Xu, Binqian, et al.
Publicado: (2024)
por: Xu, Binqian, et al.
Publicado: (2024)
PA-HOI: A Physics-Aware Human and Object Interaction Dataset
por: Wang, Ruiyan, et al.
Publicado: (2025)
por: Wang, Ruiyan, et al.
Publicado: (2025)
Combating Noisy Labels through Fostering Self- and Neighbor-Consistency
por: Sun, Zeren, et al.
Publicado: (2026)
por: Sun, Zeren, et al.
Publicado: (2026)
Spatial Structure Constraints for Weakly Supervised Semantic Segmentation
por: Chen, Tao, et al.
Publicado: (2024)
por: Chen, Tao, et al.
Publicado: (2024)
CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
por: Lei, Qinqian, et al.
Publicado: (2025)
por: Lei, Qinqian, et al.
Publicado: (2025)
Ejemplares similares
-
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
por: Xuan, Shiyu, et al.
Publicado: (2025) -
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
por: Wang, Dongkai, et al.
Publicado: (2024) -
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
por: Guo, Yixin, et al.
Publicado: (2024) -
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection
por: Lei, Ting, et al.
Publicado: (2024) -
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
por: Yang, Chanhyeong, et al.
Publicado: (2025)