HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Junwen, Xiong, Peilin, Yanai, Keiji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
di: Xiong, Peilin, et al.
Pubblicazione: (2025)
di: Xiong, Peilin, et al.
Pubblicazione: (2025)
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
di: Xiong, Peilin, et al.
Pubblicazione: (2026)
di: Xiong, Peilin, et al.
Pubblicazione: (2026)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
di: Kang, Donggoo, et al.
Pubblicazione: (2024)
di: Kang, Donggoo, et al.
Pubblicazione: (2024)
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
di: Li, Shujia, et al.
Pubblicazione: (2025)
di: Li, Shujia, et al.
Pubblicazione: (2025)
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
di: Ji, Yihong, et al.
Pubblicazione: (2025)
di: Ji, Yihong, et al.
Pubblicazione: (2025)
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
di: Yang, Panqi, et al.
Pubblicazione: (2025)
di: Yang, Panqi, et al.
Pubblicazione: (2025)
Focusing on what to decode and what to train: SOV Decoding with Specific Target Guided DeNoising and Vision Language Advisor
di: Chen, Junwen, et al.
Pubblicazione: (2023)
di: Chen, Junwen, et al.
Pubblicazione: (2023)
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
di: Benishu, Omer, et al.
Pubblicazione: (2026)
di: Benishu, Omer, et al.
Pubblicazione: (2026)
Contextual Object Detection with Multimodal Large Language Models
di: Zang, Yuhang, et al.
Pubblicazione: (2023)
di: Zang, Yuhang, et al.
Pubblicazione: (2023)
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
di: Hu, BoCheng, et al.
Pubblicazione: (2026)
di: Hu, BoCheng, et al.
Pubblicazione: (2026)
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
di: Zeng, Huajian, et al.
Pubblicazione: (2026)
di: Zeng, Huajian, et al.
Pubblicazione: (2026)
Can Multimodal Large Language Models Truly Understand Small Objects?
di: Han, Fujun, et al.
Pubblicazione: (2026)
di: Han, Fujun, et al.
Pubblicazione: (2026)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
di: Wu, Wentao, et al.
Pubblicazione: (2025)
di: Wu, Wentao, et al.
Pubblicazione: (2025)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)
di: Sapkota, Ranjan, et al.
Pubblicazione: (2025)
A Review of Human-Object Interaction Detection
di: Wang, Yuxiao, et al.
Pubblicazione: (2024)
di: Wang, Yuxiao, et al.
Pubblicazione: (2024)
2AFC Prompting of Large Multimodal Models for Image Quality Assessment
di: Zhu, Hanwei, et al.
Pubblicazione: (2024)
di: Zhu, Hanwei, et al.
Pubblicazione: (2024)
Mitigating Long-Tail Bias in HOI Detection via Adaptive Diversity Cache
di: Jiang, Yuqiu, et al.
Pubblicazione: (2025)
di: Jiang, Yuqiu, et al.
Pubblicazione: (2025)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
di: Zhang, Zhenhao, et al.
Pubblicazione: (2025)
di: Zhang, Zhenhao, et al.
Pubblicazione: (2025)
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
di: Lei, Ting, et al.
Pubblicazione: (2024)
di: Lei, Ting, et al.
Pubblicazione: (2024)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
di: Zhang, Haoshen, et al.
Pubblicazione: (2026)
di: Zhang, Haoshen, et al.
Pubblicazione: (2026)
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
di: Huo, Fushuo, et al.
Pubblicazione: (2024)
di: Huo, Fushuo, et al.
Pubblicazione: (2024)
MammothModa: Multi-Modal Large Language Model
di: She, Qi, et al.
Pubblicazione: (2024)
di: She, Qi, et al.
Pubblicazione: (2024)
Zero-Shot Human-Object Interaction Synthesis with Multimodal Priors
di: Lou, Yuke, et al.
Pubblicazione: (2025)
di: Lou, Yuke, et al.
Pubblicazione: (2025)
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization
di: Yu, JiangYong, et al.
Pubblicazione: (2025)
di: Yu, JiangYong, et al.
Pubblicazione: (2025)
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
di: Liu, Chang, et al.
Pubblicazione: (2024)
di: Liu, Chang, et al.
Pubblicazione: (2024)
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
di: Li, Jiansheng, et al.
Pubblicazione: (2025)
di: Li, Jiansheng, et al.
Pubblicazione: (2025)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
di: Xu, Sirui, et al.
Pubblicazione: (2024)
di: Xu, Sirui, et al.
Pubblicazione: (2024)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
di: Qiu, Yicheng, et al.
Pubblicazione: (2026)
di: Qiu, Yicheng, et al.
Pubblicazione: (2026)
SHOE: Semantic HOI Open-Vocabulary Evaluation Metric
di: Noack, Maja, et al.
Pubblicazione: (2026)
di: Noack, Maja, et al.
Pubblicazione: (2026)
ContextHOI: Spatial Context Learning for Human-Object Interaction Detection
di: Jia, Mingda, et al.
Pubblicazione: (2024)
di: Jia, Mingda, et al.
Pubblicazione: (2024)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
di: Gao, Jianjun, et al.
Pubblicazione: (2024)
di: Gao, Jianjun, et al.
Pubblicazione: (2024)
A Study of Failure Modes in Two-Stage Human-Object Interaction Detection
di: Wang, Lemeng, et al.
Pubblicazione: (2026)
di: Wang, Lemeng, et al.
Pubblicazione: (2026)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
di: Luo, Yuanhao, et al.
Pubblicazione: (2026)
di: Luo, Yuanhao, et al.
Pubblicazione: (2026)
TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
di: Yu, Ya-Qi, et al.
Pubblicazione: (2024)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2024)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
di: Ashqar, Huthaifa I., et al.
Pubblicazione: (2024)
di: Ashqar, Huthaifa I., et al.
Pubblicazione: (2024)
Human-Object Interaction from Human-Level Instructions
di: Wu, Zhen, et al.
Pubblicazione: (2024)
di: Wu, Zhen, et al.
Pubblicazione: (2024)
EventVL: Understand Event Streams via Multimodal Large Language Model
di: Li, Pengteng, et al.
Pubblicazione: (2025)
di: Li, Pengteng, et al.
Pubblicazione: (2025)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
di: Yao, Xincheng, et al.
Pubblicazione: (2026)
di: Yao, Xincheng, et al.
Pubblicazione: (2026)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
di: Seo, Soo Won, et al.
Pubblicazione: (2026)
di: Seo, Soo Won, et al.
Pubblicazione: (2026)
Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning
di: Hao, Pengfei, et al.
Pubblicazione: (2025)
di: Hao, Pengfei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
di: Xiong, Peilin, et al.
Pubblicazione: (2025) -
BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing
di: Xiong, Peilin, et al.
Pubblicazione: (2026) -
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
di: Kang, Donggoo, et al.
Pubblicazione: (2024) -
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
di: Li, Shujia, et al.
Pubblicazione: (2025) -
OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
di: Ji, Yihong, et al.
Pubblicazione: (2025)