A Study of Failure Modes in Two-Stage Human-Object Interaction Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Lemeng, Lei, Qinqian, Bakshi, Vidhi, Yi, Daniel, Liu, Yifan, Hou, Jiacheng, Hao, Asher Seng, Mai, Zheda, Chao, Wei-Lun, Tan, Robby T., Wang, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
by: Jeon, Sooyoung, et al.
Published: (2026)
by: Jeon, Sooyoung, et al.
Published: (2026)
CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
by: Lei, Qinqian, et al.
Published: (2025)
by: Lei, Qinqian, et al.
Published: (2025)
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
by: Lei, Qinqian, et al.
Published: (2024)
by: Lei, Qinqian, et al.
Published: (2024)
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
by: Lei, Qinqian, et al.
Published: (2025)
by: Lei, Qinqian, et al.
Published: (2025)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
by: Mai, Zheda, et al.
Published: (2025)
by: Mai, Zheda, et al.
Published: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Revisiting semi-supervised learning in the era of foundation models
by: Zhang, Ping, et al.
Published: (2025)
by: Zhang, Ping, et al.
Published: (2025)
SHOE: Semantic HOI Open-Vocabulary Evaluation Metric
by: Noack, Maja, et al.
Published: (2026)
by: Noack, Maja, et al.
Published: (2026)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)
by: Hoe, Jiun Tian, et al.
Published: (2025)
Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
by: Lee, Justin, et al.
Published: (2025)
by: Lee, Justin, et al.
Published: (2025)
3DOT: Texture Transfer for 3DGS Objects from a Single Reference Image
by: Cao, Xiao, et al.
Published: (2025)
by: Cao, Xiao, et al.
Published: (2025)
Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition
by: Mai, Zheda, et al.
Published: (2024)
by: Mai, Zheda, et al.
Published: (2024)
OneHOI: Unifying Human-Object Interaction Generation and Editing
by: Hoe, Jiun Tian, et al.
Published: (2026)
by: Hoe, Jiun Tian, et al.
Published: (2026)
White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation
by: Li, Shuwei, et al.
Published: (2026)
by: Li, Shuwei, et al.
Published: (2026)
Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image Translation
by: Li, Shuwei, et al.
Published: (2026)
by: Li, Shuwei, et al.
Published: (2026)
CAT: Exploiting Inter-Class Dynamics for Domain Adaptive Object Detection
by: Kennerley, Mikhail, et al.
Published: (2024)
by: Kennerley, Mikhail, et al.
Published: (2024)
Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions
by: Ai, Yihao, et al.
Published: (2024)
by: Ai, Yihao, et al.
Published: (2024)
UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation
by: Chen, Haopeng, et al.
Published: (2026)
by: Chen, Haopeng, et al.
Published: (2026)
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
Aggregating Diverse Cue Experts for AI-Generated Image Detection
by: Tan, Lei, et al.
Published: (2026)
by: Tan, Lei, et al.
Published: (2026)
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images
by: Wang, Benzhi, et al.
Published: (2024)
by: Wang, Benzhi, et al.
Published: (2024)
A Review of Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2024)
by: Wang, Yuxiao, et al.
Published: (2024)
SSNeRF: Sparse View Semi-supervised Neural Radiance Fields with Augmentation
by: Cao, Xiao, et al.
Published: (2024)
by: Cao, Xiao, et al.
Published: (2024)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
Revisiting Model Stitching In the Foundation Model Era
by: Mai, Zheda, et al.
Published: (2026)
by: Mai, Zheda, et al.
Published: (2026)
Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection
by: Bai, Yanbing, et al.
Published: (2025)
by: Bai, Yanbing, et al.
Published: (2025)
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
by: Mao, Aihua, et al.
Published: (2026)
by: Mao, Aihua, et al.
Published: (2026)
Bridging Annotation Gaps: Transferring Labels to Align Object Detection Datasets
by: Kennerley, Mikhail, et al.
Published: (2025)
by: Kennerley, Mikhail, et al.
Published: (2025)
HEAP: Unsupervised Object Discovery and Localization with Contrastive Grouping
by: Zhang, Xin, et al.
Published: (2023)
by: Zhang, Xin, et al.
Published: (2023)
Revisiting Code Search in a Two-Stage Paradigm
by: Hu, Fan, et al.
Published: (2022)
by: Hu, Fan, et al.
Published: (2022)
Failure of the local-global principle for isotropy of quadratic forms over function fields
by: Auel, Asher, et al.
Published: (2017)
by: Auel, Asher, et al.
Published: (2017)
Rapid mixing for high-temperature Gibbs states with arbitrary external fields
by: Bakshi, Ainesh, et al.
Published: (2026)
by: Bakshi, Ainesh, et al.
Published: (2026)
BlackBoxToBlueprint: Extracting Interpretable Logic from Legacy Systems using Reinforcement Learning and Counterfactual Analysis
by: Rathore, Vidhi
Published: (2025)
by: Rathore, Vidhi
Published: (2025)
Causal Manifold Fairness: Enforcing Geometric Invariance in Representation Learning
by: Rathore, Vidhi
Published: (2026)
by: Rathore, Vidhi
Published: (2026)
InverseMeetInsert: Robust Real Image Editing via Geometric Accumulation Inversion in Guided Diffusion Models
by: Zheng, Yan, et al.
Published: (2024)
by: Zheng, Yan, et al.
Published: (2024)
Learning Human-Object Interaction as Groups
by: Hong, Jiajun, et al.
Published: (2025)
by: Hong, Jiajun, et al.
Published: (2025)
Two-Stage Radio Map Construction with Real Environments and Sparse Measurements
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation
by: Zhu, Xiaomeng, et al.
Published: (2025)
by: Zhu, Xiaomeng, et al.
Published: (2025)
Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector
by: Wu, Qirui, et al.
Published: (2025)
by: Wu, Qirui, et al.
Published: (2025)
Similar Items
-
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
by: Jeon, Sooyoung, et al.
Published: (2026) -
CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
by: Lei, Qinqian, et al.
Published: (2025) -
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
by: Lei, Qinqian, et al.
Published: (2024) -
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
by: Lei, Qinqian, et al.
Published: (2025) -
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
by: Mai, Zheda, et al.
Published: (2025)