Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
Fuente:
arXiv
Guardado en:
| Autores principales: | Luo, Yuanhao, Wen, Di, Peng, Kunyu, Liu, Ruiping, Zheng, Junwei, Chen, Yufan, Wei, Jiale, Stiefelhage, Rainer |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
por: Wen, Di, et al.
Publicado: (2025)
por: Wen, Di, et al.
Publicado: (2025)
Graph-based Document Structure Analysis
por: Chen, Yufan, et al.
Publicado: (2025)
por: Chen, Yufan, et al.
Publicado: (2025)
Snap, Segment, Deploy: A Visual Data and Detection Pipeline for Wearable Industrial Assistants
por: Wen, Di, et al.
Publicado: (2025)
por: Wen, Di, et al.
Publicado: (2025)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
por: Liu, Ruiping, et al.
Publicado: (2026)
por: Liu, Ruiping, et al.
Publicado: (2026)
MICA: Multi-Agent Industrial Coordination Assistant
por: Wen, Di, et al.
Publicado: (2025)
por: Wen, Di, et al.
Publicado: (2025)
Towards Multi-Source Domain Generalization for Sleep Staging with Noisy Labels
por: Wang, Kening, et al.
Publicado: (2026)
por: Wang, Kening, et al.
Publicado: (2026)
SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D
por: Wang, Zirui, et al.
Publicado: (2026)
por: Wang, Zirui, et al.
Publicado: (2026)
Skeleton-Based Human Action Recognition with Noisy Labels
por: Xu, Yi, et al.
Publicado: (2024)
por: Xu, Yi, et al.
Publicado: (2024)
RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
por: Chen, Yufan, et al.
Publicado: (2024)
por: Chen, Yufan, et al.
Publicado: (2024)
HybriDLA: Hybrid Generation for Document Layout Analysis
por: Chen, Yufan, et al.
Publicado: (2025)
por: Chen, Yufan, et al.
Publicado: (2025)
Exploring Video-Based Driver Activity Recognition under Noisy Labels
por: Fan, Linjuan, et al.
Publicado: (2025)
por: Fan, Linjuan, et al.
Publicado: (2025)
$M^2$-Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs
por: Lin, Kaixin, et al.
Publicado: (2026)
por: Lin, Kaixin, et al.
Publicado: (2026)
Referring Atomic Video Action Recognition
por: Peng, Kunyu, et al.
Publicado: (2024)
por: Peng, Kunyu, et al.
Publicado: (2024)
IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory
por: Kong, Weitong, et al.
Publicado: (2026)
por: Kong, Weitong, et al.
Publicado: (2026)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
por: Zhang, Haoshen, et al.
Publicado: (2026)
por: Zhang, Haoshen, et al.
Publicado: (2026)
IMPACT-Scribe: Interactive Temporal Action Segmentation with Boundary Scribbles and Query Planning
por: Yin, Qian, et al.
Publicado: (2026)
por: Yin, Qian, et al.
Publicado: (2026)
Not an Obstacle for Dog, but a Hazard for Human: A Co-Ego Navigation System for Guide Dog Robots
por: Liu, Ruiping, et al.
Publicado: (2026)
por: Liu, Ruiping, et al.
Publicado: (2026)
Open Panoramic Segmentation
por: Zheng, Junwei, et al.
Publicado: (2024)
por: Zheng, Junwei, et al.
Publicado: (2024)
What if? Emulative Simulation with World Models for Situated Reasoning
por: Liu, Ruiping, et al.
Publicado: (2026)
por: Liu, Ruiping, et al.
Publicado: (2026)
EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
por: Peng, Kunyu, et al.
Publicado: (2025)
por: Peng, Kunyu, et al.
Publicado: (2025)
Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
por: Liu, Ruiping, et al.
Publicado: (2025)
por: Liu, Ruiping, et al.
Publicado: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
por: Wei, Yiping, et al.
Publicado: (2023)
por: Wei, Yiping, et al.
Publicado: (2023)
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
por: Peng, Kunyu, et al.
Publicado: (2025)
por: Peng, Kunyu, et al.
Publicado: (2025)
Mitigating Label Noise using Prompt-Based Hyperbolic Meta-Learning in Open-Set Domain Generalization
por: Peng, Kunyu, et al.
Publicado: (2024)
por: Peng, Kunyu, et al.
Publicado: (2024)
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
por: Materia, Daniele, et al.
Publicado: (2026)
por: Materia, Daniele, et al.
Publicado: (2026)
Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
por: Liu, Ruiping, et al.
Publicado: (2024)
por: Liu, Ruiping, et al.
Publicado: (2024)
HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios
por: Peng, Kunyu, et al.
Publicado: (2025)
por: Peng, Kunyu, et al.
Publicado: (2025)
Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
por: Li, Wenxin, et al.
Publicado: (2025)
por: Li, Wenxin, et al.
Publicado: (2025)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
por: Chen, Yifei, et al.
Publicado: (2023)
por: Chen, Yifei, et al.
Publicado: (2023)
RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization
por: Zheng, Junwei, et al.
Publicado: (2026)
por: Zheng, Junwei, et al.
Publicado: (2026)
Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
por: Wen, Di, et al.
Publicado: (2025)
por: Wen, Di, et al.
Publicado: (2025)
Comb, Prune, Distill: Towards Unified Pruning for Vision Model Compression
por: Schmitt, Jonas, et al.
Publicado: (2024)
por: Schmitt, Jonas, et al.
Publicado: (2024)
Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler
por: Peng, Kunyu, et al.
Publicado: (2024)
por: Peng, Kunyu, et al.
Publicado: (2024)
Can we Trust Unreliable Voxels? Exploring 3D Semantic Occupancy Prediction under Label Noise
por: Li, Wenxin, et al.
Publicado: (2026)
por: Li, Wenxin, et al.
Publicado: (2026)
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
por: Wei, Jiale, et al.
Publicado: (2024)
por: Wei, Jiale, et al.
Publicado: (2024)
InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing
por: Yang, Yebin, et al.
Publicado: (2026)
por: Yang, Yebin, et al.
Publicado: (2026)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
por: Liu, Ruiping, et al.
Publicado: (2024)
por: Liu, Ruiping, et al.
Publicado: (2024)
MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments
por: Zheng, Junwei, et al.
Publicado: (2023)
por: Zheng, Junwei, et al.
Publicado: (2023)
Scene-agnostic Pose Regression for Visual Localization
por: Zheng, Junwei, et al.
Publicado: (2025)
por: Zheng, Junwei, et al.
Publicado: (2025)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
por: Cho, Hyunjin, et al.
Publicado: (2024)
por: Cho, Hyunjin, et al.
Publicado: (2024)
Ejemplares similares
-
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
por: Wen, Di, et al.
Publicado: (2025) -
Graph-based Document Structure Analysis
por: Chen, Yufan, et al.
Publicado: (2025) -
Snap, Segment, Deploy: A Visual Data and Detection Pipeline for Wearable Industrial Assistants
por: Wen, Di, et al.
Publicado: (2025) -
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
por: Liu, Ruiping, et al.
Publicado: (2026) -
MICA: Multi-Agent Industrial Coordination Assistant
por: Wen, Di, et al.
Publicado: (2025)