Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuxiao, Lei, Yu, Wei, Zhenao, Xue, Weiying, Jiang, Xinyu, Zhuang, Nan, Liu, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
FreeA: Human-object Interaction Detection using Free Annotation Labels
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
Precision-Enhanced Human-Object Contact Detection via Depth-Aware Perspective Interaction and Object Texture Restoration
by: Wang, Yuxiao, et al.
Published: (2024)
by: Wang, Yuxiao, et al.
Published: (2024)
A Review of Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2024)
by: Wang, Yuxiao, et al.
Published: (2024)
What-Meets-Where: Unified Learning of Action and Contact Localization in Images
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
by: Xue, Weiying, et al.
Published: (2024)
by: Xue, Weiying, et al.
Published: (2024)
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
The Components of Collaborative Joint Perception and Prediction -- A Conceptual Framework
by: Wan, Lei, et al.
Published: (2025)
by: Wan, Lei, et al.
Published: (2025)
DA-BEV: Unsupervised Domain Adaptation for Bird's Eye View Perception
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
PromptLNet: Region-Adaptive Aesthetic Enhancement via Prompt Guidance in Low-Light Enhancement Net
by: Yin, Jun, et al.
Published: (2025)
by: Yin, Jun, et al.
Published: (2025)
Exploring Hyperspectral Anomaly Detection with Human Vision: A Small Target Aware Detector
by: Ma, Jitao, et al.
Published: (2024)
by: Ma, Jitao, et al.
Published: (2024)
Global Prompt Refinement with Non-Interfering Attention Masking for One-Shot Federated Learning
by: Qi, Zhuang, et al.
Published: (2025)
by: Qi, Zhuang, et al.
Published: (2025)
Physics Inspired Criterion for Pruning-Quantization Joint Learning
by: Xie, Weiying, et al.
Published: (2023)
by: Xie, Weiying, et al.
Published: (2023)
HAODiff: Human-Aware One-Step Diffusion via Dual-Prompt Guidance
by: Gong, Jue, et al.
Published: (2025)
by: Gong, Jue, et al.
Published: (2025)
A Proximal Algorithm for Network Slimming
by: Bui, Kevin, et al.
Published: (2023)
by: Bui, Kevin, et al.
Published: (2023)
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
Optimizing Human Pose Estimation Through Focused Human and Joint Regions
by: Jiao, Yingying, et al.
Published: (2025)
by: Jiao, Yingying, et al.
Published: (2025)
SmartEraser: Remove Anything from Images using Masked-Region Guidance
by: Jiang, Longtao, et al.
Published: (2025)
by: Jiang, Longtao, et al.
Published: (2025)
ModaVerse: Efficiently Transforming Modalities with LLMs
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Rectified Diffusion Guidance for Conditional Generation
by: Xia, Mengfei, et al.
Published: (2024)
by: Xia, Mengfei, et al.
Published: (2024)
Enhancing Image Aesthetics with Dual-Conditioned Diffusion Models Guided by Multimodal Perception
by: Nan, Xinyu, et al.
Published: (2026)
by: Nan, Xinyu, et al.
Published: (2026)
Joint Perception and Prediction for Autonomous Driving: A Survey
by: Dal'Col, Lucas, et al.
Published: (2024)
by: Dal'Col, Lucas, et al.
Published: (2024)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD Reconstruction
by: Yang, Bingchen, et al.
Published: (2024)
by: Yang, Bingchen, et al.
Published: (2024)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
by: Jiang, Zhongyu, et al.
Published: (2025)
by: Jiang, Zhongyu, et al.
Published: (2025)
Distribution-aware Interactive Attention Network and Large-scale Cloud Recognition Benchmark on FY-4A Satellite Image
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
Prompt-Free Universal Region Proposal Network
by: Tang, Qihong, et al.
Published: (2026)
by: Tang, Qihong, et al.
Published: (2026)
Joint2Human: High-quality 3D Human Generation via Compact Spherical Embedding of 3D Joints
by: Zhang, Muxin, et al.
Published: (2023)
by: Zhang, Muxin, et al.
Published: (2023)
Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
by: Yu, Qiyang, et al.
Published: (2025)
by: Yu, Qiyang, et al.
Published: (2025)
Understanding and Improving Training-free Loss-based Diffusion Guidance
by: Shen, Yifei, et al.
Published: (2024)
by: Shen, Yifei, et al.
Published: (2024)
SARD: Segmentation-Aware Anomaly Synthesis via Region-Constrained Diffusion with Discriminative Mask Guidance
by: Wang, Yanshu, et al.
Published: (2025)
by: Wang, Yanshu, et al.
Published: (2025)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration
by: Li, Chengyuan, et al.
Published: (2025)
by: Li, Chengyuan, et al.
Published: (2025)
Pinwheel-shaped Convolution and Scale-based Dynamic Loss for Infrared Small Target Detection
by: Yang, Jiangnan, et al.
Published: (2024)
by: Yang, Jiangnan, et al.
Published: (2024)
MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration
by: Peng, Jiahui, et al.
Published: (2026)
by: Peng, Jiahui, et al.
Published: (2026)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
Similar Items
-
OpenVidVRD: Open-Vocabulary Video Visual Relation Detection via Prompt-Driven Semantic Space Alignment
by: Liu, Qi, et al.
Published: (2025) -
FreeA: Human-object Interaction Detection using Free Annotation Labels
by: Liu, Qi, et al.
Published: (2024) -
Precision-Enhanced Human-Object Contact Detection via Depth-Aware Perspective Interaction and Object Texture Restoration
by: Wang, Yuxiao, et al.
Published: (2024) -
A Review of Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2024) -
What-Meets-Where: Unified Learning of Action and Contact Localization in Images
by: Wang, Yuxiao, et al.
Published: (2025)