What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Zhenlong, Wang, Yue, Zhang, Dapeng, Cui, Kejin, Chen, Rui, Tang, Jing, Sun, Lei, Yu, Hongwei, Qian, Chengxuan, Chu, Xiangxiang, Li, Shuo, Zhou, Yuyin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
par: Yuan, Zhenlong, et autres
Publié: (2025)
par: Yuan, Zhenlong, et autres
Publié: (2025)
AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
par: Yuan, Zhenlong, et autres
Publié: (2025)
par: Yuan, Zhenlong, et autres
Publié: (2025)
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
par: Tan, Rongbin, et autres
Publié: (2026)
par: Tan, Rongbin, et autres
Publié: (2026)
SHOE: Semantic HOI Open-Vocabulary Evaluation Metric
par: Noack, Maja, et autres
Publié: (2026)
par: Noack, Maja, et autres
Publié: (2026)
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
par: Lei, Ting, et autres
Publié: (2024)
par: Lei, Ting, et autres
Publié: (2024)
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
par: Lei, Ting, et autres
Publié: (2025)
par: Lei, Ting, et autres
Publié: (2025)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
par: Su, Qile, et autres
Publié: (2026)
par: Su, Qile, et autres
Publié: (2026)
SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection
par: Lin, Xin, et autres
Publié: (2025)
par: Lin, Xin, et autres
Publié: (2025)
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
par: Zhang, Dapeng, et autres
Publié: (2025)
par: Zhang, Dapeng, et autres
Publié: (2025)
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
par: Nguyen, Minh Anh, et autres
Publié: (2026)
par: Nguyen, Minh Anh, et autres
Publié: (2026)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
par: Qian, Chengxuan, et autres
Publié: (2025)
par: Qian, Chengxuan, et autres
Publié: (2025)
CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
par: Lei, Qinqian, et autres
Publié: (2025)
par: Lei, Qinqian, et autres
Publié: (2025)
FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
par: Chen, Rui, et autres
Publié: (2025)
par: Chen, Rui, et autres
Publié: (2025)
Tree Search for LLM Agent Reinforcement Learning
par: Ji, Yuxiang, et autres
Publié: (2025)
par: Ji, Yuxiang, et autres
Publié: (2025)
Protocol Agent: What If Agents Could Use Cryptography In Everyday Life?
par: De Rossi, Marco
Publié: (2026)
par: De Rossi, Marco
Publié: (2026)
EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI Detection
par: Lei, Qinqian, et autres
Publié: (2024)
par: Lei, Qinqian, et autres
Publié: (2024)
What Holds Back Open-Vocabulary Segmentation?
par: Šarić, Josip, et autres
Publié: (2025)
par: Šarić, Josip, et autres
Publié: (2025)
Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing
par: Wang, Jiyuan, et autres
Publié: (2026)
par: Wang, Jiyuan, et autres
Publié: (2026)
Adaptive Label Correction for Robust Medical Image Segmentation with Noisy Labels
par: Qian, Chengxuan, et autres
Publié: (2025)
par: Qian, Chengxuan, et autres
Publié: (2025)
DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
par: Yuan, Zhenlong, et autres
Publié: (2025)
par: Yuan, Zhenlong, et autres
Publié: (2025)
OpenMulti: Open-Vocabulary Instance-Level Multi-Agent Distributed Implicit Mapping
par: Dou, Jianyu, et autres
Publié: (2025)
par: Dou, Jianyu, et autres
Publié: (2025)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
par: Cui, Jieming, et autres
Publié: (2024)
par: Cui, Jieming, et autres
Publié: (2024)
EOV-Seg: Efficient Open-Vocabulary Panoptic Segmentation
par: Niu, Hongwei, et autres
Publié: (2024)
par: Niu, Hongwei, et autres
Publié: (2024)
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
par: Hu, BoCheng, et autres
Publié: (2026)
par: Hu, BoCheng, et autres
Publié: (2026)
EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning
par: Chen, Yinda, et autres
Publié: (2025)
par: Chen, Yinda, et autres
Publié: (2025)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
par: Lin, Jianghang, et autres
Publié: (2025)
par: Lin, Jianghang, et autres
Publié: (2025)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
par: Zhang, Haoshen, et autres
Publié: (2026)
par: Zhang, Haoshen, et autres
Publié: (2026)
From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
par: Qu, Xiangyan, et autres
Publié: (2026)
par: Qu, Xiangyan, et autres
Publié: (2026)
Imagine if We Could Start over: Designing a College from Scratch
par: Troyer, Diane
Publié: (2005)
par: Troyer, Diane
Publié: (2005)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
par: Wang, Zijun, et autres
Publié: (2026)
par: Wang, Zijun, et autres
Publié: (2026)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
par: Li, Mingxing, et autres
Publié: (2025)
par: Li, Mingxing, et autres
Publié: (2025)
Open-Vocabulary Federated Learning with Multimodal Prototyping
par: Zeng, Huimin, et autres
Publié: (2024)
par: Zeng, Huimin, et autres
Publié: (2024)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
par: Bai, Sule, et autres
Publié: (2025)
par: Bai, Sule, et autres
Publié: (2025)
What Second‐Best Epistemology Could Be
par: Marc‐Kevin Daoust
Publié: (2024)
par: Marc‐Kevin Daoust
Publié: (2024)
Open-World Reinforcement Learning over Long Short-Term Imagination
par: Li, Jiajian, et autres
Publié: (2024)
par: Li, Jiajian, et autres
Publié: (2024)
Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection
par: Sarma, Sandipan, et autres
Publié: (2025)
par: Sarma, Sandipan, et autres
Publié: (2025)
AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition
par: Lian, Zheng, et autres
Publié: (2026)
par: Lian, Zheng, et autres
Publié: (2026)
AffectGPT-R1: Leveraging Reinforcement Learning for Open-Vocabulary Multimodal Emotion Recognition
par: Lian, Zheng, et autres
Publié: (2025)
par: Lian, Zheng, et autres
Publié: (2025)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
par: Liu, Zexi, et autres
Publié: (2025)
par: Liu, Zexi, et autres
Publié: (2025)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
par: Zhang, Zhenhao, et autres
Publié: (2025)
par: Zhang, Zhenhao, et autres
Publié: (2025)
Documents similaires
-
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
par: Yuan, Zhenlong, et autres
Publié: (2025) -
AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
par: Yuan, Zhenlong, et autres
Publié: (2025) -
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
par: Tan, Rongbin, et autres
Publié: (2026) -
SHOE: Semantic HOI Open-Vocabulary Evaluation Metric
par: Noack, Maja, et autres
Publié: (2026) -
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
par: Lei, Ting, et autres
Publié: (2024)