Bidirectional Progressive Transformer for Interaction Intention Anticipation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zichen, Luo, Hongchen, Zhai, Wei, Cao, Yang, Kang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PEAR: Phrase-Based Hand-Object Interaction Anticipation
von: Zhang, Zichen, et al.
Veröffentlicht: (2024)
von: Zhang, Zichen, et al.
Veröffentlicht: (2024)
Intention-driven Ego-to-Exo Video Generation
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
von: Deng, Huilin, et al.
Veröffentlicht: (2024)
von: Deng, Huilin, et al.
Veröffentlicht: (2024)
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
von: Yang, Yuhang, et al.
Veröffentlicht: (2023)
von: Yang, Yuhang, et al.
Veröffentlicht: (2023)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
von: Shao, Yawen, et al.
Veröffentlicht: (2024)
von: Shao, Yawen, et al.
Veröffentlicht: (2024)
Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
von: Cao, Congqi, et al.
Veröffentlicht: (2025)
von: Cao, Congqi, et al.
Veröffentlicht: (2025)
Visual-Geometric Collaborative Guidance for Affordance Learning
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
Leverage Task Context for Object Affordance Ranking
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
von: Liu, Cuiyu, et al.
Veröffentlicht: (2024)
von: Liu, Cuiyu, et al.
Veröffentlicht: (2024)
Intention Action Anticipation Model with Guide-Feedback Loop Mechanism
von: Ma, Zongnan, et al.
Veröffentlicht: (2024)
von: Ma, Zongnan, et al.
Veröffentlicht: (2024)
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
von: Deng, Huilin, et al.
Veröffentlicht: (2025)
von: Deng, Huilin, et al.
Veröffentlicht: (2025)
Interaction Region Visual Transformer for Egocentric Action Anticipation
von: Roy, Debaditya, et al.
Veröffentlicht: (2022)
von: Roy, Debaditya, et al.
Veröffentlicht: (2022)
Lightweight Vision Transformer with Bidirectional Interaction
von: Fan, Qihang, et al.
Veröffentlicht: (2023)
von: Fan, Qihang, et al.
Veröffentlicht: (2023)
HOI4ABOT: Human-Object Interaction Anticipation for Human Intention Reading Collaborative roBOTs
von: Mascaro, Esteve Valls, et al.
Veröffentlicht: (2023)
von: Mascaro, Esteve Valls, et al.
Veröffentlicht: (2023)
Visual Context Window Extension: A New Perspective for Long Video Understanding
von: Wei, Hongchen, et al.
Veröffentlicht: (2024)
von: Wei, Hongchen, et al.
Veröffentlicht: (2024)
Training-Free Reasoning and Reflection in MLLMs
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion
von: Lin, Xiaotong, et al.
Veröffentlicht: (2025)
von: Lin, Xiaotong, et al.
Veröffentlicht: (2025)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
von: Cho, Hyunjin, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjin, et al.
Veröffentlicht: (2024)
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
von: Sato, Yuji, et al.
Veröffentlicht: (2025)
von: Sato, Yuji, et al.
Veröffentlicht: (2025)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
Sparse Fine-Tuning of Transformers for Generative Tasks
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of Transformer-Based Question Answering Models and RAG-Enhanced Design
von: Zhang, Zichen, et al.
Veröffentlicht: (2025)
von: Zhang, Zichen, et al.
Veröffentlicht: (2025)
MambaPupil: Bidirectional Selective Recurrent model for Event-based Eye tracking
von: Wang, Zhong, et al.
Veröffentlicht: (2024)
von: Wang, Zhong, et al.
Veröffentlicht: (2024)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
von: Han, Guangyi, et al.
Veröffentlicht: (2025)
von: Han, Guangyi, et al.
Veröffentlicht: (2025)
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
PiT: Progressive Diffusion Transformer
von: Wu, Jiafu, et al.
Veröffentlicht: (2025)
von: Wu, Jiafu, et al.
Veröffentlicht: (2025)
SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention Localization
von: Hu, Xinyi, et al.
Veröffentlicht: (2025)
von: Hu, Xinyi, et al.
Veröffentlicht: (2025)
Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
von: Cao, Mang, et al.
Veröffentlicht: (2025)
von: Cao, Mang, et al.
Veröffentlicht: (2025)
ACIT: Attention-Guided Cross-Modal Interaction Transformer for Pedestrian Crossing Intention Prediction
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025)
von: Li, Yuanzhe, et al.
Veröffentlicht: (2025)
Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention
von: Ozdel, Suleyman, et al.
Veröffentlicht: (2024)
von: Ozdel, Suleyman, et al.
Veröffentlicht: (2024)
Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation
von: Labadia, Lorenzo Mur, et al.
Veröffentlicht: (2026)
von: Labadia, Lorenzo Mur, et al.
Veröffentlicht: (2026)
BCTR: Bidirectional Conditioning Transformer for Scene Graph Generation
von: Hao, Peng, et al.
Veröffentlicht: (2024)
von: Hao, Peng, et al.
Veröffentlicht: (2024)
MatE: Material Extraction from Single-Image via Geometric Prior
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
von: Luo, Yuanhao, et al.
Veröffentlicht: (2026)
von: Luo, Yuanhao, et al.
Veröffentlicht: (2026)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2024)
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PEAR: Phrase-Based Hand-Object Interaction Anticipation
von: Zhang, Zichen, et al.
Veröffentlicht: (2024) -
Intention-driven Ego-to-Exo Video Generation
von: Luo, Hongchen, et al.
Veröffentlicht: (2024) -
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
von: Deng, Huilin, et al.
Veröffentlicht: (2024) -
LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
von: Yang, Yuhang, et al.
Veröffentlicht: (2023) -
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
von: Shao, Yawen, et al.
Veröffentlicht: (2024)