HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Vo, Khoa, Phan, Thinh, Yamazaki, Kashu, Tran, Minh, Le, Ngan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AISFormer: Amodal Instance Segmentation with Transformer
di: Tran, Minh, et al.
Pubblicazione: (2022)
di: Tran, Minh, et al.
Pubblicazione: (2022)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
di: Vo, Hao, et al.
Pubblicazione: (2026)
di: Vo, Hao, et al.
Pubblicazione: (2026)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
di: Tran, Kim Hoang, et al.
Pubblicazione: (2024)
di: Tran, Kim Hoang, et al.
Pubblicazione: (2024)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
di: Hanyu, Taisei, et al.
Pubblicazione: (2025)
di: Hanyu, Taisei, et al.
Pubblicazione: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
di: Chung, Nhat, et al.
Pubblicazione: (2025)
di: Chung, Nhat, et al.
Pubblicazione: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
di: Le-Phan, Minh-Khoa, et al.
Pubblicazione: (2026)
HAtt-Flow: Hierarchical Attention-Flow Mechanism for Group Activity Scene Graph Generation in Videos
di: Chappa, Naga VS Raviteja, et al.
Pubblicazione: (2023)
di: Chappa, Naga VS Raviteja, et al.
Pubblicazione: (2023)
TrackMe:A Simple and Effective Multiple Object Tracking Annotation Tool
di: Phan, Thinh, et al.
Pubblicazione: (2024)
di: Phan, Thinh, et al.
Pubblicazione: (2024)
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
di: Nguyen, Trong-Thuan, et al.
Pubblicazione: (2025)
di: Nguyen, Trong-Thuan, et al.
Pubblicazione: (2025)
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
Z-GMOT: Zero-shot Generic Multiple Object Tracking
di: Tran, Kim Hoang, et al.
Pubblicazione: (2023)
di: Tran, Kim Hoang, et al.
Pubblicazione: (2023)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
di: Le, Minh Khoa, et al.
Pubblicazione: (2026)
di: Le, Minh Khoa, et al.
Pubblicazione: (2026)
Language-driven Grasp Detection with Mask-guided Attention
di: Van Vo, Tuan, et al.
Pubblicazione: (2024)
di: Van Vo, Tuan, et al.
Pubblicazione: (2024)
Lightweight Language-driven Grasp Detection using Conditional Consistency Model
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia, et al.
Pubblicazione: (2024)
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
di: Tran, Tuyen, et al.
Pubblicazione: (2025)
di: Tran, Tuyen, et al.
Pubblicazione: (2025)
Learning Human Motion with Temporally Conditional Mamba
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Quang, et al.
Pubblicazione: (2025)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
di: Li, Lei-lei, et al.
Pubblicazione: (2025)
di: Li, Lei-lei, et al.
Pubblicazione: (2025)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
di: Le, Quang-Hung, et al.
Pubblicazione: (2024)
di: Le, Quang-Hung, et al.
Pubblicazione: (2024)
DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
di: Tran, Minh, et al.
Pubblicazione: (2025)
di: Tran, Minh, et al.
Pubblicazione: (2025)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
di: Vo, Thanh-Nhan, et al.
Pubblicazione: (2025)
di: Vo, Thanh-Nhan, et al.
Pubblicazione: (2025)
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
di: Tran, Tuyen, et al.
Pubblicazione: (2025)
di: Tran, Tuyen, et al.
Pubblicazione: (2025)
VENUS: Visual Editing with Noise Inversion Using Scene Graphs
di: Vo, Thanh-Nhan, et al.
Pubblicazione: (2026)
di: Vo, Thanh-Nhan, et al.
Pubblicazione: (2026)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
di: Le, Huy, et al.
Pubblicazione: (2025)
di: Le, Huy, et al.
Pubblicazione: (2025)
Detecting Precise Hand Touch Moments in Egocentric Video
di: Nguyen, Huy Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Huy Anh, et al.
Pubblicazione: (2026)
Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
di: Schoonbeek, Tim J., et al.
Pubblicazione: (2025)
di: Schoonbeek, Tim J., et al.
Pubblicazione: (2025)
Instance Tracking in 3D Scenes from Egocentric Videos
di: Zhao, Yunhan, et al.
Pubblicazione: (2023)
di: Zhao, Yunhan, et al.
Pubblicazione: (2023)
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
di: Shibata, Yuto, et al.
Pubblicazione: (2026)
di: Shibata, Yuto, et al.
Pubblicazione: (2026)
SAMJAM: Zero-Shot Video Scene Graph Generation for Egocentric Kitchen Videos
di: Li, Joshua, et al.
Pubblicazione: (2025)
di: Li, Joshua, et al.
Pubblicazione: (2025)
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
di: Vo, Hao, et al.
Pubblicazione: (2026)
di: Vo, Hao, et al.
Pubblicazione: (2026)
Static Scene Reconstruction from Dynamic Egocentric Videos
di: Cui, Qifei, et al.
Pubblicazione: (2026)
di: Cui, Qifei, et al.
Pubblicazione: (2026)
Facial Chick Sexing: An Automated Chick Sexing System From Chick Facial Image
di: Rodriguez, Marta Veganzones, et al.
Pubblicazione: (2024)
di: Rodriguez, Marta Veganzones, et al.
Pubblicazione: (2024)
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
di: Truong, Thanh-Dat, et al.
Pubblicazione: (2023)
di: Truong, Thanh-Dat, et al.
Pubblicazione: (2023)
Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input
di: Hoang, Trung-Hieu, et al.
Pubblicazione: (2023)
di: Hoang, Trung-Hieu, et al.
Pubblicazione: (2023)
CarcassFormer: An End-to-end Transformer-based Framework for Simultaneous Localization, Segmentation and Classification of Poultry Carcass Defect
di: Tran, Minh, et al.
Pubblicazione: (2024)
di: Tran, Minh, et al.
Pubblicazione: (2024)
DINTR: Tracking via Diffusion-based Interpolation
di: Nguyen, Pha, et al.
Pubblicazione: (2024)
di: Nguyen, Pha, et al.
Pubblicazione: (2024)
Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
di: Nguyen, Toan, et al.
Pubblicazione: (2024)
di: Nguyen, Toan, et al.
Pubblicazione: (2024)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
di: Le, Huy, et al.
Pubblicazione: (2025)
di: Le, Huy, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AISFormer: Amodal Instance Segmentation with Transformer
di: Tran, Minh, et al.
Pubblicazione: (2022) -
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
di: Tran, Minh, et al.
Pubblicazione: (2024) -
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
di: Vo, Hao, et al.
Pubblicazione: (2026) -
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
di: Tran, Minh, et al.
Pubblicazione: (2024) -
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
di: Tran, Kim Hoang, et al.
Pubblicazione: (2024)