Object Aware Egocentric Online Action Detection
Fuente:
arXiv
Saved in:
| Main Authors: | An, Joungbin, Park, Yunsu, Kang, Hyolim, Kim, Seon Joo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
by: Kang, Hyolim, et al.
Published: (2024)
by: Kang, Hyolim, et al.
Published: (2024)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
by: Kang, Hyolim, et al.
Published: (2025)
by: Kang, Hyolim, et al.
Published: (2025)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025)
by: Lall, Vishakha, et al.
Published: (2025)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
by: Jung, Ji Hyeok, et al.
Published: (2024)
by: Jung, Ji Hyeok, et al.
Published: (2024)
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
by: Tse, Tze Ho Elden, et al.
Published: (2025)
by: Tse, Tze Ho Elden, et al.
Published: (2025)
Continual Multimodal Egocentric Activity Recognition via Modality-Aware Novel Detection
by: Lim, Wonseon, et al.
Published: (2026)
by: Lim, Wonseon, et al.
Published: (2026)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
Representing 3D Shapes With 64 Latent Vectors for 3D Diffusion Models
by: Cho, In, et al.
Published: (2025)
by: Cho, In, et al.
Published: (2025)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
Efficient Egocentric Action Recognition with Multimodal Data
by: Calzavara, Marco, et al.
Published: (2025)
by: Calzavara, Marco, et al.
Published: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
MALT: Multi-scale Action Learning Transformer for Online Action Detection
by: Yang, Zhipeng, et al.
Published: (2024)
by: Yang, Zhipeng, et al.
Published: (2024)
Investigating Long-term Training for Remote Sensing Object Detection
by: Park, JongHyun, et al.
Published: (2024)
by: Park, JongHyun, et al.
Published: (2024)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
by: Chu, Qiaohui, et al.
Published: (2025)
by: Chu, Qiaohui, et al.
Published: (2025)
GenOL: Generating Diverse Examples for Name-only Online Learning
by: Seo, Minhyuk, et al.
Published: (2024)
by: Seo, Minhyuk, et al.
Published: (2024)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
by: Li, Yuan-Ming, et al.
Published: (2024)
by: Li, Yuan-Ming, et al.
Published: (2024)
LEAP:D -- A Novel Prompt-based Approach for Domain-Generalized Aerial Object Detection
by: Park, Chanyeong, et al.
Published: (2024)
by: Park, Chanyeong, et al.
Published: (2024)
SFUOD: Source-Free Unknown Object Detection
by: Park, Keon-Hee, et al.
Published: (2025)
by: Park, Keon-Hee, et al.
Published: (2025)
Visual Accommodation: Rethinking Image Scale as a Learnable Variable for Object Detection
by: Seo, Daeun, et al.
Published: (2024)
by: Seo, Daeun, et al.
Published: (2024)
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
by: Haneji, Yuto, et al.
Published: (2024)
by: Haneji, Yuto, et al.
Published: (2024)
Towards Streaming LiDAR Object Detection with Point Clouds as Egocentric Sequences
by: Zhang, Mellon M., et al.
Published: (2025)
by: Zhang, Mellon M., et al.
Published: (2025)
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025)
by: Song, Min Geun, et al.
Published: (2025)
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
by: Park, Konyul, et al.
Published: (2024)
by: Park, Konyul, et al.
Published: (2024)
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025)
by: An, Sojung, et al.
Published: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
Background-Aware Defect Generation for Robust Industrial Anomaly Detection
by: Cho, Youngjae, et al.
Published: (2024)
by: Cho, Youngjae, et al.
Published: (2024)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
by: Messina, Nicola, et al.
Published: (2025)
by: Messina, Nicola, et al.
Published: (2025)
Trajectory-Aware Adaptive Inference in Object Detection Models
by: Papanikolaou, Grigorios, et al.
Published: (2026)
by: Papanikolaou, Grigorios, et al.
Published: (2026)
Spatial-Frequency Aware for Object Detection in RAW Image
by: Ye, Zhuohua, et al.
Published: (2025)
by: Ye, Zhuohua, et al.
Published: (2025)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
FC-Track: Overlap-Aware Post-Association Correction for Online Multi-Object Tracking
by: Ju, Cheng, et al.
Published: (2026)
by: Ju, Cheng, et al.
Published: (2026)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
by: Park, Junho, et al.
Published: (2025)
by: Park, Junho, et al.
Published: (2025)
Uncertainty Aware Human-machine Collaboration in Camouflaged Object Detection
by: Yang, Ziyue, et al.
Published: (2025)
by: Yang, Ziyue, et al.
Published: (2025)
Similar Items
-
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
by: Kang, Hyolim, et al.
Published: (2024) -
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024) -
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
by: Kang, Hyolim, et al.
Published: (2025) -
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025) -
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)