ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Hyolim, Hyun, Jeongseok, An, Joungbin, Yu, Youngjae, Kim, Seon Joo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
by: Kang, Hyolim, et al.
Published: (2025)
by: Kang, Hyolim, et al.
Published: (2025)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025)
by: Han, Su Ho, et al.
Published: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
by: An, Joungbin, et al.
Published: (2025)
by: An, Joungbin, et al.
Published: (2025)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
by: Lee, Jinsung, et al.
Published: (2024)
by: Lee, Jinsung, et al.
Published: (2024)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition
by: Li, Xunsong, et al.
Published: (2024)
by: Li, Xunsong, et al.
Published: (2024)
VISAGE: Video Instance Segmentation with Appearance-Guided Enhancement
by: Kim, Hanjung, et al.
Published: (2023)
by: Kim, Hanjung, et al.
Published: (2023)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
by: Shi, Yiran, et al.
Published: (2026)
by: Shi, Yiran, et al.
Published: (2026)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Dual-Stream Alignment for Action Segmentation
by: Gammulle, Harshala, et al.
Published: (2025)
by: Gammulle, Harshala, et al.
Published: (2025)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
by: Jeon, Subin, et al.
Published: (2024)
by: Jeon, Subin, et al.
Published: (2024)
Attentive Illumination Decomposition Model for Multi-Illuminant White Balancing
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
by: Lang, Xiaolei, et al.
Published: (2026)
by: Lang, Xiaolei, et al.
Published: (2026)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
Action100M: A Large-scale Video Action Dataset
by: Chen, Delong, et al.
Published: (2026)
by: Chen, Delong, et al.
Published: (2026)
A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
by: Rahaman, Md. Afzalur, et al.
Published: (2026)
Learning to Enhance Aperture Phasor Field for Non-Line-of-Sight Imaging
by: Cho, In, et al.
Published: (2024)
by: Cho, In, et al.
Published: (2024)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
MMAD: Multi-label Micro-Action Detection in Videos
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Semi-supervised Active Learning for Video Action Detection
by: Singh, Ayush, et al.
Published: (2023)
by: Singh, Ayush, et al.
Published: (2023)
DANCE: Density-agnostic and Class-aware Network for Point Cloud Completion
by: Kim, Da-Yeong, et al.
Published: (2025)
by: Kim, Da-Yeong, et al.
Published: (2025)
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
by: Son, Taein, et al.
Published: (2024)
by: Son, Taein, et al.
Published: (2024)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
Action Emergence from Streaming Intent
by: Jing, Pengfei, et al.
Published: (2026)
by: Jing, Pengfei, et al.
Published: (2026)
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
by: Xu, Haochuan, et al.
Published: (2025)
by: Xu, Haochuan, et al.
Published: (2025)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
by: Nazarenus, Eric, et al.
Published: (2026)
by: Nazarenus, Eric, et al.
Published: (2026)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
by: Bae, Kyungho, et al.
Published: (2023)
by: Bae, Kyungho, et al.
Published: (2023)
Stable Mean Teacher for Semi-supervised Video Action Detection
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Similar Items
-
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024) -
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024) -
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
by: Kang, Hyolim, et al.
Published: (2025) -
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025) -
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025)