DEVIAS: Learning Disentangled Video Representations of Action and Scene
Fuente:
arXiv
Saved in:
| Main Authors: | Bae, Kyungho, Ahn, Geo, Kim, Youngrae, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026)
by: Han, Jiwook, et al.
Published: (2026)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)
by: Lee, Jongseo, et al.
Published: (2024)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Learning Group Actions In Disentangled Latent Image Representations
by: Swarnali, Farhana Hossain, et al.
Published: (2025)
by: Swarnali, Farhana Hossain, et al.
Published: (2025)
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
by: Kim, Youngrae, et al.
Published: (2026)
by: Kim, Youngrae, et al.
Published: (2026)
Forecasting Future Videos from Novel Views via Disentangled 3D Scene Representation
by: Yarram, Sudhir, et al.
Published: (2024)
by: Yarram, Sudhir, et al.
Published: (2024)
Prompt-guided Disentangled Representation for Action Recognition
by: Wu, Tianci, et al.
Published: (2025)
by: Wu, Tianci, et al.
Published: (2025)
A Review of Image Retrieval Techniques: Data Augmentation and Adversarial Learning Approaches
by: Jinwoo, Kim
Published: (2024)
by: Jinwoo, Kim
Published: (2024)
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
by: Park, Seong Hyeon, et al.
Published: (2025)
by: Park, Seong Hyeon, et al.
Published: (2025)
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
Rascene: High-Fidelity 3D Scene Imaging with mmWave Communication Signals
by: Song, Kunzhe, et al.
Published: (2026)
by: Song, Kunzhe, et al.
Published: (2026)
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
by: Zhang, Boqiang, et al.
Published: (2024)
by: Zhang, Boqiang, et al.
Published: (2024)
Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
by: Lee, Jongseo, et al.
Published: (2026)
by: Lee, Jongseo, et al.
Published: (2026)
Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
by: Lim, Geuntaek, et al.
Published: (2024)
by: Lim, Geuntaek, et al.
Published: (2024)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
by: Kothandaraman, Divya, et al.
Published: (2022)
by: Kothandaraman, Divya, et al.
Published: (2022)
What Happens When: Learning Temporal Orders of Events in Videos
by: Ahn, Daechul, et al.
Published: (2025)
by: Ahn, Daechul, et al.
Published: (2025)
JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling
by: Lee, Seok Hwan, et al.
Published: (2024)
by: Lee, Seok Hwan, et al.
Published: (2024)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
by: Lee, Sanghyeon, et al.
Published: (2026)
by: Lee, Sanghyeon, et al.
Published: (2026)
Learning 3D Scene Analogies with Neural Contextual Scene Maps
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Feature Augmentation based Test-Time Adaptation
by: Cho, Younggeol, et al.
Published: (2024)
by: Cho, Younggeol, et al.
Published: (2024)
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation
by: Ahn, Jinwoo, et al.
Published: (2024)
by: Ahn, Jinwoo, et al.
Published: (2024)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
by: Kim, Jinwoo, et al.
Published: (2023)
by: Kim, Jinwoo, et al.
Published: (2023)
MetaWeather: Few-Shot Weather-Degraded Image Restoration
by: Kim, Youngrae, et al.
Published: (2023)
by: Kim, Youngrae, et al.
Published: (2023)
CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting
by: Desai, Nishq Poorav, et al.
Published: (2026)
by: Desai, Nishq Poorav, et al.
Published: (2026)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Unsupervised Learning of Disentangled Representations from Video
by: Denton, Remi, et al.
Published: (2017)
by: Denton, Remi, et al.
Published: (2017)
Similar Items
-
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
by: Han, Jiwook, et al.
Published: (2026) -
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025) -
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026) -
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
by: Lee, Jongseo, et al.
Published: (2025) -
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)