Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jongseo, Lee, Wooil, Park, Gyeong-Moon, Kim, Seong Tae, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025)
by: Lee, Jongseo, et al.
Published: (2025)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)
by: Lee, Jongseo, et al.
Published: (2024)
Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
by: Lee, Jongseo, et al.
Published: (2026)
by: Lee, Jongseo, et al.
Published: (2026)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
by: Bae, Kyungho, et al.
Published: (2023)
by: Bae, Kyungho, et al.
Published: (2023)
SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition
by: Kim, Ka Young, et al.
Published: (2025)
by: Kim, Ka Young, et al.
Published: (2025)
ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention Disentanglement
by: Lim, Habin, et al.
Published: (2025)
by: Lim, Habin, et al.
Published: (2025)
ESC: Erasing Space Concept for Knowledge Deletion
by: Lee, Tae-Young, et al.
Published: (2025)
by: Lee, Tae-Young, et al.
Published: (2025)
Universal Domain Adaptation for Semantic Segmentation
by: Choe, Seun-An, et al.
Published: (2025)
by: Choe, Seun-An, et al.
Published: (2025)
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
by: Shao, Jie, et al.
Published: (2025)
by: Shao, Jie, et al.
Published: (2025)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
by: Lee, Sanghyeon, et al.
Published: (2026)
by: Lee, Sanghyeon, et al.
Published: (2026)
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
by: Chang, Joohyun, et al.
Published: (2025)
by: Chang, Joohyun, et al.
Published: (2025)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Generative Unlearning for Any Identity
by: Seo, Juwon, et al.
Published: (2024)
by: Seo, Juwon, et al.
Published: (2024)
Multispectral Pedestrian Detection with Sparsely Annotated Label
by: Lee, Chan, et al.
Published: (2025)
by: Lee, Chan, et al.
Published: (2025)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
Open-Set Domain Adaptation for Semantic Segmentation
by: Choe, Seun-An, et al.
Published: (2024)
by: Choe, Seun-An, et al.
Published: (2024)
Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models
by: Lee, Tae-Young, et al.
Published: (2025)
by: Lee, Tae-Young, et al.
Published: (2025)
Online Continuous Generalized Category Discovery
by: Park, Keon-Hee, et al.
Published: (2024)
by: Park, Keon-Hee, et al.
Published: (2024)
Towards Model-Agnostic Dataset Condensation by Heterogeneous Models
by: Moon, Jun-Yeong, et al.
Published: (2024)
by: Moon, Jun-Yeong, et al.
Published: (2024)
Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
by: Lin, Rifen, et al.
Published: (2025)
by: Lin, Rifen, et al.
Published: (2025)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
EPSilon: Efficient Point Sampling for Lightening of Hybrid-based 3D Avatar Generation
by: Moon, Seungjun, et al.
Published: (2025)
by: Moon, Seungjun, et al.
Published: (2025)
Versatile Incremental Learning: Towards Class and Domain-Agnostic Incremental Learning
by: Park, Min-Yeong, et al.
Published: (2024)
by: Park, Min-Yeong, et al.
Published: (2024)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
by: Kothandaraman, Divya, et al.
Published: (2022)
by: Kothandaraman, Divya, et al.
Published: (2022)
ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization
by: Kim, Minseo, et al.
Published: (2026)
by: Kim, Minseo, et al.
Published: (2026)
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
by: Park, Seong Hyeon, et al.
Published: (2025)
by: Park, Seong Hyeon, et al.
Published: (2025)
Exploring Explainability in Video Action Recognition
by: Saha, Avinab, et al.
Published: (2024)
by: Saha, Avinab, et al.
Published: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
by: Choi, June Suk, et al.
Published: (2025)
by: Choi, June Suk, et al.
Published: (2025)
Disentangling Disentangled Representations: Towards Improved Latent Units via Diffusion Models
by: Jun, Youngjun, et al.
Published: (2024)
by: Jun, Youngjun, et al.
Published: (2024)
M2Former: Multi-Scale Patch Selection for Fine-Grained Visual Recognition
by: Moon, Jiyong, et al.
Published: (2023)
by: Moon, Jiyong, et al.
Published: (2023)
HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
by: Woo, Suhan, et al.
Published: (2025)
by: Woo, Suhan, et al.
Published: (2025)
Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
by: Kim, Hyeon Bae, et al.
Published: (2024)
by: Kim, Hyeon Bae, et al.
Published: (2024)
Similar Items
-
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
by: Lee, Jongseo, et al.
Published: (2025) -
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023) -
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
by: Lee, Jongseo, et al.
Published: (2025) -
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
by: Lee, Jongseo, et al.
Published: (2025) -
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
by: Lee, Jongseo, et al.
Published: (2024)