Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jun, Youngjun, Kang, Seil, Han, Woojung, Hwang, Seong Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Real-Time Visual Attribution Streaming in Thinking Model
von: Kang, Seil, et al.
Veröffentlicht: (2026)
von: Kang, Seil, et al.
Veröffentlicht: (2026)
CoBra: Complementary Branch Fusing Class and Semantic Knowledge for Robust Weakly Supervised Semantic Segmentation
von: Han, Woojung, et al.
Veröffentlicht: (2024)
von: Han, Woojung, et al.
Veröffentlicht: (2024)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
Rare Text Semantics Were Always There in Your Diffusion Transformer
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
Disentangling Disentangled Representations: Towards Improved Latent Units via Diffusion Models
von: Jun, Youngjun, et al.
Veröffentlicht: (2024)
von: Jun, Youngjun, et al.
Veröffentlicht: (2024)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models
von: Wang, Zerui, et al.
Veröffentlicht: (2024)
von: Wang, Zerui, et al.
Veröffentlicht: (2024)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
von: Lee, Junsung, et al.
Veröffentlicht: (2025)
von: Lee, Junsung, et al.
Veröffentlicht: (2025)
FALCON: Frequency Adjoint Link with CONtinuous Density Mask for Fast Single Image Dehazing
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2025)
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2025)
Context-based Interpretable Spatio-Temporal Graph Convolutional Network for Human Motion Forecasting
von: Medina, Edgar, et al.
Veröffentlicht: (2024)
von: Medina, Edgar, et al.
Veröffentlicht: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
von: Lee, SuBeen, et al.
Veröffentlicht: (2025)
von: Lee, SuBeen, et al.
Veröffentlicht: (2025)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
von: Wu, Peng, et al.
Veröffentlicht: (2024)
von: Wu, Peng, et al.
Veröffentlicht: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement
von: Xie, Xinyu, et al.
Veröffentlicht: (2025)
von: Xie, Xinyu, et al.
Veröffentlicht: (2025)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
von: Han, Woojung, et al.
Veröffentlicht: (2025)
von: Han, Woojung, et al.
Veröffentlicht: (2025)
Interpreting vision transformers via residual replacement model
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jinyeong, et al.
Veröffentlicht: (2025)
Human-Centric Video Anomaly Detection Through Spatio-Temporal Pose Tokenization and Transformer
von: Noghre, Ghazal Alinezhad, et al.
Veröffentlicht: (2024)
von: Noghre, Ghazal Alinezhad, et al.
Veröffentlicht: (2024)
Boosting Camera Motion Control for Video Diffusion Transformers
von: Cheong, Soon Yau, et al.
Veröffentlicht: (2024)
von: Cheong, Soon Yau, et al.
Veröffentlicht: (2024)
Self-Attentive Spatio-Temporal Calibration for Precise Intermediate Layer Matching in ANN-to-SNN Distillation
von: Hong, Di, et al.
Veröffentlicht: (2025)
von: Hong, Di, et al.
Veröffentlicht: (2025)
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
von: Lee, Uichan, et al.
Veröffentlicht: (2026)
von: Lee, Uichan, et al.
Veröffentlicht: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation
von: Kim, Chanyoung, et al.
Veröffentlicht: (2024)
von: Kim, Chanyoung, et al.
Veröffentlicht: (2024)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
von: Li, Bing, et al.
Veröffentlicht: (2022)
von: Li, Bing, et al.
Veröffentlicht: (2022)
Video Motion Transfer with Diffusion Transformers
von: Pondaven, Alexander, et al.
Veröffentlicht: (2024)
von: Pondaven, Alexander, et al.
Veröffentlicht: (2024)
MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024)
von: Meral, Tuna Han Salih, et al.
Veröffentlicht: (2024)
Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
von: Li, Baiqi, et al.
Veröffentlicht: (2026)
von: Li, Baiqi, et al.
Veröffentlicht: (2026)
MotionShop: Zero-Shot Motion Transfer in Video Diffusion Models with Mixture of Score Guidance
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2024)
DragText: Rethinking Text Embedding in Point-based Image Editing
von: Choi, Gayoon, et al.
Veröffentlicht: (2024)
von: Choi, Gayoon, et al.
Veröffentlicht: (2024)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
von: Dokme, Atahan, et al.
Veröffentlicht: (2026)
von: Dokme, Atahan, et al.
Veröffentlicht: (2026)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
von: Zhang, Runze, et al.
Veröffentlicht: (2025)
von: Zhang, Runze, et al.
Veröffentlicht: (2025)
FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Real-Time Visual Attribution Streaming in Thinking Model
von: Kang, Seil, et al.
Veröffentlicht: (2026) -
CoBra: Complementary Branch Fusing Class and Semantic Knowledge for Robust Weakly Supervised Semantic Segmentation
von: Han, Woojung, et al.
Veröffentlicht: (2024) -
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025) -
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025) -
Rare Text Semantics Were Always There in Your Diffusion Transformer
von: Kang, Seil, et al.
Veröffentlicht: (2025)