Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts
Fuente:
arXiv
Saved in:
| Main Authors: | Cole, Adam, Grierson, Mick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proceedings of The third international workshop on eXplainable AI for the Arts (XAIxArts)
by: Ford, Corey, et al.
Published: (2025)
by: Ford, Corey, et al.
Published: (2025)
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
by: Cole, Adam, et al.
Published: (2026)
by: Cole, Adam, et al.
Published: (2026)
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
XAIxArts Manifesto: Explainable AI for the Arts
by: Bryan-Kinns, Nick, et al.
Published: (2025)
by: Bryan-Kinns, Nick, et al.
Published: (2025)
Coral Model Generation from Single Images for Virtual Reality Applications
by: Fu, Jie, et al.
Published: (2024)
by: Fu, Jie, et al.
Published: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
by: Cai, Minghong, et al.
Published: (2024)
by: Cai, Minghong, et al.
Published: (2024)
Focus360: Guiding User Attention in Immersive Videos for VR
by: Silva, Paulo Vitor S., et al.
Published: (2026)
by: Silva, Paulo Vitor S., et al.
Published: (2026)
Wireless Video Semantic Communication with Decoupled Diffusion Multi-frame Compensation
by: Xie, Bingyan, et al.
Published: (2025)
by: Xie, Bingyan, et al.
Published: (2025)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
by: Zuo, Heda, et al.
Published: (2025)
by: Zuo, Heda, et al.
Published: (2025)
Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks
by: Zhang, Hailong, et al.
Published: (2025)
by: Zhang, Hailong, et al.
Published: (2025)
BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind
by: Mao, Yuanyuan, et al.
Published: (2024)
by: Mao, Yuanyuan, et al.
Published: (2024)
AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
by: Sajid, M., et al.
Published: (2025)
by: Sajid, M., et al.
Published: (2025)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
by: Wu, Yihang, et al.
Published: (2024)
by: Wu, Yihang, et al.
Published: (2024)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis
by: Feng, Xinyu, et al.
Published: (2024)
by: Feng, Xinyu, et al.
Published: (2024)
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
by: Nareti, Utsav Kumar, et al.
Published: (2024)
by: Nareti, Utsav Kumar, et al.
Published: (2024)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
by: Sun, Jiahui, et al.
Published: (2025)
by: Sun, Jiahui, et al.
Published: (2025)
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
by: Menon, Aditya Srinivas, et al.
Published: (2025)
by: Menon, Aditya Srinivas, et al.
Published: (2025)
Semantic-Guided Unsupervised Video Summarization
by: Liu, Haizhou, et al.
Published: (2026)
by: Liu, Haizhou, et al.
Published: (2026)
Towards Open-Vocabulary Video Semantic Segmentation
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
SFANet: Spatial-Frequency Attention Network for Deepfake Detection
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
MM-HSD: Multi-Modal Hate Speech Detection in Videos
by: Céspedes-Sarrias, Berta, et al.
Published: (2025)
by: Céspedes-Sarrias, Berta, et al.
Published: (2025)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Generative Semantic Communication: Diffusion Models Beyond Bit Recovery
by: Grassucci, Eleonora, et al.
Published: (2023)
by: Grassucci, Eleonora, et al.
Published: (2023)
Harmonizing Attention: Training-free Texture-aware Geometry Transfer
by: Ikuta, Eito, et al.
Published: (2024)
by: Ikuta, Eito, et al.
Published: (2024)
Cross Modification Attention Based Deliberation Model for Image Captioning
by: Lian, Zheng, et al.
Published: (2021)
by: Lian, Zheng, et al.
Published: (2021)
Band-Attention Modulated RetNet for Face Forgery Detection
by: Zhang, Zhida, et al.
Published: (2024)
by: Zhang, Zhida, et al.
Published: (2024)
Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services
by: Jin, Liuyi, et al.
Published: (2025)
by: Jin, Liuyi, et al.
Published: (2025)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
by: Lin, Zixing, et al.
Published: (2026)
by: Lin, Zixing, et al.
Published: (2026)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
by: Tian, Chong, et al.
Published: (2026)
by: Tian, Chong, et al.
Published: (2026)
LL-GABR: Energy Efficient Live Video Streaming Using Reinforcement Learning
by: Raman, Adithya, et al.
Published: (2024)
by: Raman, Adithya, et al.
Published: (2024)
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
by: Wang, Dali, et al.
Published: (2026)
by: Wang, Dali, et al.
Published: (2026)
End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
by: Artioli, Emanuele, et al.
Published: (2025)
by: Artioli, Emanuele, et al.
Published: (2025)
Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming
by: He, Zhiqiang, et al.
Published: (2025)
by: He, Zhiqiang, et al.
Published: (2025)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
by: Dai, Xiangxiang, et al.
Published: (2024)
by: Dai, Xiangxiang, et al.
Published: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline
by: Oh, Minwoo, et al.
Published: (2025)
by: Oh, Minwoo, et al.
Published: (2025)
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
by: Jin, Xin, et al.
Published: (2026)
by: Jin, Xin, et al.
Published: (2026)
Promptus: Can Prompts Streaming Replace Video Streaming with Stable Diffusion
by: Wu, Jiangkai, et al.
Published: (2024)
by: Wu, Jiangkai, et al.
Published: (2024)
Similar Items
-
Proceedings of The third international workshop on eXplainable AI for the Arts (XAIxArts)
by: Ford, Corey, et al.
Published: (2025) -
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
by: Cole, Adam, et al.
Published: (2026) -
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
by: Bryan-Kinns, Nick, et al.
Published: (2024) -
XAIxArts Manifesto: Explainable AI for the Arts
by: Bryan-Kinns, Nick, et al.
Published: (2025) -
Coral Model Generation from Single Images for Virtual Reality Applications
by: Fu, Jie, et al.
Published: (2024)