TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Victor Shea-Jay, Zhuo, Le, Xin, Yi, Wang, Zhaokai, Wang, Fu-Yun, Wang, Yuchi, Zhang, Renrui, Gao, Peng, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision-to-Music Generation: A Survey
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
TSC-PCAC: Voxel Transformer and Sparse Convolution Based Point Cloud Attribute Compression for 3D Broadcasting
von: Guo, Zixi, et al.
Veröffentlicht: (2024)
von: Guo, Zixi, et al.
Veröffentlicht: (2024)
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026)
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2025)
von: Cai, Qi, et al.
Veröffentlicht: (2025)
Block Erasure-Aware Semantic Multimedia Compression via JSCC Autoencoder
von: Esfahanizadeh, Homa, et al.
Veröffentlicht: (2026)
von: Esfahanizadeh, Homa, et al.
Veröffentlicht: (2026)
PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024)
von: Wang, Sen, et al.
Veröffentlicht: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
von: Yan, Mingxuan, et al.
Veröffentlicht: (2024)
von: Yan, Mingxuan, et al.
Veröffentlicht: (2024)
Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
AMD: Autoregressive Motion Diffusion
von: Han, Bo, et al.
Veröffentlicht: (2023)
von: Han, Bo, et al.
Veröffentlicht: (2023)
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
SpikEmo: Enhancing Emotion Recognition With Spiking Temporal Dynamics in Conversations
von: Yu, Xiaomin, et al.
Veröffentlicht: (2024)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2024)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025)
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025)
Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
von: Wei, Weixing, et al.
Veröffentlicht: (2025)
von: Wei, Weixing, et al.
Veröffentlicht: (2025)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Zhu, Xiaofei, et al.
Veröffentlicht: (2024)
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation
von: Wang, Baisen, et al.
Veröffentlicht: (2024)
von: Wang, Baisen, et al.
Veröffentlicht: (2024)
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
von: Wang, Chengzhi, et al.
Veröffentlicht: (2025)
von: Wang, Chengzhi, et al.
Veröffentlicht: (2025)
Interest-Aware Joint Caching, Computing, and Communication Optimization for Mobile VR Delivery in MEC Networks
von: Fu, Baojie, et al.
Veröffentlicht: (2024)
von: Fu, Baojie, et al.
Veröffentlicht: (2024)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance
von: Zhu, Fengbin, et al.
Veröffentlicht: (2025)
von: Zhu, Fengbin, et al.
Veröffentlicht: (2025)
Rethink Web Service Resilience in Space: A Radiation-Aware and Sustainable Transmission Solution
von: Chen, Long, et al.
Veröffentlicht: (2026)
von: Chen, Long, et al.
Veröffentlicht: (2026)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
von: Yang, Quanwei, et al.
Veröffentlicht: (2025)
von: Yang, Quanwei, et al.
Veröffentlicht: (2025)
Hallucination Localization in Video Captioning
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
Period-conscious Time-series Reconstruction under Local Differential Privacy
von: Wang, Yaxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2026)
Harmony-Aware Music-driven Motion Synthesis with Perceptual Constraint on UGC Datasets
von: Wu, Xinyi, et al.
Veröffentlicht: (2025)
von: Wu, Xinyi, et al.
Veröffentlicht: (2025)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
Music Grounding by Short Video
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
von: Xie, Tianyidan, et al.
Veröffentlicht: (2026)
von: Xie, Tianyidan, et al.
Veröffentlicht: (2026)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
von: Joshi, Aastha, et al.
Veröffentlicht: (2026)
von: Joshi, Aastha, et al.
Veröffentlicht: (2026)
StreamOptix: A Cross-layer Adaptive Video Delivery Scheme
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
A Novel FACS-Aligned Anatomical Text Description Paradigm for Fine-Grained Facial Behavior Synthesis
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
Voxel-GS: Quantized Scaffold Gaussian Splatting Compression with Run-Length Coding
von: Fu, Chunyang, et al.
Veröffentlicht: (2025)
von: Fu, Chunyang, et al.
Veröffentlicht: (2025)
Perceptual-oriented Learned Image Compression with Dynamic Kernel
von: Fu, Nianxiang, et al.
Veröffentlicht: (2024)
von: Fu, Nianxiang, et al.
Veröffentlicht: (2024)
Editing on the Generative Manifold: A Theoretical and Empirical Study of General Diffusion-Based Image Editing Trade-offs
von: Hu, Yi, et al.
Veröffentlicht: (2026)
von: Hu, Yi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Vision-to-Music Generation: A Survey
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025) -
TSC-PCAC: Voxel Transformer and Sparse Convolution Based Point Cloud Attribute Compression for 3D Broadcasting
von: Guo, Zixi, et al.
Veröffentlicht: (2024) -
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026) -
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2025) -
Block Erasure-Aware Semantic Multimedia Compression via JSCC Autoencoder
von: Esfahanizadeh, Homa, et al.
Veröffentlicht: (2026)