Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
Fuente:
arXiv
Saved in:
| Main Authors: | Taghipour, Ashkan, Ghahremani, Morteza, Bennamoun, Mohammed, Rekavandi, Aref Miri, Li, Zinuo, Laga, Hamid, Boussaid, Farid |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
LatentMove: Towards Complex Human Movement Video Generation
by: Taghipour, Ashkan, et al.
Published: (2025)
by: Taghipour, Ashkan, et al.
Published: (2025)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026)
by: Taghipour, Ashkan, et al.
Published: (2026)
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
by: Taghipour, Ashkan, et al.
Published: (2025)
by: Taghipour, Ashkan, et al.
Published: (2025)
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
by: Jospin, Laurent Valentin, et al.
Published: (2021)
by: Jospin, Laurent Valentin, et al.
Published: (2021)
Towards Adaptive Subspace Detection in Heterogeneous Environment
by: Rekavandi, Aref Miri
Published: (2024)
by: Rekavandi, Aref Miri
Published: (2024)
Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis
by: Nizamani, Awais, et al.
Published: (2025)
by: Nizamani, Awais, et al.
Published: (2025)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
by: Rajapaksha, Uchitha, et al.
Published: (2024)
by: Rajapaksha, Uchitha, et al.
Published: (2024)
A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-shaped Structures
by: Khanam, Tahmina, et al.
Published: (2024)
by: Khanam, Tahmina, et al.
Published: (2024)
A Riemannian Framework for the Elastic Analysis of the Spatiotemporal Variability in the Shape and Structure of Tree-like 4D Objects
by: Khanam, Tahmina, et al.
Published: (2025)
by: Khanam, Tahmina, et al.
Published: (2025)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
3D Brain and Heart Volume Generative Models: A Survey
by: Liu, Yanbin, et al.
Published: (2022)
by: Liu, Yanbin, et al.
Published: (2022)
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
by: Li, Zinuo, et al.
Published: (2025)
by: Li, Zinuo, et al.
Published: (2025)
RS-Reg: Probabilistic and Robust Certified Regression Through Randomized Smoothing
by: Rekavandi, Aref Miri, et al.
Published: (2024)
by: Rekavandi, Aref Miri, et al.
Published: (2024)
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
by: Li, Zinuo, et al.
Published: (2026)
by: Li, Zinuo, et al.
Published: (2026)
Conditional Diffusion for 3D CT Volume Reconstruction from 2D X-rays
by: Rath, Martin, et al.
Published: (2026)
by: Rath, Martin, et al.
Published: (2026)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Hybrid Transformer-Mamba Architecture for Weakly Supervised Volumetric Medical Segmentation
by: Lyu, Yiheng, et al.
Published: (2025)
by: Lyu, Yiheng, et al.
Published: (2025)
Auxiliary Tasks Enhanced Dual-affinity Learning for Weakly Supervised Semantic Segmentation
by: Xu, Lian, et al.
Published: (2024)
by: Xu, Lian, et al.
Published: (2024)
Adversarial Distortion Learning for Medical Image Denoising
by: Ghahremani, Morteza, et al.
Published: (2022)
by: Ghahremani, Morteza, et al.
Published: (2022)
DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
by: Zhu, Jingmin, et al.
Published: (2025)
by: Zhu, Jingmin, et al.
Published: (2025)
Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection
by: Sagar, A S M Sharifuzzaman, et al.
Published: (2026)
by: Sagar, A S M Sharifuzzaman, et al.
Published: (2026)
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
by: Wang, Ning, et al.
Published: (2026)
by: Wang, Ning, et al.
Published: (2026)
No-Clean-Reference Image Super-Resolution: Application to Electron Microscopy
by: Khateri, Mohammad, et al.
Published: (2024)
by: Khateri, Mohammad, et al.
Published: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
by: Zhou, Yupeng, et al.
Published: (2023)
by: Zhou, Yupeng, et al.
Published: (2023)
UIFormer: A Unified Transformer-based Framework for Incremental Few-Shot Object Detection and Instance Segmentation
by: Zhang, Chengyuan, et al.
Published: (2024)
by: Zhang, Chengyuan, et al.
Published: (2024)
A Closer Look at the Explainability of Contrastive Language-Image Pre-training
by: Li, Yi, et al.
Published: (2023)
by: Li, Yi, et al.
Published: (2023)
Finite Element and Computational Fluid Dynamics Analysis of a Biodegradable Implant for Large Femoral Bone Defects
by: Sina Taghipour, et al.
Published: (2026)
by: Sina Taghipour, et al.
Published: (2026)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation
by: Wang, Jiajun, et al.
Published: (2024)
by: Wang, Jiajun, et al.
Published: (2024)
A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification
by: Marks, Markus, et al.
Published: (2024)
by: Marks, Markus, et al.
Published: (2024)
Mamba? Catch The Hype Or Rethink What Really Helps for Image Registration
by: Jian, Bailiang, et al.
Published: (2024)
by: Jian, Bailiang, et al.
Published: (2024)
Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
by: Li, Yitong, et al.
Published: (2025)
by: Li, Yitong, et al.
Published: (2025)
A Closer Look at Edema Area Segmentation in SD-OCT Images Using Adversarial Framework
by: Tao, Yuhui, et al.
Published: (2025)
by: Tao, Yuhui, et al.
Published: (2025)
LMM-Regularized CLIP Embeddings for Image Classification
by: Tzelepi, Maria, et al.
Published: (2024)
by: Tzelepi, Maria, et al.
Published: (2024)
A Closer Look at Claim Decomposition
by: Wanner, Miriam, et al.
Published: (2024)
by: Wanner, Miriam, et al.
Published: (2024)
A Closer Look at the Russell Paradox
by: Sheridan, Flash
Published: (2021)
by: Sheridan, Flash
Published: (2021)
A Closer Look at Constrained Instantons
by: Aoki, Takafumi, et al.
Published: (2026)
by: Aoki, Takafumi, et al.
Published: (2026)
Bookmobiles: A Somewhat Closer Look
by: Healy, Eugene
Published: (1971)
by: Healy, Eugene
Published: (1971)
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
Similar Items
-
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
by: Taghipour, Ashkan, et al.
Published: (2024) -
LatentMove: Towards Complex Human Movement Video Generation
by: Taghipour, Ashkan, et al.
Published: (2025) -
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026) -
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
by: Taghipour, Ashkan, et al.
Published: (2025) -
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
by: Jospin, Laurent Valentin, et al.
Published: (2021)