Your Image is My Video: Reshaping the Receptive Field via Image-To-Video Differentiable AutoAugmentation and Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Casarin, Sofia, Ugwu, Cynthia I., Escalera, Sergio, Lanz, Oswald |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
L-SWAG: Layer-Sample Wise Activation with Gradients information for Zero-Shot NAS on Vision Transformers
by: Casarin, Sofia, et al.
Published: (2025)
by: Casarin, Sofia, et al.
Published: (2025)
GRASP-GCN: Graph-Shape Prioritization for Neural Architecture Search under Distribution Shifts
by: Casarin, Sofia, et al.
Published: (2024)
by: Casarin, Sofia, et al.
Published: (2024)
Fractals as Pre-training Datasets for Anomaly Detection and Localization
by: Ugwu, C. I., et al.
Published: (2024)
by: Ugwu, C. I., et al.
Published: (2024)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
HyFusion: Enhanced Reception Field Transformer for Hyperspectral Image Fusion
by: Lee, Chia-Ming, et al.
Published: (2025)
by: Lee, Chia-Ming, et al.
Published: (2025)
Gate-Shift-Pose: Enhancing Action Recognition in Sports with Skeleton Information
by: Bianchi, Edoardo, et al.
Published: (2025)
by: Bianchi, Edoardo, et al.
Published: (2025)
Temporal Differential Fields for 4D Motion Modeling via Image-to-Video Synthesis
by: You, Xin, et al.
Published: (2025)
by: You, Xin, et al.
Published: (2025)
Sparse-Dense Side-Tuner for efficient Video Temporal Grounding
by: Pujol-Perich, David, et al.
Published: (2025)
by: Pujol-Perich, David, et al.
Published: (2025)
ASTRA: An Action Spotting TRAnsformer for Soccer Videos
by: Xarles, Artur, et al.
Published: (2024)
by: Xarles, Artur, et al.
Published: (2024)
Efficient Single Image Super-Resolution with Entropy Attention and Receptive Field Augmentation
by: Zhao, Xiaole, et al.
Published: (2024)
by: Zhao, Xiaole, et al.
Published: (2024)
Enhancing Histopathological Image Classification via Integrated HOG and Deep Features with Robust Noise Performance
by: Ezuma, Ifeanyi, et al.
Published: (2026)
by: Ezuma, Ifeanyi, et al.
Published: (2026)
Animate Your Motion: Turning Still Images into Dynamic Videos
by: Li, Mingxiao, et al.
Published: (2024)
by: Li, Mingxiao, et al.
Published: (2024)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
by: Erregue, Iñaki, et al.
Published: (2026)
by: Erregue, Iñaki, et al.
Published: (2026)
An Efficient Aerial Image Detection with Variable Receptive Fields
by: Wenbin, Liu
Published: (2025)
by: Wenbin, Liu
Published: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
T-DEED: Temporal-Discriminability Enhancer Encoder-Decoder for Precise Event Spotting in Sports Videos
by: Xarles, Artur, et al.
Published: (2024)
by: Xarles, Artur, et al.
Published: (2024)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
by: Liu, Shuming, et al.
Published: (2026)
by: Liu, Shuming, et al.
Published: (2026)
3D Wavelet Convolutions with Extended Receptive Fields for Hyperspectral Image Classification
by: Li, Guandong, et al.
Published: (2025)
by: Li, Guandong, et al.
Published: (2025)
Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization
by: Su, Weijian, et al.
Published: (2026)
by: Su, Weijian, et al.
Published: (2026)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)
by: Mur-Labadia, Lorenzo, et al.
Published: (2025)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
Your Image is Secretly the Last Frame of a Pseudo Video
by: Chen, Wenlong, et al.
Published: (2024)
by: Chen, Wenlong, et al.
Published: (2024)
Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection
by: Jeong, Sungheon, et al.
Published: (2025)
by: Jeong, Sungheon, et al.
Published: (2025)
HST-MRF: Heterogeneous Swin Transformer with Multi-Receptive Field for Medical Image Segmentation
by: Huang, Xiaofei, et al.
Published: (2023)
by: Huang, Xiaofei, et al.
Published: (2023)
Exploring AI-based Anonymization of Industrial Image and Video Data in the Context of Feature Preservation
by: Triess, Sabrina Cynthia, et al.
Published: (2024)
by: Triess, Sabrina Cynthia, et al.
Published: (2024)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformer
by: Dong, Qingji, et al.
Published: (2026)
by: Dong, Qingji, et al.
Published: (2026)
Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image Classification
by: Wu, Xiao, et al.
Published: (2025)
by: Wu, Xiao, et al.
Published: (2025)
AtomoVideo: High Fidelity Image-to-Video Generation
by: Gong, Litong, et al.
Published: (2024)
by: Gong, Litong, et al.
Published: (2024)
TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video Generation
by: Kwon, Gihyun, et al.
Published: (2024)
by: Kwon, Gihyun, et al.
Published: (2024)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
Distill Video Datasets into Images
by: Zhao, Zhenghao, et al.
Published: (2025)
by: Zhao, Zhenghao, et al.
Published: (2025)
Restricted Receptive Fields for Face Verification
by: Ozturk, Kagan, et al.
Published: (2025)
by: Ozturk, Kagan, et al.
Published: (2025)
Wavelet Convolutions for Large Receptive Fields
by: Finder, Shahaf E., et al.
Published: (2024)
by: Finder, Shahaf E., et al.
Published: (2024)
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
SVAD: From Single Image to 3D Avatar via Synthetic Data Generation with Video Diffusion and Data Augmentation
by: Choi, Yonwoo
Published: (2025)
by: Choi, Yonwoo
Published: (2025)
Light-A-Video: Training-free Video Relighting via Progressive Light Fusion
by: Zhou, Yujie, et al.
Published: (2025)
by: Zhou, Yujie, et al.
Published: (2025)
WaterWave: Bridging Underwater Image Enhancement into Video Streams via Wavelet-based Temporal Consistency Field
by: Zhu, Qi, et al.
Published: (2025)
by: Zhu, Qi, et al.
Published: (2025)
Auto-Vocabulary Semantic Segmentation
by: Ülger, Osman, et al.
Published: (2023)
by: Ülger, Osman, et al.
Published: (2023)
Similar Items
-
L-SWAG: Layer-Sample Wise Activation with Gradients information for Zero-Shot NAS on Vision Transformers
by: Casarin, Sofia, et al.
Published: (2025) -
GRASP-GCN: Graph-Shape Prioritization for Neural Architecture Search under Distribution Shifts
by: Casarin, Sofia, et al.
Published: (2024) -
Fractals as Pre-training Datasets for Anomaly Detection and Localization
by: Ugwu, C. I., et al.
Published: (2024) -
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026) -
HyFusion: Enhanced Reception Field Transformer for Hyperspectral Image Fusion
by: Lee, Chia-Ming, et al.
Published: (2025)