Conditioning GAN Without Training Dataset
Fuente:
arXiv
Saved in:
| Main Author: | Mekonnen, Kidist Amde |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024)
by: Wang, Zhecan, et al.
Published: (2024)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
by: Abbasi, Mehryar, et al.
Published: (2024)
by: Abbasi, Mehryar, et al.
Published: (2024)
Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
by: Karaca, Ali Can, et al.
Published: (2025)
by: Karaca, Ali Can, et al.
Published: (2025)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
ExDDV: A New Dataset for Explainable Deepfake Detection in Video
by: Hondru, Vlad, et al.
Published: (2025)
by: Hondru, Vlad, et al.
Published: (2025)
Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
by: Singer, Assaf, et al.
Published: (2025)
by: Singer, Assaf, et al.
Published: (2025)
ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
by: Yu, Lijun
Published: (2024)
by: Yu, Lijun
Published: (2024)
Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification on the DAIC-WOZ
by: Patapati, Santosh V.
Published: (2024)
by: Patapati, Santosh V.
Published: (2024)
Vlogger: Make Your Dream A Vlog
by: Zhuang, Shaobin, et al.
Published: (2024)
by: Zhuang, Shaobin, et al.
Published: (2024)
Diffusion Model-Based Video Editing: A Survey
by: Sun, Wenhao, et al.
Published: (2024)
by: Sun, Wenhao, et al.
Published: (2024)
Diffusion Models, Image Super-Resolution And Everything: A Survey
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
Continuous Sign Language Recognition System using Deep Learning with MediaPipe Holistic
by: Srivastava, Sharvani, et al.
Published: (2024)
by: Srivastava, Sharvani, et al.
Published: (2024)
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
by: Liao, Yi, et al.
Published: (2024)
by: Liao, Yi, et al.
Published: (2024)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
Evaluating the Impact of Point Cloud Colorization on Semantic Segmentation Accuracy
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
by: Liu, Sheng, et al.
Published: (2024)
by: Liu, Sheng, et al.
Published: (2024)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
by: Sun, Zeyi, et al.
Published: (2024)
by: Sun, Zeyi, et al.
Published: (2024)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
by: Li, Yumeng, et al.
Published: (2024)
by: Li, Yumeng, et al.
Published: (2024)
HOIN: High-Order Implicit Neural Representations
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Flow Generator Matching
by: Huang, Zemin, et al.
Published: (2024)
by: Huang, Zemin, et al.
Published: (2024)
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
by: Shi, Ruixin, et al.
Published: (2024)
by: Shi, Ruixin, et al.
Published: (2024)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
by: Ntrougkas, Mariano V., et al.
Published: (2024)
by: Ntrougkas, Mariano V., et al.
Published: (2024)
A Low-Resolution Image is Worth 1x1 Words: Enabling Fine Image Super-Resolution with Transformers and TaylorShift
by: Nagaraju, Sanath Budakegowdanadoddi, et al.
Published: (2024)
by: Nagaraju, Sanath Budakegowdanadoddi, et al.
Published: (2024)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
by: Zhang, Kangkai, et al.
Published: (2024)
by: Zhang, Kangkai, et al.
Published: (2024)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024)
by: Wu, Xiaoping, et al.
Published: (2024)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
by: Luo, Weijian, et al.
Published: (2024)
by: Luo, Weijian, et al.
Published: (2024)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
ReconBoost: Boosting Can Achieve Modality Reconcilement
by: Hua, Cong, et al.
Published: (2024)
by: Hua, Cong, et al.
Published: (2024)
Multi-layer Learnable Attention Mask for Multimodal Tasks
by: Barrios, Wayner, et al.
Published: (2024)
by: Barrios, Wayner, et al.
Published: (2024)
Similar Items
-
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024) -
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024) -
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024) -
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
by: Abbasi, Mehryar, et al.
Published: (2024) -
Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
by: Karaca, Ali Can, et al.
Published: (2025)