Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pu, Yujiang, Huang, Zhanbo, Boddeti, Vishnu, Kong, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
Procedural Mistake Detection via Action Effect Modeling
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
by: Huang, Zhanbo, et al.
Published: (2025)
by: Huang, Zhanbo, et al.
Published: (2025)
H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
by: Huang, Zhanbo, et al.
Published: (2026)
by: Huang, Zhanbo, et al.
Published: (2026)
CryptoFace: End-to-End Encrypted Face Recognition
by: Ao, Wei, et al.
Published: (2025)
by: Ao, Wei, et al.
Published: (2025)
SEAL: Semantic Attention Learning for Long Video Representation
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
OASIS Uncovers: High-Quality T2I Models, Same Old Stereotypes
by: Dehdashtian, Sepehr, et al.
Published: (2025)
by: Dehdashtian, Sepehr, et al.
Published: (2025)
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding
by: Cheng, Zixu, et al.
Published: (2024)
by: Cheng, Zixu, et al.
Published: (2024)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
Learning Prompt-Enhanced Context Features for Weakly-Supervised Video Anomaly Detection
by: Pu, Yujiang, et al.
Published: (2023)
by: Pu, Yujiang, et al.
Published: (2023)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
Utility-Fairness Trade-Offs and How to Find Them
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
by: Weng, Zejia, et al.
Published: (2024)
by: Weng, Zejia, et al.
Published: (2024)
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
by: Gao, Mingqi, et al.
Published: (2026)
by: Gao, Mingqi, et al.
Published: (2026)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026)
by: Zhang, Yabo, et al.
Published: (2026)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
by: Zhou, Chao, et al.
Published: (2025)
by: Zhou, Chao, et al.
Published: (2025)
VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models
by: Zheng, Bowen, et al.
Published: (2026)
by: Zheng, Bowen, et al.
Published: (2026)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
by: Cao, Pu, et al.
Published: (2024)
by: Cao, Pu, et al.
Published: (2024)
ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
Versatile Transition Generation with Image-to-Video Diffusion
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
SMOL-MapSeg: Show Me One Label as prompt
by: Yuan, Yunshuang, et al.
Published: (2025)
by: Yuan, Yunshuang, et al.
Published: (2025)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024)
by: Souček, Tomáš, et al.
Published: (2024)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
Enhancing Privacy in Face Analytics Using Fully Homomorphic Encryption
by: Yalavarthi, Bharat, et al.
Published: (2024)
by: Yalavarthi, Bharat, et al.
Published: (2024)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
by: Xie, Jinheng, et al.
Published: (2024)
by: Xie, Jinheng, et al.
Published: (2024)
Show-o2: Improved Native Unified Multimodal Models
by: Xie, Jinheng, et al.
Published: (2025)
by: Xie, Jinheng, et al.
Published: (2025)
Make Me Happier: Evoking Emotions Through Image Diffusion Models
by: Lin, Qing, et al.
Published: (2024)
by: Lin, Qing, et al.
Published: (2024)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
by: Sun, Yang-Tian, et al.
Published: (2025)
by: Sun, Yang-Tian, et al.
Published: (2025)
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image Models
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
by: Shi, Lei, et al.
Published: (2024)
by: Shi, Lei, et al.
Published: (2024)
Anisotropic Diffusion Probabilistic Model for Imbalanced Image Classification
by: Kong, Jingyu, et al.
Published: (2024)
by: Kong, Jingyu, et al.
Published: (2024)
USV: Unified Sparsification for Accelerating Video Diffusion Models
by: Wu, Xinjian, et al.
Published: (2025)
by: Wu, Xinjian, et al.
Published: (2025)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
by: Bhattacharya, Uttaran, et al.
Published: (2022)
by: Bhattacharya, Uttaran, et al.
Published: (2022)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
by: Chen, Houyuan, et al.
Published: (2026)
by: Chen, Houyuan, et al.
Published: (2026)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
Similar Items
-
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024) -
Procedural Mistake Detection via Action Effect Modeling
by: Guo, Wenliang, et al.
Published: (2025) -
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
by: Huang, Zhanbo, et al.
Published: (2025) -
H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
by: Huang, Zhanbo, et al.
Published: (2026) -
CryptoFace: End-to-End Encrypted Face Recognition
by: Ao, Wei, et al.
Published: (2025)