Generating Illustrated Instructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Menon, Sachit, Misra, Ishan, Girdhar, Rohit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
por: Girdhar, Rohit, et al.
Publicado: (2023)
por: Girdhar, Rohit, et al.
Publicado: (2023)
InstanceDiffusion: Instance-level Control for Image Generation
por: Wang, Xudong, et al.
Publicado: (2024)
por: Wang, Xudong, et al.
Publicado: (2024)
Diffusion Autoencoders are Scalable Image Tokenizers
por: Chen, Yinbo, et al.
Publicado: (2025)
por: Chen, Yinbo, et al.
Publicado: (2025)
LLMs can see and hear without any training
por: Ashutosh, Kumar, et al.
Publicado: (2025)
por: Ashutosh, Kumar, et al.
Publicado: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024)
por: Zhang, Yiyuan, et al.
Publicado: (2024)
Instruction-Driven 3D Facial Expression Generation and Transition
por: Vo, Anh H., et al.
Publicado: (2026)
por: Vo, Anh H., et al.
Publicado: (2026)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
por: Barrios, Wayner, et al.
Publicado: (2025)
por: Barrios, Wayner, et al.
Publicado: (2025)
Flow Generator Matching
por: Huang, Zemin, et al.
Publicado: (2024)
por: Huang, Zemin, et al.
Publicado: (2024)
Improving Visual Representation Alignment Generation with GRPO
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
por: Zhang, Shiyi, et al.
Publicado: (2026)
por: Zhang, Shiyi, et al.
Publicado: (2026)
STIV: Scalable Text and Image Conditioned Video Generation
por: Lin, Zongyu, et al.
Publicado: (2024)
por: Lin, Zongyu, et al.
Publicado: (2024)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
por: Li, Yumeng, et al.
Publicado: (2024)
por: Li, Yumeng, et al.
Publicado: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
por: Tang, Shixiang, et al.
Publicado: (2025)
por: Tang, Shixiang, et al.
Publicado: (2025)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
por: Abbasi, Mehryar, et al.
Publicado: (2024)
por: Abbasi, Mehryar, et al.
Publicado: (2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
Latent Space Probing for Adult Content Detection in Video Generative Models
por: Khatri, Alizishaan, et al.
Publicado: (2026)
por: Khatri, Alizishaan, et al.
Publicado: (2026)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
por: Li, Quanhao, et al.
Publicado: (2025)
por: Li, Quanhao, et al.
Publicado: (2025)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
por: Li, Quanhao, et al.
Publicado: (2026)
por: Li, Quanhao, et al.
Publicado: (2026)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
por: Lin, Yuanze, et al.
Publicado: (2025)
por: Lin, Yuanze, et al.
Publicado: (2025)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
por: Dong, Hao, et al.
Publicado: (2026)
por: Dong, Hao, et al.
Publicado: (2026)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
por: Yu, Lijun
Publicado: (2024)
por: Yu, Lijun
Publicado: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
por: Zheng, Sixiao, et al.
Publicado: (2025)
por: Zheng, Sixiao, et al.
Publicado: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
por: Li, Guangzhao, et al.
Publicado: (2025)
por: Li, Guangzhao, et al.
Publicado: (2025)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
por: Wang, Haoming, et al.
Publicado: (2025)
por: Wang, Haoming, et al.
Publicado: (2025)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
por: Li, Junzhe, et al.
Publicado: (2025)
por: Li, Junzhe, et al.
Publicado: (2025)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
por: Sun, Zeyi, et al.
Publicado: (2024)
por: Sun, Zeyi, et al.
Publicado: (2024)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
por: Xing, Zhen, et al.
Publicado: (2024)
por: Xing, Zhen, et al.
Publicado: (2024)
Video-based Music Generation
por: Sulun, Serkan
Publicado: (2026)
por: Sulun, Serkan
Publicado: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
por: Lin, Zhiqiu, et al.
Publicado: (2024)
por: Lin, Zhiqiu, et al.
Publicado: (2024)
Instant3D: Instant Text-to-3D Generation
por: Li, Ming, et al.
Publicado: (2023)
por: Li, Ming, et al.
Publicado: (2023)
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
por: Panambur, Tejas, et al.
Publicado: (2025)
por: Panambur, Tejas, et al.
Publicado: (2025)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
por: Muqeet, Abdul, et al.
Publicado: (2023)
por: Muqeet, Abdul, et al.
Publicado: (2023)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
por: Shang, Ziqiao, et al.
Publicado: (2023)
por: Shang, Ziqiao, et al.
Publicado: (2023)
Deep ReLU Networks Have Surprisingly Simple Polytopes
por: Fan, Feng-Lei, et al.
Publicado: (2023)
por: Fan, Feng-Lei, et al.
Publicado: (2023)
TbExplain: A Text-based Explanation Method for Scene Classification Models with the Statistical Prediction Correction
por: Aminimehr, Amirhossein, et al.
Publicado: (2023)
por: Aminimehr, Amirhossein, et al.
Publicado: (2023)
Enhancing multimodal cooperation via sample-level modality valuation
por: Wei, Yake, et al.
Publicado: (2023)
por: Wei, Yake, et al.
Publicado: (2023)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
por: Li, Xuewei, et al.
Publicado: (2023)
por: Li, Xuewei, et al.
Publicado: (2023)
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
por: Long, Zijun, et al.
Publicado: (2023)
por: Long, Zijun, et al.
Publicado: (2023)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
por: Yang, Chenglin, et al.
Publicado: (2023)
por: Yang, Chenglin, et al.
Publicado: (2023)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
por: Chen, Weifeng, et al.
Publicado: (2023)
por: Chen, Weifeng, et al.
Publicado: (2023)
Ejemplares similares
-
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
por: Girdhar, Rohit, et al.
Publicado: (2023) -
InstanceDiffusion: Instance-level Control for Image Generation
por: Wang, Xudong, et al.
Publicado: (2024) -
Diffusion Autoencoders are Scalable Image Tokenizers
por: Chen, Yinbo, et al.
Publicado: (2025) -
LLMs can see and hear without any training
por: Ashutosh, Kumar, et al.
Publicado: (2025) -
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024)