VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Sixiao, Peng, Zimian, Zhou, Yanpeng, Zhu, Yi, Xu, Hang, Huang, Xiangru, Fu, Yanwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
di: Zheng, Sixiao, et al.
Pubblicazione: (2024)
di: Zheng, Sixiao, et al.
Pubblicazione: (2024)
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
di: Zheng, Sixiao, et al.
Pubblicazione: (2024)
di: Zheng, Sixiao, et al.
Pubblicazione: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
di: Liang, Feng, et al.
Pubblicazione: (2023)
di: Liang, Feng, et al.
Pubblicazione: (2023)
VidCtx: Context-aware Video Question Answering with Image Models
di: Goulas, Andreas, et al.
Pubblicazione: (2024)
di: Goulas, Andreas, et al.
Pubblicazione: (2024)
Feedback-Driven Rate Control for Learned Video Compression
di: Xu, Zhiheng, et al.
Pubblicazione: (2026)
di: Xu, Zhiheng, et al.
Pubblicazione: (2026)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
di: Xiao, Xinyu, et al.
Pubblicazione: (2026)
di: Xiao, Xinyu, et al.
Pubblicazione: (2026)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
di: Huh, Mina, et al.
Pubblicazione: (2026)
di: Huh, Mina, et al.
Pubblicazione: (2026)
Adaptive Offloading and Enhancement for Low-Light Video Analytics on Mobile Devices
di: He, Yuanyi, et al.
Pubblicazione: (2024)
di: He, Yuanyi, et al.
Pubblicazione: (2024)
ConCLVD: Controllable Chinese Landscape Video Generation via Diffusion Model
di: Liu, Dingming, et al.
Pubblicazione: (2024)
di: Liu, Dingming, et al.
Pubblicazione: (2024)
VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
di: Tian, Zeyue, et al.
Pubblicazione: (2024)
di: Tian, Zeyue, et al.
Pubblicazione: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
A Video Steganography for H.265/HEVC Based on Multiple CU Size and Block Structure Distortion
di: Zhang, Xiang, et al.
Pubblicazione: (2026)
di: Zhang, Xiang, et al.
Pubblicazione: (2026)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
di: Li, Fu, et al.
Pubblicazione: (2025)
di: Li, Fu, et al.
Pubblicazione: (2025)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
di: Zhou, Yang-Hao, et al.
Pubblicazione: (2026)
di: Zhou, Yang-Hao, et al.
Pubblicazione: (2026)
Can We Hear from Events? Generating Speech from Event Camera
di: Fang, Jingping, et al.
Pubblicazione: (2026)
di: Fang, Jingping, et al.
Pubblicazione: (2026)
Joint Optimization of Buffer Delay and HARQ for Video Communications
di: Cheng, Baoping, et al.
Pubblicazione: (2024)
di: Cheng, Baoping, et al.
Pubblicazione: (2024)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
di: Tong, Xinyi, et al.
Pubblicazione: (2025)
di: Tong, Xinyi, et al.
Pubblicazione: (2025)
Editing on the Generative Manifold: A Theoretical and Empirical Study of General Diffusion-Based Image Editing Trade-offs
di: Hu, Yi, et al.
Pubblicazione: (2026)
di: Hu, Yi, et al.
Pubblicazione: (2026)
Stage Light is Sequence$^2$: Multi-Light Control via Imitation Learning
di: Zhao, Zijian, et al.
Pubblicazione: (2026)
di: Zhao, Zijian, et al.
Pubblicazione: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
di: Chen, Tong, et al.
Pubblicazione: (2024)
di: Chen, Tong, et al.
Pubblicazione: (2024)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
di: Kavediya, Harsh, et al.
Pubblicazione: (2025)
di: Kavediya, Harsh, et al.
Pubblicazione: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
di: Zhang, Zhongwei, et al.
Pubblicazione: (2025)
di: Zhang, Zhongwei, et al.
Pubblicazione: (2025)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
di: Xie, Tianyidan, et al.
Pubblicazione: (2026)
di: Xie, Tianyidan, et al.
Pubblicazione: (2026)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
di: Cheng, Hongye, et al.
Pubblicazione: (2025)
di: Cheng, Hongye, et al.
Pubblicazione: (2025)
Deep Bi-directional Attention Network for Image Super-Resolution Quality Assessment
di: Li, Yixiao, et al.
Pubblicazione: (2024)
di: Li, Yixiao, et al.
Pubblicazione: (2024)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
di: Deng, Weihui, et al.
Pubblicazione: (2024)
di: Deng, Weihui, et al.
Pubblicazione: (2024)
Resi-VidTok: An Efficient and Decomposed Progressive Tokenization Framework for Ultra-Low-Rate and Lightweight Video Transmission
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
di: Wang, Chengzhi, et al.
Pubblicazione: (2025)
di: Wang, Chengzhi, et al.
Pubblicazione: (2025)
Accelerating Multi-Condition T2I Generation via Adaptive Condition Offloading and Pruning
di: Kong, Yuxin, et al.
Pubblicazione: (2026)
di: Kong, Yuxin, et al.
Pubblicazione: (2026)
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
di: Qin, Yang, et al.
Pubblicazione: (2025)
di: Qin, Yang, et al.
Pubblicazione: (2025)
Music Grounding by Short Video
di: Xin, Zijie, et al.
Pubblicazione: (2024)
di: Xin, Zijie, et al.
Pubblicazione: (2024)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
di: Li, Qingcao, et al.
Pubblicazione: (2026)
di: Li, Qingcao, et al.
Pubblicazione: (2026)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
di: Yang, Quanwei, et al.
Pubblicazione: (2025)
di: Yang, Quanwei, et al.
Pubblicazione: (2025)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
di: Li, Haitian, et al.
Pubblicazione: (2026)
di: Li, Haitian, et al.
Pubblicazione: (2026)
Automatic Camera Trajectory Control with Enhanced Immersion for Virtual Cinematography
di: Wu, Xinyi, et al.
Pubblicazione: (2023)
di: Wu, Xinyi, et al.
Pubblicazione: (2023)
H.265/HEVC Video Steganalysis Based on CU Block Structure Gradients and IPM Mapping
di: Zhang, Xiang, et al.
Pubblicazione: (2026)
di: Zhang, Xiang, et al.
Pubblicazione: (2026)
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
di: Pan, Mianzhi, et al.
Pubblicazione: (2023)
di: Pan, Mianzhi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
di: Zheng, Sixiao, et al.
Pubblicazione: (2024) -
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
di: Zheng, Sixiao, et al.
Pubblicazione: (2024) -
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
di: Qin, Bosheng, et al.
Pubblicazione: (2023) -
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
di: Liang, Feng, et al.
Pubblicazione: (2023) -
VidCtx: Context-aware Video Question Answering with Image Models
di: Goulas, Andreas, et al.
Pubblicazione: (2024)