Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Liu, Song, Yizhi, Huang, Hejun, Liu, Pinxin, Tang, Yunlong, Aliaga, Daniel, Zhou, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
von: Wu, Yuheng, et al.
Veröffentlicht: (2026)
von: Wu, Yuheng, et al.
Veröffentlicht: (2026)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
von: Xu, Zhen, et al.
Veröffentlicht: (2024)
von: Xu, Zhen, et al.
Veröffentlicht: (2024)
Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
von: Gong, Shucheng, et al.
Veröffentlicht: (2025)
von: Gong, Shucheng, et al.
Veröffentlicht: (2025)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
von: Hu, Xiaowei, et al.
Veröffentlicht: (2024)
von: Hu, Xiaowei, et al.
Veröffentlicht: (2024)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
von: Chai, Zenghao, et al.
Veröffentlicht: (2024)
von: Chai, Zenghao, et al.
Veröffentlicht: (2024)
SVGS: Enhancing Gaussian Splatting Using Primitives with Spatially Varying Colors
von: Xu, Rui, et al.
Veröffentlicht: (2024)
von: Xu, Rui, et al.
Veröffentlicht: (2024)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
von: Wang, Yuze, et al.
Veröffentlicht: (2025)
von: Wang, Yuze, et al.
Veröffentlicht: (2025)
MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching
von: Xie, Shuzhao, et al.
Veröffentlicht: (2026)
von: Xie, Shuzhao, et al.
Veröffentlicht: (2026)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
von: Lin, Jiantao, et al.
Veröffentlicht: (2025)
von: Lin, Jiantao, et al.
Veröffentlicht: (2025)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
Neural Network-Based Tracking and 3D Reconstruction of Baseball Pitch Trajectories from Single-View 2D Video
von: Hsieh, Jhen
Veröffentlicht: (2024)
von: Hsieh, Jhen
Veröffentlicht: (2024)
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing
von: Lionar, Stefan, et al.
Veröffentlicht: (2025)
von: Lionar, Stefan, et al.
Veröffentlicht: (2025)
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
von: Chen, Weiliang, et al.
Veröffentlicht: (2024)
von: Chen, Weiliang, et al.
Veröffentlicht: (2024)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
von: Dong, Wenqi, et al.
Veröffentlicht: (2025)
von: Dong, Wenqi, et al.
Veröffentlicht: (2025)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
Laplacian Analysis Meets Dynamics Modelling: Gaussian Splatting for 4D Reconstruction
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
GS-ProCams: Gaussian Splatting-based Projector-Camera Systems
von: Deng, Qingyue, et al.
Veröffentlicht: (2024)
von: Deng, Qingyue, et al.
Veröffentlicht: (2024)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
von: Xu, Yu, et al.
Veröffentlicht: (2024)
von: Xu, Yu, et al.
Veröffentlicht: (2024)
COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Perceptual Visual Quality Assessment: Principles, Methods, and Future Directions
von: Zhou, Wei, et al.
Veröffentlicht: (2025)
von: Zhou, Wei, et al.
Veröffentlicht: (2025)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2025)
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2025)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
von: Flynn, John, et al.
Veröffentlicht: (2026)
von: Flynn, John, et al.
Veröffentlicht: (2026)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025)
von: S, Sridhar, et al.
Veröffentlicht: (2025)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
von: Cha, SeungJu, et al.
Veröffentlicht: (2025)
von: Cha, SeungJu, et al.
Veröffentlicht: (2025)
Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
von: Wang, Zijian, et al.
Veröffentlicht: (2025)
Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2025)
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2025)
altiro3D: Scene representation from single image and novel view synthesis
von: Canessa, E., et al.
Veröffentlicht: (2023)
von: Canessa, E., et al.
Veröffentlicht: (2023)
Exploring Palette based Color Guidance in Diffusion Models
von: Qiu, Qianru, et al.
Veröffentlicht: (2025)
von: Qiu, Qianru, et al.
Veröffentlicht: (2025)
Real-Time Position-Aware View Synthesis from Single-View Input
von: Gond, Manu, et al.
Veröffentlicht: (2024)
von: Gond, Manu, et al.
Veröffentlicht: (2024)
ImagenHub: Standardizing the evaluation of conditional image generation models
von: Ku, Max, et al.
Veröffentlicht: (2023)
von: Ku, Max, et al.
Veröffentlicht: (2023)
SAGE: Semantic-Driven Adaptive Gaussian Splatting in Extended Reality
von: Schiavo, Chiara, et al.
Veröffentlicht: (2025)
von: Schiavo, Chiara, et al.
Veröffentlicht: (2025)
Textured mesh Quality Assessment using Geometry and Color Field Similarity
von: Yang, Kaifa, et al.
Veröffentlicht: (2025)
von: Yang, Kaifa, et al.
Veröffentlicht: (2025)
InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2023)
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2023)
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
von: Gupta, Prerit, et al.
Veröffentlicht: (2025)
von: Gupta, Prerit, et al.
Veröffentlicht: (2025)
DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
Freehand Sketch Generation from Mechanical Components
von: Liao, Zhichao, et al.
Veröffentlicht: (2024)
von: Liao, Zhichao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025) -
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026) -
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024) -
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
von: Wu, Yuheng, et al.
Veröffentlicht: (2026) -
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
von: Xu, Zhen, et al.
Veröffentlicht: (2024)