Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
Fuente:
arXiv
Saved in:
| Main Authors: | NVIDIA, :, Alhaija, Hassan Abu, Alvarez, Jose, Bala, Maciej, Cai, Tiffany, Cao, Tianshi, Cha, Liz, Chen, Joshua, Chen, Mike, Ferroni, Francesco, Fidler, Sanja, Fox, Dieter, Ge, Yunhao, Gu, Jinwei, Hassani, Ali, Isaev, Michael, Jannaty, Pooya, Lan, Shiyi, Lasser, Tobias, Ling, Huan, Liu, Ming-Yu, Liu, Xian, Lu, Yifan, Luo, Alice, Ma, Qianli, Mao, Hanzi, Ramos, Fabio, Ren, Xuanchi, Shen, Tianchang, Sun, Xinglong, Tang, Shitao, Wang, Ting-Chun, Wu, Jay, Xu, Jiashu, Xu, Stella, Xie, Kevin, Ye, Yuchong, Yang, Xiaodong, Zeng, Xiaohui, Zeng, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
by: Ren, Xuanchi, et al.
Published: (2025)
by: Ren, Xuanchi, et al.
Published: (2025)
XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
by: Ren, Xuanchi, et al.
Published: (2023)
by: Ren, Xuanchi, et al.
Published: (2023)
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026)
by: Liu, Shaowei, et al.
Published: (2026)
Cosmos World Foundation Model Platform for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
APE: Agentic Prompt Enhancer for Image Generation and Editing
by: Huang, Zijian, et al.
Published: (2026)
by: Huang, Zijian, et al.
Published: (2026)
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
by: Ren, Xuanchi, et al.
Published: (2025)
by: Ren, Xuanchi, et al.
Published: (2025)
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
by: Lu, Yifan, et al.
Published: (2024)
by: Lu, Yifan, et al.
Published: (2024)
LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis
by: Xie, Kevin, et al.
Published: (2024)
by: Xie, Kevin, et al.
Published: (2024)
Lyra 2.0: Explorable Generative 3D Worlds
by: Shen, Tianchang, et al.
Published: (2026)
by: Shen, Tianchang, et al.
Published: (2026)
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
by: Lu, Yifan, et al.
Published: (2026)
by: Lu, Yifan, et al.
Published: (2026)
SpaceMesh: A Continuous Representation for Learning Manifold Surface Meshes
by: Shen, Tianchang, et al.
Published: (2024)
by: Shen, Tianchang, et al.
Published: (2024)
Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
by: Bahmani, Sherwin, et al.
Published: (2025)
by: Bahmani, Sherwin, et al.
Published: (2025)
ViR: Towards Efficient Vision Retention Backbones
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
by: Sabour, Amirmojtaba, et al.
Published: (2025)
by: Sabour, Amirmojtaba, et al.
Published: (2025)
Align Your Steps: Optimizing Sampling Schedules in Diffusion Models
by: Sabour, Amirmojtaba, et al.
Published: (2024)
by: Sabour, Amirmojtaba, et al.
Published: (2024)
On Data Engineering for Scaling LLM Terminal Capabilities
by: Pi, Renjie, et al.
Published: (2026)
by: Pi, Renjie, et al.
Published: (2026)
SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
by: Ren, Xuanchi, et al.
Published: (2024)
by: Ren, Xuanchi, et al.
Published: (2024)
Trajeglish: Traffic Modeling as Next-Token Prediction
by: Philion, Jonah, et al.
Published: (2023)
by: Philion, Jonah, et al.
Published: (2023)
EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
by: Zeng, Wenxuan, et al.
Published: (2024)
by: Zeng, Wenxuan, et al.
Published: (2024)
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
by: Wang, Zhengyi, et al.
Published: (2024)
by: Wang, Zhengyi, et al.
Published: (2024)
FAN-Unet: Enhancing Unet with vision Fourier Analysis Block for Biomedical Image Segmentation
by: Xu, Jiashu
Published: (2024)
by: Xu, Jiashu
Published: (2024)
Med-TTT: Vision Test-Time Training model for Medical Image Segmentation
by: Xu, Jiashu
Published: (2024)
by: Xu, Jiashu
Published: (2024)
HC-Mamba: Vision MAMBA with Hybrid Convolutional Techniques for Medical Image Segmentation
by: Xu, Jiashu
Published: (2024)
by: Xu, Jiashu
Published: (2024)
ViPE: Video Pose Engine for 3D Geometric Perception
by: Huang, Jiahui, et al.
Published: (2025)
by: Huang, Jiahui, et al.
Published: (2025)
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
ReMatching Dynamic Reconstruction Flow
by: Oblak, Sara, et al.
Published: (2024)
by: Oblak, Sara, et al.
Published: (2024)
NeRF-XL: Scaling NeRFs with Multiple GPUs
by: Li, Ruilong, et al.
Published: (2024)
by: Li, Ruilong, et al.
Published: (2024)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
by: NVIDIA, et al.
Published: (2024)
by: NVIDIA, et al.
Published: (2024)
A Local Gaussian Process Regression Approach to Frequency Response Function Estimation
by: Fang, Xiaozhu, et al.
Published: (2024)
by: Fang, Xiaozhu, et al.
Published: (2024)
On Kernel Design for Regularized Volterra Series Identification of Wiener-Hammerstein Systems
by: Xu, Yu, et al.
Published: (2025)
by: Xu, Yu, et al.
Published: (2025)
Die Ordnung des Theaters. Eine Soziologie der Regie
by: Hänzi, Denis
Published: (2015)
by: Hänzi, Denis
Published: (2015)
CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
by: Xu, Haoxuan, et al.
Published: (2026)
by: Xu, Haoxuan, et al.
Published: (2026)
RadarGen: Automotive Radar Point Cloud Generation from Cameras
by: Borreda, Tomer, et al.
Published: (2025)
by: Borreda, Tomer, et al.
Published: (2025)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
by: Liao, Yuan-Hong, et al.
Published: (2024)
by: Liao, Yuan-Hong, et al.
Published: (2024)
EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models
by: Namekata, Koichi, et al.
Published: (2024)
by: Namekata, Koichi, et al.
Published: (2024)
SuperPADL: Scaling Language-Directed Physics-Based Control with Progressive Supervised Distillation
by: Juravsky, Jordan, et al.
Published: (2024)
by: Juravsky, Jordan, et al.
Published: (2024)
Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?
by: Liao, Yuan-Hong, et al.
Published: (2024)
by: Liao, Yuan-Hong, et al.
Published: (2024)
Similar Items
-
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026) -
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
by: Ren, Xuanchi, et al.
Published: (2025) -
XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
by: Ren, Xuanchi, et al.
Published: (2023) -
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026) -
Cosmos World Foundation Model Platform for Physical AI
by: NVIDIA, et al.
Published: (2025)