Advancing Open-source World Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Robbyant Team, Gao, Zelin, Wang, Qiuyu, Zeng, Yanhong, Zhu, Jiapeng, Cheng, Ka Leong, Li, Yixuan, Wang, Hanlin, Xu, Yinghao, Ma, Shuailei, Chen, Yihang, Liu, Jie, Cheng, Yansong, Yao, Yao, Zhu, Jiayi, Meng, Yihao, Zheng, Kecheng, Bai, Qingyan, Chen, Jingye, Shen, Zehong, Yu, Yue, Zhu, Xing, Shen, Yujun, Ouyang, Hao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
por: Wang, Hanlin, et al.
Publicado: (2025)
por: Wang, Hanlin, et al.
Publicado: (2025)
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
por: Bai, Qingyan, et al.
Publicado: (2025)
por: Bai, Qingyan, et al.
Publicado: (2025)
Edicho: Consistent Image Editing in the Wild
por: Bai, Qingyan, et al.
Publicado: (2024)
por: Bai, Qingyan, et al.
Publicado: (2024)
Learning Naturally Aggregated Appearance for Efficient 3D Editing
por: Cheng, Ka Leong, et al.
Publicado: (2023)
por: Cheng, Ka Leong, et al.
Publicado: (2023)
MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
por: Liu, Zichen, et al.
Publicado: (2025)
por: Liu, Zichen, et al.
Publicado: (2025)
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
por: Meng, Yihao, et al.
Publicado: (2026)
por: Meng, Yihao, et al.
Publicado: (2026)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
por: Meng, Yihao, et al.
Publicado: (2025)
por: Meng, Yihao, et al.
Publicado: (2025)
Calligrapher: Freestyle Text Image Customization
por: Ma, Yue, et al.
Publicado: (2025)
por: Ma, Yue, et al.
Publicado: (2025)
Geometric Context Transformer for Streaming 3D Reconstruction
por: Chen, Lin-Zhuo, et al.
Publicado: (2026)
por: Chen, Lin-Zhuo, et al.
Publicado: (2026)
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
por: Lu, Yunhong, et al.
Publicado: (2025)
por: Lu, Yunhong, et al.
Publicado: (2025)
AniDoc: Animation Creation Made Easier
por: Meng, Yihao, et al.
Publicado: (2024)
por: Meng, Yihao, et al.
Publicado: (2024)
LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis
por: Wang, Hanlin, et al.
Publicado: (2024)
por: Wang, Hanlin, et al.
Publicado: (2024)
Learning Visual Generative Priors without Text
por: Ma, Shuailei, et al.
Publicado: (2024)
por: Ma, Shuailei, et al.
Publicado: (2024)
DepthLab: From Partial to Complete
por: Liu, Zhiheng, et al.
Publicado: (2024)
por: Liu, Zhiheng, et al.
Publicado: (2024)
CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
por: Ouyang, Hao, et al.
Publicado: (2023)
por: Ouyang, Hao, et al.
Publicado: (2023)
UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
por: Xia, Mengfei, et al.
Publicado: (2025)
por: Xia, Mengfei, et al.
Publicado: (2025)
InFusion: Inpainting 3D Gaussians via Learning Depth Completion from Diffusion Prior
por: Liu, Zhiheng, et al.
Publicado: (2024)
por: Liu, Zhiheng, et al.
Publicado: (2024)
MagicQuill: An Intelligent Interactive Image Editing System
por: Liu, Zichen, et al.
Publicado: (2024)
por: Liu, Zichen, et al.
Publicado: (2024)
Contextual AD Narration with Interleaved Multimodal Sequence
por: Wang, Hanlin, et al.
Publicado: (2024)
por: Wang, Hanlin, et al.
Publicado: (2024)
DreamLIP: Language-Image Pre-training with Long Captions
por: Zheng, Kecheng, et al.
Publicado: (2024)
por: Zheng, Kecheng, et al.
Publicado: (2024)
Real-time 3D-aware Portrait Editing from a Single Image
por: Bai, Qingyan, et al.
Publicado: (2024)
por: Bai, Qingyan, et al.
Publicado: (2024)
MangaNinja: Line Art Colorization with Precise Reference Following
por: Liu, Zhiheng, et al.
Publicado: (2025)
por: Liu, Zhiheng, et al.
Publicado: (2025)
Framer: Interactive Frame Interpolation
por: Wang, Wen, et al.
Publicado: (2024)
por: Wang, Wen, et al.
Publicado: (2024)
Human Cognition Inspired RAG with Knowledge Graph for Complex Problem Solving
por: Cheng, Yao, et al.
Publicado: (2025)
por: Cheng, Yao, et al.
Publicado: (2025)
Observation and Simulation of Runoff During an Extreme Heatwave in a Glacial Basin on the Central Tibetan Plateau
por: Fei Zhu, et al.
Publicado: (2024)
por: Fei Zhu, et al.
Publicado: (2024)
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
por: Wang, Yunnan, et al.
Publicado: (2026)
por: Wang, Yunnan, et al.
Publicado: (2026)
Embedding Integer Lattices as Ideals into Polynomial Rings
por: Cheng, Yihang, et al.
Publicado: (2023)
por: Cheng, Yihang, et al.
Publicado: (2023)
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
por: Wu, Wei, et al.
Publicado: (2024)
por: Wu, Wei, et al.
Publicado: (2024)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
por: Yang, Jiahao, et al.
Publicado: (2026)
por: Yang, Jiahao, et al.
Publicado: (2026)
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
por: Cui, Jiahao, et al.
Publicado: (2024)
por: Cui, Jiahao, et al.
Publicado: (2024)
Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos
por: Wei, Dingkun, et al.
Publicado: (2026)
por: Wei, Dingkun, et al.
Publicado: (2026)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
por: Lu, Fan, et al.
Publicado: (2024)
por: Lu, Fan, et al.
Publicado: (2024)
ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
por: Li, Ao, et al.
Publicado: (2025)
por: Li, Ao, et al.
Publicado: (2025)
Generative Organizational Behavior Simulation using Large Language Model based Autonomous Agents: A Holacracy Perspective
por: Zhu, Chen, et al.
Publicado: (2024)
por: Zhu, Chen, et al.
Publicado: (2024)
HierPromptLM: A Pure PLM-based Framework for Representation Learning on Heterogeneous Text-rich Networks
por: Zhu, Qiuyu, et al.
Publicado: (2025)
por: Zhu, Qiuyu, et al.
Publicado: (2025)
Masked Depth Modeling for Spatial Perception
por: Tan, Bin, et al.
Publicado: (2026)
por: Tan, Bin, et al.
Publicado: (2026)
Worst-case generation via minimax optimization in Wasserstein space
por: Cheng, Xiuyuan, et al.
Publicado: (2025)
por: Cheng, Xiuyuan, et al.
Publicado: (2025)
Generative models for decision-making under distributional shift
por: Cheng, Xiuyuan, et al.
Publicado: (2026)
por: Cheng, Xiuyuan, et al.
Publicado: (2026)
Out-of-distribution detection based on subspace projection of high-dimensional features output by the last convolutional layer
por: Zhu, Qiuyu, et al.
Publicado: (2024)
por: Zhu, Qiuyu, et al.
Publicado: (2024)
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
por: Chen, Ziyang, et al.
Publicado: (2026)
por: Chen, Ziyang, et al.
Publicado: (2026)
Ejemplares similares
-
The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
por: Wang, Hanlin, et al.
Publicado: (2025) -
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
por: Bai, Qingyan, et al.
Publicado: (2025) -
Edicho: Consistent Image Editing in the Wild
por: Bai, Qingyan, et al.
Publicado: (2024) -
Learning Naturally Aggregated Appearance for Efficient 3D Editing
por: Cheng, Ka Leong, et al.
Publicado: (2023) -
MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
por: Liu, Zichen, et al.
Publicado: (2025)