Endless World: Real-Time 3D-Aware Long Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ke, Mei, Yiqun, Xu, Jiacong, Patel, Vishal M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wild-GS: Real-Time Novel View Synthesis from Unconstrained Photo Collections
von: Xu, Jiacong, et al.
Veröffentlicht: (2024)
von: Xu, Jiacong, et al.
Veröffentlicht: (2024)
FreeViS: Training-free Video Stylization with Inconsistent References
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
Reference-based Controllable Scene Stylization with Gaussian Splatting
von: Mei, Yiqun, et al.
Veröffentlicht: (2024)
von: Mei, Yiqun, et al.
Veröffentlicht: (2024)
InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction
von: Ye, Dingqiang, et al.
Veröffentlicht: (2026)
von: Ye, Dingqiang, et al.
Veröffentlicht: (2026)
Leveraging Thermal Modality to Enhance Reconstruction in Low-Light Conditions
von: Xu, Jiacong, et al.
Veröffentlicht: (2024)
von: Xu, Jiacong, et al.
Veröffentlicht: (2024)
MedCL: Learning Consistent Anatomy Distribution for Scribble-supervised Medical Image Segmentation
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
von: Zhang, Ke, et al.
Veröffentlicht: (2025)
ModelMix: A New Model-Mixup Strategy to Minimize Vicinal Risk across Tasks for Few-scribble based Cardiac Segmentation
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation
von: Chu, Ernie, et al.
Veröffentlicht: (2026)
von: Chu, Ernie, et al.
Veröffentlicht: (2026)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
Helios: Real Real-Time Long Video Generation Model
von: Yuan, Shenghai, et al.
Veröffentlicht: (2026)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2026)
RELIC: Interactive Video World Model with Long-Horizon Memory
von: Hong, Yicong, et al.
Veröffentlicht: (2025)
von: Hong, Yicong, et al.
Veröffentlicht: (2025)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs
von: Zhou, Teng, et al.
Veröffentlicht: (2024)
von: Zhou, Teng, et al.
Veröffentlicht: (2024)
From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution
von: Chen, Zheng, et al.
Veröffentlicht: (2026)
von: Chen, Zheng, et al.
Veröffentlicht: (2026)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
SegFace: Face Segmentation of Long-Tail Classes
von: Narayan, Kartik, et al.
Veröffentlicht: (2024)
von: Narayan, Kartik, et al.
Veröffentlicht: (2024)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
von: Liu, Kunhao, et al.
Veröffentlicht: (2025)
von: Liu, Kunhao, et al.
Veröffentlicht: (2025)
ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
von: Xing, Ke, et al.
Veröffentlicht: (2025)
von: Xing, Ke, et al.
Veröffentlicht: (2025)
LongDPM: Overlap-Aware 4D Reconstruction from Long Monocular Videos
von: Xu, Chenyi, et al.
Veröffentlicht: (2026)
von: Xu, Chenyi, et al.
Veröffentlicht: (2026)
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
von: Huang, Tianyu, et al.
Veröffentlicht: (2025)
von: Huang, Tianyu, et al.
Veröffentlicht: (2025)
Face-to-Face: A Video Dataset for Multi-Person Interaction Modeling
von: Chu, Ernie, et al.
Veröffentlicht: (2026)
von: Chu, Ernie, et al.
Veröffentlicht: (2026)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
von: Rahman, Aimon, et al.
Veröffentlicht: (2024)
von: Rahman, Aimon, et al.
Veröffentlicht: (2024)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
von: Lu, Beijia, et al.
Veröffentlicht: (2025)
von: Lu, Beijia, et al.
Veröffentlicht: (2025)
Dreamguider: Improved Training free Diffusion-based Conditional Generation
von: Nair, Nithin Gopalakrishnan, et al.
Veröffentlicht: (2024)
von: Nair, Nithin Gopalakrishnan, et al.
Veröffentlicht: (2024)
Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image
von: Mei, Yiqun, et al.
Veröffentlicht: (2024)
von: Mei, Yiqun, et al.
Veröffentlicht: (2024)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
3D Surface Reconstruction with Enhanced High-Frequency Details
von: Zhang, Shikun, et al.
Veröffentlicht: (2025)
von: Zhang, Shikun, et al.
Veröffentlicht: (2025)
LongLive: Real-time Interactive Long Video Generation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Return of Frustratingly Easy Unsupervised Video Domain Adaptation
von: Wei, Pengfei, et al.
Veröffentlicht: (2026)
von: Wei, Pengfei, et al.
Veröffentlicht: (2026)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
EAvatar: Expression-Aware Head Avatar Reconstruction with Generative Geometry Priors
von: Zhang, Shikun, et al.
Veröffentlicht: (2025)
von: Zhang, Shikun, et al.
Veröffentlicht: (2025)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
von: Fang, Zhixue, et al.
Veröffentlicht: (2026)
von: Fang, Zhixue, et al.
Veröffentlicht: (2026)
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
von: Hsin-Ying, Lee, et al.
Veröffentlicht: (2026)
von: Hsin-Ying, Lee, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Wild-GS: Real-Time Novel View Synthesis from Unconstrained Photo Collections
von: Xu, Jiacong, et al.
Veröffentlicht: (2024) -
FreeViS: Training-free Video Stylization with Inconsistent References
von: Xu, Jiacong, et al.
Veröffentlicht: (2025) -
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
von: Zhang, Ke, et al.
Veröffentlicht: (2025) -
Reference-based Controllable Scene Stylization with Gaussian Splatting
von: Mei, Yiqun, et al.
Veröffentlicht: (2024) -
InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction
von: Ye, Dingqiang, et al.
Veröffentlicht: (2026)