Sora Generates Videos with Stunning Geometrical Consistency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xuanyi, Zhou, Daquan, Zhang, Chenxu, Wei, Shaodong, Hou, Qibin, Cheng, Ming-Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
von: Zhang, Xuying, et al.
Veröffentlicht: (2025)
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2024)
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2024)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2026)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2026)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
von: Ouyang, Ziheng, et al.
Veröffentlicht: (2025)
von: Ouyang, Ziheng, et al.
Veröffentlicht: (2025)
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
von: Su, Zihan, et al.
Veröffentlicht: (2025)
von: Su, Zihan, et al.
Veröffentlicht: (2025)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
von: Wan, Yuhao, et al.
Veröffentlicht: (2024)
von: Wan, Yuhao, et al.
Veröffentlicht: (2024)
Open-Sora Plan: Open-Source Large Video Generation Model
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
von: Zhang, Hongyu, et al.
Veröffentlicht: (2026)
von: Zhang, Hongyu, et al.
Veröffentlicht: (2026)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
von: Li, Yunheng, et al.
Veröffentlicht: (2025)
von: Li, Yunheng, et al.
Veröffentlicht: (2025)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
von: Yin, Bowen, et al.
Veröffentlicht: (2023)
von: Yin, Bowen, et al.
Veröffentlicht: (2023)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
von: Zhou, Yupeng, et al.
Veröffentlicht: (2023)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2023)
Open-Sora: Democratizing Efficient Video Production for All
von: Zheng, Zangwei, et al.
Veröffentlicht: (2024)
von: Zheng, Zangwei, et al.
Veröffentlicht: (2024)
Simple Visual Artifact Detection in Sora-Generated Videos
von: Sugiyama, Misora, et al.
Veröffentlicht: (2025)
von: Sugiyama, Misora, et al.
Veröffentlicht: (2025)
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
von: Zhang, Shi-Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Shi-Chen, et al.
Veröffentlicht: (2025)
GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation
von: Wan, Yuhao, et al.
Veröffentlicht: (2025)
von: Wan, Yuhao, et al.
Veröffentlicht: (2025)
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
Zone Evaluation: Revealing Spatial Bias in Object Detection
von: Zheng, Zhaohui, et al.
Veröffentlicht: (2023)
von: Zheng, Zhaohui, et al.
Veröffentlicht: (2023)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
von: Wang, Jiabao, et al.
Veröffentlicht: (2023)
von: Wang, Jiabao, et al.
Veröffentlicht: (2023)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
Referring Camouflaged Object Detection
von: Zhang, Xuying, et al.
Veröffentlicht: (2023)
von: Zhang, Xuying, et al.
Veröffentlicht: (2023)
HumanNet: Scaling Human-centric Video Learning to One Million Hours
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
From Sora What We Can See: A Survey of Text-to-Video Generation
von: Sun, Rui, et al.
Veröffentlicht: (2024)
von: Sun, Rui, et al.
Veröffentlicht: (2024)
Mixture of Style Experts for Diverse Image Stylization
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
von: Zhu, Shihao, et al.
Veröffentlicht: (2026)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
von: Chen, Yuming, et al.
Veröffentlicht: (2023)
von: Chen, Yuming, et al.
Veröffentlicht: (2023)
Traffic Scene Parsing through the TSP6K Dataset
von: Jiang, Peng-Tao, et al.
Veröffentlicht: (2023)
von: Jiang, Peng-Tao, et al.
Veröffentlicht: (2023)
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
von: Yuan, Xinbin, et al.
Veröffentlicht: (2025)
von: Yuan, Xinbin, et al.
Veröffentlicht: (2025)
Measuring 3D Spatial Geometric Consistency in Dynamic Generated Videos
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
Towards Stable 3D Object Detection
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
A Simple Detector with Frame Dynamics is a Strong Tracker
von: Peng, Chenxu, et al.
Veröffentlicht: (2025)
von: Peng, Chenxu, et al.
Veröffentlicht: (2025)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
von: Kim, Bosung, et al.
Veröffentlicht: (2025)
von: Kim, Bosung, et al.
Veröffentlicht: (2025)
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
von: Kim, Bosung, et al.
Veröffentlicht: (2025)
von: Kim, Bosung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2024) -
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
von: Zhang, Xuying, et al.
Veröffentlicht: (2025) -
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2024) -
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2026) -
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)