Astra: General Interactive World Model with Autoregressive Denoising
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Yixuan, Feng, Jiaqi, Zheng, Wenzhao, Gao, Yuan, Tao, Xin, Wan, Pengfei, Zhou, Jie, Lu, Jiwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
SpectralAR: Spectral Autoregressive Visual Generation
by: Huang, Yuanhui, et al.
Published: (2025)
by: Huang, Yuanhui, et al.
Published: (2025)
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024)
by: Zuo, Sicheng, et al.
Published: (2024)
GlobalMamba: Global Image Serialization for Vision Mamba
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image
by: Zhang, Yanran, et al.
Published: (2025)
by: Zhang, Yanran, et al.
Published: (2025)
$\bf{D^3}$QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
by: Zhang, Yanran, et al.
Published: (2025)
by: Zhang, Yanran, et al.
Published: (2025)
Path Choice Matters for Clear Attribution in Path Methods
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
Preventing Local Pitfalls in Vector Quantization via Optimal Transport
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
Doe-1: Closed-Loop Autonomous Driving with Large World Model
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Latent Diffusion Model without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Measuring 3D Spatial Geometric Consistency in Dynamic Generated Videos
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
GenWorld: Towards Detecting AI-generated Real-world Simulation Videos
by: Chen, Weiliang, et al.
Published: (2025)
by: Chen, Weiliang, et al.
Published: (2025)
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
by: Wu, Yuqi, et al.
Published: (2025)
by: Wu, Yuqi, et al.
Published: (2025)
OGGSplat: Open Gaussian Growing for Generalizable Reconstruction with Expanded Field-of-View
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
V2M: Visual 2-Dimensional Mamba for Image Representation Learning
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
SFTok: Bridging the Performance Gap in Discrete Tokenizers
by: Rao, Qihang, et al.
Published: (2025)
by: Rao, Qihang, et al.
Published: (2025)
GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
Quantize-then-Rectify: Efficient VQ-VAE Training
by: Zhang, Borui, et al.
Published: (2025)
by: Zhang, Borui, et al.
Published: (2025)
Fast Shapley Value Estimation: A Unified Approach
by: Zhang, Borui, et al.
Published: (2023)
by: Zhang, Borui, et al.
Published: (2023)
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
by: Zhang, Yanran, et al.
Published: (2026)
by: Zhang, Yanran, et al.
Published: (2026)
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
by: Wang, Lening, et al.
Published: (2024)
by: Wang, Lening, et al.
Published: (2024)
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Learning Counterfactually Decoupled Attention for Open-World Model Attribution
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
AstraNav-World: World Model for Foresight Control and Consistency
by: Chen, Jintao, et al.
Published: (2025)
by: Chen, Jintao, et al.
Published: (2025)
4Dynamic: Text-to-4D Generation with Hybrid Priors
by: Yuan, Yu-Jie, et al.
Published: (2024)
by: Yuan, Yu-Jie, et al.
Published: (2024)
Moaw: Unleashing Motion Awareness for Video Diffusion Models
by: Zhang, Tianqi, et al.
Published: (2026)
by: Zhang, Tianqi, et al.
Published: (2026)
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
by: Chen, Ming, et al.
Published: (2025)
by: Chen, Ming, et al.
Published: (2025)
GameFactory: Creating New Games with Generative Interactive Videos
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
Streaming 4D Visual Geometry Transformer
by: Zhuo, Dong, et al.
Published: (2025)
by: Zhuo, Dong, et al.
Published: (2025)
Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detection
by: Zeng, Shuai, et al.
Published: (2024)
by: Zeng, Shuai, et al.
Published: (2024)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
FlowIE: Efficient Image Enhancement via Rectified Flow
by: Zhu, Yixuan, et al.
Published: (2024)
by: Zhu, Yixuan, et al.
Published: (2024)
Similar Items
-
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024) -
Terra: Explorable Native 3D World Model with Point Latents
by: Huang, Yuanhui, et al.
Published: (2025) -
SpectralAR: Spectral Autoregressive Visual Generation
by: Huang, Yuanhui, et al.
Published: (2025) -
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024) -
GlobalMamba: Global Image Serialization for Vision Mamba
by: Wang, Chengkun, et al.
Published: (2024)