Infinite Gaze Generation for Videos with Autoregressive Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Jenna, Groth, Colin, Wu, Tong, Torrens, Finley, Sangkloy, Patsorn, Wetzstein, Gordon, Sun, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
von: Kang, Jenna, et al.
Veröffentlicht: (2025)
von: Kang, Jenna, et al.
Veröffentlicht: (2025)
Cost-Aware Routing for Efficient Text-To-Image Generation
von: Li, Qinchan, et al.
Veröffentlicht: (2025)
von: Li, Qinchan, et al.
Veröffentlicht: (2025)
GazeFusion: Saliency-Guided Image Generation
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2024)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
von: Deng, Boyang, et al.
Veröffentlicht: (2024)
von: Deng, Boyang, et al.
Veröffentlicht: (2024)
BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
von: Po, Ryan, et al.
Veröffentlicht: (2025)
von: Po, Ryan, et al.
Veröffentlicht: (2025)
Spectral Progressive Diffusion for Efficient Image and Video Generation
von: Xiao, Howard, et al.
Veröffentlicht: (2026)
von: Xiao, Howard, et al.
Veröffentlicht: (2026)
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
von: Chao, Brian, et al.
Veröffentlicht: (2026)
von: Chao, Brian, et al.
Veröffentlicht: (2026)
Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
von: Xie, Linxi, et al.
Veröffentlicht: (2026)
von: Xie, Linxi, et al.
Veröffentlicht: (2026)
GazeGPT: Augmenting Human Capabilities using Gaze-contingent Contextual AI for Smart Eyewear
von: Konrad, Robert, et al.
Veröffentlicht: (2024)
von: Konrad, Robert, et al.
Veröffentlicht: (2024)
Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
von: Zhang, Lvmin, et al.
Veröffentlicht: (2025)
von: Zhang, Lvmin, et al.
Veröffentlicht: (2025)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
von: Ackermann, Jan, et al.
Veröffentlicht: (2025)
von: Ackermann, Jan, et al.
Veröffentlicht: (2025)
Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
von: Xiao, Steven, et al.
Veröffentlicht: (2025)
von: Xiao, Steven, et al.
Veröffentlicht: (2025)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
von: Zhang, Lvmin, et al.
Veröffentlicht: (2025)
von: Zhang, Lvmin, et al.
Veröffentlicht: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
von: Fang, Ye, et al.
Veröffentlicht: (2025)
von: Fang, Ye, et al.
Veröffentlicht: (2025)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2024)
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2024)
Video World Models with Long-term Spatial Memory
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
Orthogonal Adaptation for Modular Customization of Diffusion Models
von: Po, Ryan, et al.
Veröffentlicht: (2023)
von: Po, Ryan, et al.
Veröffentlicht: (2023)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
von: Zhang, Mengchen, et al.
Veröffentlicht: (2025)
von: Zhang, Mengchen, et al.
Veröffentlicht: (2025)
JoyStreamer-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion
von: Li, Chaochao, et al.
Veröffentlicht: (2025)
von: Li, Chaochao, et al.
Veröffentlicht: (2025)
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
von: Ackermann, Jan, et al.
Veröffentlicht: (2026)
von: Ackermann, Jan, et al.
Veröffentlicht: (2026)
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video Generator
von: Zhu, Liyuan, et al.
Veröffentlicht: (2026)
von: Zhu, Liyuan, et al.
Veröffentlicht: (2026)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
von: He, Hao, et al.
Veröffentlicht: (2024)
von: He, Hao, et al.
Veröffentlicht: (2024)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
von: He, Hao, et al.
Veröffentlicht: (2025)
von: He, Hao, et al.
Veröffentlicht: (2025)
Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)
Real-Time Motion-Controllable Autoregressive Video Diffusion
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
Interspatial Attention for Efficient 4D Human Video Generation
von: Shao, Ruizhi, et al.
Veröffentlicht: (2025)
von: Shao, Ruizhi, et al.
Veröffentlicht: (2025)
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Autoregressive Video Generation without Vector Quantization
von: Deng, Haoge, et al.
Veröffentlicht: (2024)
von: Deng, Haoge, et al.
Veröffentlicht: (2024)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
von: Li, Zongyi, et al.
Veröffentlicht: (2024)
von: Li, Zongyi, et al.
Veröffentlicht: (2024)
Pathwise Test-Time Correction for Autoregressive Long Video Generation
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2026)
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2026)
Efficient Autoregressive Video Diffusion with Dummy Head
von: Guo, Hang, et al.
Veröffentlicht: (2026)
von: Guo, Hang, et al.
Veröffentlicht: (2026)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
von: Cheng, Tianle, et al.
Veröffentlicht: (2025)
von: Cheng, Tianle, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
von: Kang, Jenna, et al.
Veröffentlicht: (2025) -
Cost-Aware Routing for Efficient Text-To-Image Generation
von: Li, Qinchan, et al.
Veröffentlicht: (2025) -
GazeFusion: Saliency-Guided Image Generation
von: Zhang, Yunxiang, et al.
Veröffentlicht: (2024) -
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
von: Deng, Boyang, et al.
Veröffentlicht: (2024) -
BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
von: Po, Ryan, et al.
Veröffentlicht: (2025)