RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yanzuo, Zuo, Ronglai, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
Improving Continuous Sign Language Recognition with Consistency Constraints and Signer Removal
by: Zuo, Ronglai, et al.
Published: (2022)
by: Zuo, Ronglai, et al.
Published: (2022)
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
Towards Online Continuous Sign Language Recognition and Translation
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
by: Yuan, Shihao, et al.
Published: (2025)
by: Yuan, Shihao, et al.
Published: (2025)
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
by: Zhang, Ruicheng, et al.
Published: (2026)
by: Zhang, Ruicheng, et al.
Published: (2026)
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News
by: Niu, Zhe, et al.
Published: (2024)
by: Niu, Zhe, et al.
Published: (2024)
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2025)
by: Ma, Xiaoxiao, et al.
Published: (2025)
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
by: Fang, Xueji, et al.
Published: (2025)
by: Fang, Xueji, et al.
Published: (2025)
DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis
by: Zuo, Yi, et al.
Published: (2026)
by: Zuo, Yi, et al.
Published: (2026)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
by: Liu, Kunhao, et al.
Published: (2025)
by: Liu, Kunhao, et al.
Published: (2025)
Novel View Extrapolation with Video Diffusion Priors
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis
by: Ren, Yuxi, et al.
Published: (2024)
by: Ren, Yuxi, et al.
Published: (2024)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
Not All Frames Deserve Full Computation: Accelerating Autoregressive Video Generation via Selective Computation and Predictive Extrapolation
by: Cui, Hanshuai, et al.
Published: (2026)
by: Cui, Hanshuai, et al.
Published: (2026)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
by: Deng, Boyang, et al.
Published: (2024)
by: Deng, Boyang, et al.
Published: (2024)
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
by: Yan, Feihong, et al.
Published: (2026)
by: Yan, Feihong, et al.
Published: (2026)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
by: Ghosh, Partha, et al.
Published: (2024)
by: Ghosh, Partha, et al.
Published: (2024)
Autoregressive Video Generation without Vector Quantization
by: Deng, Haoge, et al.
Published: (2024)
by: Deng, Haoge, et al.
Published: (2024)
Real-Time Motion-Controllable Autoregressive Video Diffusion
by: Zhao, Kesen, et al.
Published: (2025)
by: Zhao, Kesen, et al.
Published: (2025)
Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
Unleashing Vision-Language Semantics for Deepfake Video Detection
by: Zhu, Jiawen, et al.
Published: (2026)
by: Zhu, Jiawen, et al.
Published: (2026)
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
by: Xiao, Steven, et al.
Published: (2025)
by: Xiao, Steven, et al.
Published: (2025)
RAVEN: Erasing Invisible Watermarks via Novel View Synthesis
by: Shamshad, Fahad, et al.
Published: (2026)
by: Shamshad, Fahad, et al.
Published: (2026)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
Arc2Face: A Foundation Model for ID-Consistent Human Faces
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
by: Guo, Ying, et al.
Published: (2025)
by: Guo, Ying, et al.
Published: (2025)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
by: Meng, Yihao, et al.
Published: (2026)
by: Meng, Yihao, et al.
Published: (2026)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)
by: Lyu, Hengye, et al.
Published: (2026)
A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
by: Long, Do Xuan, et al.
Published: (2026)
by: Long, Do Xuan, et al.
Published: (2026)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Efficient Autoregressive Video Diffusion with Dummy Head
by: Guo, Hang, et al.
Published: (2026)
by: Guo, Hang, et al.
Published: (2026)
MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation
by: Ye, Xinyan, et al.
Published: (2026)
by: Ye, Xinyan, et al.
Published: (2026)
Boosting Object Detection with Zero-Shot Day-Night Domain Adaptation
by: Du, Zhipeng, et al.
Published: (2023)
by: Du, Zhipeng, et al.
Published: (2023)
Similar Items
-
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
by: Zhao, Zengqun, et al.
Published: (2026) -
Improving Continuous Sign Language Recognition with Consistency Constraints and Signer Removal
by: Zuo, Ronglai, et al.
Published: (2022) -
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
by: Zuo, Ronglai, et al.
Published: (2024) -
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026) -
Towards Online Continuous Sign Language Recognition and Translation
by: Zuo, Ronglai, et al.
Published: (2024)