AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Dailan, Feng, Guanlin, Ge, Xingtong, Zhang, Yi, Ma, Bingqi, Song, Guanglu, Liu, Yu, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
von: He, Dailan, et al.
Veröffentlicht: (2025)
von: He, Dailan, et al.
Veröffentlicht: (2025)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
von: Ma, Bingqi, et al.
Veröffentlicht: (2026)
von: Ma, Bingqi, et al.
Veröffentlicht: (2026)
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
von: Ge, Xingtong, et al.
Veröffentlicht: (2026)
von: Ge, Xingtong, et al.
Veröffentlicht: (2026)
High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
von: He, Dailan, et al.
Veröffentlicht: (2025)
von: He, Dailan, et al.
Veröffentlicht: (2025)
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
von: Ma, Bingqi, et al.
Veröffentlicht: (2024)
von: Ma, Bingqi, et al.
Veröffentlicht: (2024)
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
von: Zong, Zhuofan, et al.
Veröffentlicht: (2024)
von: Zong, Zhuofan, et al.
Veröffentlicht: (2024)
MoVA: Adapting Mixture of Vision Experts to Multimodal Context
von: Zong, Zhuofan, et al.
Veröffentlicht: (2024)
von: Zong, Zhuofan, et al.
Veröffentlicht: (2024)
Boosting Neural Representations for Videos with a Conditional Decoder
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting
von: Li, Tiantian, et al.
Veröffentlicht: (2025)
von: Li, Tiantian, et al.
Veröffentlicht: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
von: Wu, Yi, et al.
Veröffentlicht: (2025)
von: Wu, Yi, et al.
Veröffentlicht: (2025)
AR4D: Autoregressive 4D Generation from Monocular Videos
von: Zhu, Hanxin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2025)
Task-Aware Encoder Control for Deep Video Compression
von: Ge, Xingtong, et al.
Veröffentlicht: (2024)
von: Ge, Xingtong, et al.
Veröffentlicht: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
von: Ji, Longbin, et al.
Veröffentlicht: (2026)
Group Critical-token Policy Optimization for Autoregressive Image Generation
von: Zhang, Guohui, et al.
Veröffentlicht: (2025)
von: Zhang, Guohui, et al.
Veröffentlicht: (2025)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
von: Zhou, Junkang, et al.
Veröffentlicht: (2026)
von: Zhou, Junkang, et al.
Veröffentlicht: (2026)
AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
von: Wang, Fu-Yun, et al.
Veröffentlicht: (2024)
von: Wang, Fu-Yun, et al.
Veröffentlicht: (2024)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
von: Zou, Kai, et al.
Veröffentlicht: (2026)
von: Zou, Kai, et al.
Veröffentlicht: (2026)
Fast Training-free Perceptual Image Compression
von: Zhu, Ziran, et al.
Veröffentlicht: (2025)
von: Zhu, Ziran, et al.
Veröffentlicht: (2025)
ADT: Tuning Diffusion Models with Adversarial Supervision
von: Shen, Dazhong, et al.
Veröffentlicht: (2025)
von: Shen, Dazhong, et al.
Veröffentlicht: (2025)
LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
von: Zou, Zhen, et al.
Veröffentlicht: (2026)
von: Zou, Zhen, et al.
Veröffentlicht: (2026)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation
von: Wang, Fu-Yun, et al.
Veröffentlicht: (2024)
von: Wang, Fu-Yun, et al.
Veröffentlicht: (2024)
MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
von: Wu, Fangtai, et al.
Veröffentlicht: (2025)
von: Wu, Fangtai, et al.
Veröffentlicht: (2025)
Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior
von: Xu, Tongda, et al.
Veröffentlicht: (2025)
von: Xu, Tongda, et al.
Veröffentlicht: (2025)
MixAR: Mixture Autoregressive Image Generation
von: Hu, Jinyuan, et al.
Veröffentlicht: (2025)
von: Hu, Jinyuan, et al.
Veröffentlicht: (2025)
Temporally Aligned Audio for Video with Autoregression
von: Viertola, Ilpo, et al.
Veröffentlicht: (2024)
von: Viertola, Ilpo, et al.
Veröffentlicht: (2024)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
von: Qi, Jingyuan, et al.
Veröffentlicht: (2025)
von: Qi, Jingyuan, et al.
Veröffentlicht: (2025)
ControlAR: Controllable Image Generation with Autoregressive Models
von: Li, Zongming, et al.
Veröffentlicht: (2024)
von: Li, Zongming, et al.
Veröffentlicht: (2024)
EditAR: Unified Conditional Generation with Autoregressive Models
von: Mu, Jiteng, et al.
Veröffentlicht: (2025)
von: Mu, Jiteng, et al.
Veröffentlicht: (2025)
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
von: Han, Feng, et al.
Veröffentlicht: (2025)
von: Han, Feng, et al.
Veröffentlicht: (2025)
SJD-VP: Speculative Jacobi Decoding with Verification Prediction for Autoregressive Image Generation
von: Shan, Bingqi, et al.
Veröffentlicht: (2026)
von: Shan, Bingqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
von: He, Dailan, et al.
Veröffentlicht: (2025) -
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
von: Ma, Bingqi, et al.
Veröffentlicht: (2026) -
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
von: Ge, Xingtong, et al.
Veröffentlicht: (2026) -
High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
von: He, Dailan, et al.
Veröffentlicht: (2025) -
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
von: Shao, Hao, et al.
Veröffentlicht: (2024)