AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Yuchao, Fang, Guian, Jiang, Yuxin, Mao, Weijia, Han, Song, Cai, Han, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
Mitty: Diffusion-based Human-to-Robot Video Generation
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Personalized Vision via Visual In-Context Learning
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
by: Gao, Heyuan, et al.
Published: (2026)
by: Gao, Heyuan, et al.
Published: (2026)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
Edit Transfer: Learning Image Editing via Vision In-Context Relations
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
by: Liang, Zhanhao, et al.
Published: (2026)
by: Liang, Zhanhao, et al.
Published: (2026)
FramePrompt: In-context Controllable Animation with Zero Structural Changes
by: Fang, Guian, et al.
Published: (2025)
by: Fang, Guian, et al.
Published: (2025)
Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
by: Mao, Qi, et al.
Published: (2025)
by: Mao, Qi, et al.
Published: (2025)
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
by: Yeo, Juan, et al.
Published: (2025)
by: Yeo, Juan, et al.
Published: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
by: Liu, Yunze, et al.
Published: (2026)
by: Liu, Yunze, et al.
Published: (2026)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
Place Anything into Any Video
by: Liu, Ziling, et al.
Published: (2024)
by: Liu, Ziling, et al.
Published: (2024)
P-Flow: Prompting Visual Effects Generation
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
Stereo Any Video: Temporally Consistent Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)
by: Jing, Junpeng, et al.
Published: (2025)
AVID: Any-Length Video Inpainting with Diffusion Model
by: Zhang, Zhixing, et al.
Published: (2023)
by: Zhang, Zhixing, et al.
Published: (2023)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
by: Zhao, Chengshu, et al.
Published: (2025)
by: Zhao, Chengshu, et al.
Published: (2025)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
FreeTuner: Any Subject in Any Style with Training-free Diffusion
by: Xu, Youcan, et al.
Published: (2024)
by: Xu, Youcan, et al.
Published: (2024)
AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
AnyTop: Character Animation Diffusion with Any Topology
by: Gat, Inbar, et al.
Published: (2025)
by: Gat, Inbar, et al.
Published: (2025)
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
by: Chen, Yutian, et al.
Published: (2026)
by: Chen, Yutian, et al.
Published: (2026)
Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Multi-human Interactive Talking Dataset
by: Zhu, Zeyu, et al.
Published: (2025)
by: Zhu, Zeyu, et al.
Published: (2025)
Automated Movie Generation via Multi-Agent CoT Planning
by: Wu, Weijia, et al.
Published: (2025)
by: Wu, Weijia, et al.
Published: (2025)
Towards A Better Metric for Text-to-Video Generation
by: Wu, Jay Zhangjie, et al.
Published: (2024)
by: Wu, Jay Zhangjie, et al.
Published: (2024)
Dereflection Any Image with Diffusion Priors and Diversified Data
by: Hu, Jichen, et al.
Published: (2025)
by: Hu, Jichen, et al.
Published: (2025)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
by: Wu, Changwei, et al.
Published: (2025)
by: Wu, Changwei, et al.
Published: (2025)
Segment Any Change
by: Zheng, Zhuo, et al.
Published: (2024)
by: Zheng, Zhuo, et al.
Published: (2024)
Similar Items
-
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025) -
Mitty: Diffusion-based Human-to-Robot Video Generation
by: Song, Yiren, et al.
Published: (2025) -
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026) -
Personalized Vision via Visual In-Context Learning
by: Jiang, Yuxin, et al.
Published: (2025) -
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
by: Gao, Heyuan, et al.
Published: (2026)