Pyramidal Flow Matching for Efficient Video Generative Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Yang, Sun, Zhicheng, Li, Ningyuan, Xu, Kun, Jiang, Hao, Zhuang, Nan, Huang, Quzhe, Song, Yang, Mu, Yadong, Lin, Zhouchen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance
by: Sun, Zhicheng, et al.
Published: (2024)
by: Sun, Zhicheng, et al.
Published: (2024)
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
by: Li, Jinghan, et al.
Published: (2025)
by: Li, Jinghan, et al.
Published: (2025)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
by: Jin, Yang, et al.
Published: (2023)
by: Jin, Yang, et al.
Published: (2023)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Flow Matching Posterior Sampling: A Training-free Conditional Generation for Flow Matching
by: Song, Kaiyu, et al.
Published: (2024)
by: Song, Kaiyu, et al.
Published: (2024)
OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation
by: Lin, Yuchen, et al.
Published: (2025)
by: Lin, Yuchen, et al.
Published: (2025)
Generating Attribute-Aware Human Motions from Textual Prompt
by: Wang, Xinghan, et al.
Published: (2025)
by: Wang, Xinghan, et al.
Published: (2025)
Real-Time Video Generation with Pyramid Attention Broadcast
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
by: Huang, Quzhe, et al.
Published: (2024)
by: Huang, Quzhe, et al.
Published: (2024)
Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
Learning Inverse Laplacian Pyramid for Progressive Depth Completion
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026)
by: Korzhenkov, Denis, et al.
Published: (2026)
Spatiotemporal Pyramid Flow Matching for Climate Emulation
by: Irvin, Jeremy Andrew, et al.
Published: (2025)
by: Irvin, Jeremy Andrew, et al.
Published: (2025)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
by: Lin, Chenguo, et al.
Published: (2024)
by: Lin, Chenguo, et al.
Published: (2024)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
DyCrowd: Towards Dynamic Crowd Reconstruction from a Large-scene Video
by: Wen, Hao, et al.
Published: (2025)
by: Wen, Hao, et al.
Published: (2025)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
by: Wang, Zanyi, et al.
Published: (2025)
by: Wang, Zanyi, et al.
Published: (2025)
UniFlowRestore: A General Video Restoration Framework via Flow Matching and Prompt Guidance
by: Sun, Shuning, et al.
Published: (2025)
by: Sun, Shuning, et al.
Published: (2025)
Beyond Fixed Inference: Quantitative Flow Matching for Adaptive Image Denoising
by: Duan, Jigang, et al.
Published: (2026)
by: Duan, Jigang, et al.
Published: (2026)
Dynamic Pyramid Network for Efficient Multimodal Large Language Model
by: Ai, Hao, et al.
Published: (2025)
by: Ai, Hao, et al.
Published: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
by: Du, Yang, et al.
Published: (2025)
by: Du, Yang, et al.
Published: (2025)
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
by: Xu, Peiran, et al.
Published: (2026)
by: Xu, Peiran, et al.
Published: (2026)
Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View Images
by: Yan, Hongyu, et al.
Published: (2024)
by: Yan, Hongyu, et al.
Published: (2024)
Global Feature Pyramid Network
by: Xiao, Weilin, et al.
Published: (2023)
by: Xiao, Weilin, et al.
Published: (2023)
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
by: You, Yuyang, et al.
Published: (2026)
by: You, Yuyang, et al.
Published: (2026)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
by: Xu, Xuancheng, et al.
Published: (2025)
by: Xu, Xuancheng, et al.
Published: (2025)
CurveFlow: Curvature-Guided Flow Matching for Image Generation
by: Luo, Yan, et al.
Published: (2025)
by: Luo, Yan, et al.
Published: (2025)
SuperFlow: Training Flow Matching Models with RL on the Fly
by: Chen, Kaijie, et al.
Published: (2025)
by: Chen, Kaijie, et al.
Published: (2025)
MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
by: Gao, Zhiting, et al.
Published: (2025)
by: Gao, Zhiting, et al.
Published: (2025)
KAN-FPN-Stem:A KAN-Enhanced Feature Pyramid Stem for Boosting ViT-based Pose Estimation
by: Tang, HaoNan
Published: (2025)
by: Tang, HaoNan
Published: (2025)
DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation
by: Hu, Jie, et al.
Published: (2026)
by: Hu, Jie, et al.
Published: (2026)
Efficient Pyramid Channel Attention Network for Pathological Myopia Recognition
by: Zhang, Xiaoqing, et al.
Published: (2023)
by: Zhang, Xiaoqing, et al.
Published: (2023)
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
by: Li, Jiahua, et al.
Published: (2025)
by: Li, Jiahua, et al.
Published: (2025)
Similar Items
-
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
by: Jin, Yang, et al.
Published: (2024) -
RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance
by: Sun, Zhicheng, et al.
Published: (2024) -
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
by: Li, Jinghan, et al.
Published: (2025) -
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
by: Jin, Yang, et al.
Published: (2023) -
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)