Flow Generator Matching
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Zemin, Geng, Zhengyang, Luo, Weijian, Qi, Guo-jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
One-Step Diffusion Distillation through Score Implicit Matching
di: Luo, Weijian, et al.
Pubblicazione: (2024)
di: Luo, Weijian, et al.
Pubblicazione: (2024)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
di: Luo, Weijian, et al.
Pubblicazione: (2024)
di: Luo, Weijian, et al.
Pubblicazione: (2024)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2025)
di: Li, Quanhao, et al.
Pubblicazione: (2025)
FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait
di: Ki, Taekyung, et al.
Pubblicazione: (2024)
di: Ki, Taekyung, et al.
Pubblicazione: (2024)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2026)
di: Li, Quanhao, et al.
Pubblicazione: (2026)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
LayerT2V: A Unified Multi-Layer Video Generation Framework
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
di: Lin, Xiang, et al.
Pubblicazione: (2025)
di: Lin, Xiang, et al.
Pubblicazione: (2025)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
di: Liao, Yi, et al.
Pubblicazione: (2024)
di: Liao, Yi, et al.
Pubblicazione: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
di: Li, Junzhe, et al.
Pubblicazione: (2025)
di: Li, Junzhe, et al.
Pubblicazione: (2025)
Generating Illustrated Instructions
di: Menon, Sachit, et al.
Pubblicazione: (2023)
di: Menon, Sachit, et al.
Pubblicazione: (2023)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
di: Luo, Qiuming, et al.
Pubblicazione: (2026)
di: Luo, Qiuming, et al.
Pubblicazione: (2026)
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
di: Shi, Ruixin, et al.
Pubblicazione: (2024)
di: Shi, Ruixin, et al.
Pubblicazione: (2024)
Improving Visual Representation Alignment Generation with GRPO
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
STIV: Scalable Text and Image Conditioned Video Generation
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
di: Li, Yumeng, et al.
Pubblicazione: (2024)
di: Li, Yumeng, et al.
Pubblicazione: (2024)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
di: Abbasi, Mehryar, et al.
Pubblicazione: (2024)
di: Abbasi, Mehryar, et al.
Pubblicazione: (2024)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
di: Li, Xuewei, et al.
Pubblicazione: (2023)
di: Li, Xuewei, et al.
Pubblicazione: (2023)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
Latent Space Probing for Adult Content Detection in Video Generative Models
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
di: Yu, Lijun
Pubblicazione: (2024)
di: Yu, Lijun
Pubblicazione: (2024)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
di: Dong, Hao, et al.
Pubblicazione: (2026)
di: Dong, Hao, et al.
Pubblicazione: (2026)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
di: Wang, Haoming, et al.
Pubblicazione: (2025)
di: Wang, Haoming, et al.
Pubblicazione: (2025)
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
di: Qin, Jie, et al.
Pubblicazione: (2025)
di: Qin, Jie, et al.
Pubblicazione: (2025)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
ReconBoost: Boosting Can Achieve Modality Reconcilement
di: Hua, Cong, et al.
Pubblicazione: (2024)
di: Hua, Cong, et al.
Pubblicazione: (2024)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
di: Huang, Fanding, et al.
Pubblicazione: (2025)
di: Huang, Fanding, et al.
Pubblicazione: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
Deep ReLU Networks Have Surprisingly Simple Polytopes
di: Fan, Feng-Lei, et al.
Pubblicazione: (2023)
di: Fan, Feng-Lei, et al.
Pubblicazione: (2023)
DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation
di: Liu, Ziyuan, et al.
Pubblicazione: (2026)
di: Liu, Ziyuan, et al.
Pubblicazione: (2026)
Video-based Music Generation
di: Sulun, Serkan
Pubblicazione: (2026)
di: Sulun, Serkan
Pubblicazione: (2026)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
PhyWorld: Physics-Faithful World Model for Video Generation
di: Zhao, Pu, et al.
Pubblicazione: (2026)
di: Zhao, Pu, et al.
Pubblicazione: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
di: Lin, Zhiqiu, et al.
Pubblicazione: (2024)
di: Lin, Zhiqiu, et al.
Pubblicazione: (2024)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
Documenti analoghi
-
One-Step Diffusion Distillation through Score Implicit Matching
di: Luo, Weijian, et al.
Pubblicazione: (2024) -
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
di: Luo, Weijian, et al.
Pubblicazione: (2024) -
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2025) -
FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait
di: Ki, Taekyung, et al.
Pubblicazione: (2024) -
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2026)