Uniform Discrete Diffusion with Metric Path for Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Haoge, Pan, Ting, Zhang, Fan, Liu, Yang, Luo, Zhuoyan, Cui, Yufeng, Wang, Wenxuan, Shen, Chunhua, Shan, Shiguang, Zhang, Zhaoxiang, Wang, Xinlong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autoregressive Video Generation without Vector Quantization
by: Deng, Haoge, et al.
Published: (2024)
by: Deng, Haoge, et al.
Published: (2024)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
Tokenize Anything via Prompting
by: Pan, Ting, et al.
Published: (2023)
by: Pan, Ting, et al.
Published: (2023)
Emu3.5: Native Multimodal Models are World Learners
by: Cui, Yufeng, et al.
Published: (2025)
by: Cui, Yufeng, et al.
Published: (2025)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
by: Ma, Baorui, et al.
Published: (2024)
by: Ma, Baorui, et al.
Published: (2024)
CI-VID: A Coherent Interleaved Text-Video Dataset
by: Ju, Yiming, et al.
Published: (2025)
by: Ju, Yiming, et al.
Published: (2025)
Diffusion Feedback Helps CLIP See Better
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models
by: Wang, Wen, et al.
Published: (2023)
by: Wang, Wen, et al.
Published: (2023)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2024)
by: Wang, Zhongqi, et al.
Published: (2024)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
by: Li, Zizun, et al.
Published: (2026)
by: Li, Zizun, et al.
Published: (2026)
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
by: Sun, Quan, et al.
Published: (2024)
by: Sun, Quan, et al.
Published: (2024)
Generalized Face Liveness Detection via De-fake Face Generator
by: Long, Xingming, et al.
Published: (2024)
by: Long, Xingming, et al.
Published: (2024)
T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
by: Li, Changzhen, et al.
Published: (2025)
by: Li, Changzhen, et al.
Published: (2025)
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
by: Gao, Huanlin, et al.
Published: (2025)
by: Gao, Huanlin, et al.
Published: (2025)
Generative Multimodal Models are In-Context Learners
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
Design-based Estimation Theory for Complex Experiments
by: Chang, Haoge
Published: (2023)
by: Chang, Haoge
Published: (2023)
Emu: Generative Pretraining in Multimodality
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
by: Wan, Zhifan, et al.
Published: (2024)
by: Wan, Zhifan, et al.
Published: (2024)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
Optimizing for the Shortest Path in Denoising Diffusion Model
by: Chen, Ping, et al.
Published: (2025)
by: Chen, Ping, et al.
Published: (2025)
Computer-Aided Design Generation by Cascaded Discrete Diffusion Model
by: Pan, Honghu, et al.
Published: (2026)
by: Pan, Honghu, et al.
Published: (2026)
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models
by: Cai, Yufei, et al.
Published: (2025)
by: Cai, Yufei, et al.
Published: (2025)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
Generalized Discrete Diffusion with Self-Correction
by: Wang, Linxuan, et al.
Published: (2026)
by: Wang, Linxuan, et al.
Published: (2026)
CapsFusion: Rethinking Image-Text Data at Scale
by: Yu, Qiying, et al.
Published: (2023)
by: Yu, Qiying, et al.
Published: (2023)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
by: Li, Keliang, et al.
Published: (2026)
by: Li, Keliang, et al.
Published: (2026)
Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
by: Yang, Junqi, et al.
Published: (2026)
by: Yang, Junqi, et al.
Published: (2026)
Confidence Aware Learning for Reliable Face Anti-spoofing
by: Long, Xingming, et al.
Published: (2024)
by: Long, Xingming, et al.
Published: (2024)
Revisiting Backdoor Attacks on Time Series Classification in the Frequency Domain
by: Huang, Yuanmin, et al.
Published: (2025)
by: Huang, Yuanmin, et al.
Published: (2025)
Similar Items
-
Autoregressive Video Generation without Vector Quantization
by: Deng, Haoge, et al.
Published: (2024) -
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026) -
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025) -
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025) -
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
by: Wang, Jiahao, et al.
Published: (2026)