RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiaoping, Hu, Jie, Wei, Xiaoming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023)
by: Chen, Weifeng, et al.
Published: (2023)
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
by: Qin, Jie, et al.
Published: (2025)
by: Qin, Jie, et al.
Published: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
Enhancing multimodal cooperation via sample-level modality valuation
by: Wei, Yake, et al.
Published: (2023)
by: Wei, Yake, et al.
Published: (2023)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)
by: Xing, Zhen, et al.
Published: (2024)
Diffusion Model-Based Video Editing: A Survey
by: Sun, Wenhao, et al.
Published: (2024)
by: Sun, Wenhao, et al.
Published: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Diffusion Models, Image Super-Resolution And Everything: A Survey
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
by: Luo, Weijian, et al.
Published: (2024)
by: Luo, Weijian, et al.
Published: (2024)
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
by: Jin, Xin, et al.
Published: (2026)
by: Jin, Xin, et al.
Published: (2026)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
TbExplain: A Text-based Explanation Method for Scene Classification Models with the Statistical Prediction Correction
by: Aminimehr, Amirhossein, et al.
Published: (2023)
by: Aminimehr, Amirhossein, et al.
Published: (2023)
On-the-fly Modulation for Balanced Multimodal Learning
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
by: Sun, Zeyi, et al.
Published: (2024)
by: Sun, Zeyi, et al.
Published: (2024)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
by: Liu, Sheng, et al.
Published: (2024)
by: Liu, Sheng, et al.
Published: (2024)
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025)
by: Lin, Yuanze, et al.
Published: (2025)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
by: Liu, Huijie, et al.
Published: (2025)
by: Liu, Huijie, et al.
Published: (2025)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
by: Zhang, Zewei, et al.
Published: (2024)
by: Zhang, Zewei, et al.
Published: (2024)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
by: Yu, Lijun, et al.
Published: (2023)
by: Yu, Lijun, et al.
Published: (2023)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
by: Sun, Zeyi, et al.
Published: (2024)
by: Sun, Zeyi, et al.
Published: (2024)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation
by: Shi, Ruixin, et al.
Published: (2024)
by: Shi, Ruixin, et al.
Published: (2024)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
by: Shi, Tiancheng, et al.
Published: (2024)
by: Shi, Tiancheng, et al.
Published: (2024)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
HOIN: High-Order Implicit Neural Representations
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
by: Ghosh, Dhruba, et al.
Published: (2026)
by: Ghosh, Dhruba, et al.
Published: (2026)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
by: Lin, Xiang, et al.
Published: (2025)
by: Lin, Xiang, et al.
Published: (2025)
Latent Space Probing for Adult Content Detection in Video Generative Models
by: Khatri, Alizishaan, et al.
Published: (2026)
by: Khatri, Alizishaan, et al.
Published: (2026)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
by: Jiang, Qinting, et al.
Published: (2024)
by: Jiang, Qinting, et al.
Published: (2024)
Similar Items
-
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023) -
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
by: Qin, Jie, et al.
Published: (2025) -
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
by: Liu, Jiajun, et al.
Published: (2024) -
Enhancing multimodal cooperation via sample-level modality valuation
by: Wei, Yake, et al.
Published: (2023) -
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)