Raccoon: Multi-stage Diffusion Training with Coarse-to-Fine Curating Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Zhiyu, Wang, Junyan, Yang, Hao, Qin, Luozheng, Chen, Hesen, Zhou, Qiang, Li, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SARA: Structural and Adversarial Representation Alignment for Training-efficient Diffusion Models
von: Chen, Hesen, et al.
Veröffentlicht: (2025)
von: Chen, Hesen, et al.
Veröffentlicht: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
Omni-Video: Democratizing Unified Video Understanding and Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
von: Yang, Mengping, et al.
Veröffentlicht: (2026)
von: Yang, Mengping, et al.
Veröffentlicht: (2026)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
E2ED^2:Direct Mapping from Noise to Data for Enhanced Diffusion Models
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
von: Qin, Luozheng, et al.
Veröffentlicht: (2026)
von: Qin, Luozheng, et al.
Veröffentlicht: (2026)
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
Unraveling MMDiT Blocks: Training-free Analysis and Enhancement of Text-conditioned Diffusion
von: Li, Binglei, et al.
Veröffentlicht: (2026)
von: Li, Binglei, et al.
Veröffentlicht: (2026)
Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
von: Wang, Zhao, et al.
Veröffentlicht: (2025)
von: Wang, Zhao, et al.
Veröffentlicht: (2025)
Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation
von: Wang, Junyan, et al.
Veröffentlicht: (2024)
von: Wang, Junyan, et al.
Veröffentlicht: (2024)
ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack
von: Gao, Ziyi, et al.
Veröffentlicht: (2024)
von: Gao, Ziyi, et al.
Veröffentlicht: (2024)
Active Coarse-to-Fine Segmentation of Moveable Parts from Real Images
von: Wang, Ruiqi, et al.
Veröffentlicht: (2023)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2023)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
von: Zhang, Angang, et al.
Veröffentlicht: (2025)
von: Zhang, Angang, et al.
Veröffentlicht: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation Maps
von: Wu, Hang, et al.
Veröffentlicht: (2024)
von: Wu, Hang, et al.
Veröffentlicht: (2024)
Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
von: Maduabuchi, Chika, et al.
Veröffentlicht: (2025)
von: Maduabuchi, Chika, et al.
Veröffentlicht: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
Bidirectional Sparse Attention for Faster Video Diffusion Training
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
CoCoDiff: Diversifying Skeleton Action Features via Coarse-Fine Text-Co-Guided Latent Diffusion
von: Zhao, Zhifu, et al.
Veröffentlicht: (2025)
von: Zhao, Zhifu, et al.
Veröffentlicht: (2025)
Distilling Multi-view Diffusion Models into 3D Generators
von: Qin, Hao, et al.
Veröffentlicht: (2025)
von: Qin, Hao, et al.
Veröffentlicht: (2025)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
Coarse-to-Fine Hierarchical Alignment for UAV-based Human Detection using Diffusion Models
von: Li, Wenda, et al.
Veröffentlicht: (2025)
von: Li, Wenda, et al.
Veröffentlicht: (2025)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
Coarse-to-Fine Structure-Aware Artistic Style Transfer
von: Liu, Kunxiao, et al.
Veröffentlicht: (2025)
von: Liu, Kunxiao, et al.
Veröffentlicht: (2025)
CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
von: Guo, Yaowei, et al.
Veröffentlicht: (2025)
von: Guo, Yaowei, et al.
Veröffentlicht: (2025)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
Motion Control for Enhanced Complex Action Video Generation
von: Zhou, Qiang, et al.
Veröffentlicht: (2024)
von: Zhou, Qiang, et al.
Veröffentlicht: (2024)
Coarse-To-Fine Tensor Trains for Compact Visual Representations
von: Loeschcke, Sebastian, et al.
Veröffentlicht: (2024)
von: Loeschcke, Sebastian, et al.
Veröffentlicht: (2024)
Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning
von: Yang, Shuonan, et al.
Veröffentlicht: (2026)
von: Yang, Shuonan, et al.
Veröffentlicht: (2026)
SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
von: Wang, Jianyi, et al.
Veröffentlicht: (2025)
von: Wang, Jianyi, et al.
Veröffentlicht: (2025)
Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation
von: Li, Binglei, et al.
Veröffentlicht: (2026)
von: Li, Binglei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SARA: Structural and Adversarial Representation Alignment for Training-efficient Diffusion Models
von: Chen, Hesen, et al.
Veröffentlicht: (2025) -
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
von: Yang, Hao, et al.
Veröffentlicht: (2026) -
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
von: Qin, Luozheng, et al.
Veröffentlicht: (2025) -
Omni-Video: Democratizing Unified Video Understanding and Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025) -
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)