Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Zitong, Zhang, Kaidong, Ding, Yukang, Gao, Chao, Ding, Rui, Chen, Ying, Zuo, Wangmeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
von: Qu, Yunpeng, et al.
Veröffentlicht: (2026)
von: Qu, Yunpeng, et al.
Veröffentlicht: (2026)
Gaze into the Details: Locality-Sensitive Enhancement for OCTA Retinal Vessel Segmentation
von: Huang, Tuopusen, et al.
Veröffentlicht: (2026)
von: Huang, Tuopusen, et al.
Veröffentlicht: (2026)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
von: Gan, Chaofan, et al.
Veröffentlicht: (2025)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
von: Tang, Changli, et al.
Veröffentlicht: (2024)
von: Tang, Changli, et al.
Veröffentlicht: (2024)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
StructSR: Refuse Spurious Details in Real-World Image Super-Resolution
von: Li, Yachao, et al.
Veröffentlicht: (2025)
von: Li, Yachao, et al.
Veröffentlicht: (2025)
DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
von: Liu, Yiheng, et al.
Veröffentlicht: (2025)
von: Liu, Yiheng, et al.
Veröffentlicht: (2025)
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
von: Yu, Haodong, et al.
Veröffentlicht: (2026)
von: Yu, Haodong, et al.
Veröffentlicht: (2026)
FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation
von: Bai, Yunpeng, et al.
Veröffentlicht: (2024)
von: Bai, Yunpeng, et al.
Veröffentlicht: (2024)
Mind the Detail: Uncovering Clinically Relevant Image Details in Accelerated MRI with Semantically Diverse Reconstructions
von: Morshuis, Jan Nikolas, et al.
Veröffentlicht: (2025)
von: Morshuis, Jan Nikolas, et al.
Veröffentlicht: (2025)
MR-GDINO: Efficient Open-World Continual Object Detection
von: Dong, Bowen, et al.
Veröffentlicht: (2024)
von: Dong, Bowen, et al.
Veröffentlicht: (2024)
DetailSemNet: Elevating Signature Verification through Detail-Semantic Integration
von: Shih, Meng-Cheng, et al.
Veröffentlicht: (2025)
von: Shih, Meng-Cheng, et al.
Veröffentlicht: (2025)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
von: Zhang, Yabo, et al.
Veröffentlicht: (2024)
von: Zhang, Yabo, et al.
Veröffentlicht: (2024)
Discriminator-Free Direct Preference Optimization for Video Diffusion
von: Cheng, Haoran, et al.
Veröffentlicht: (2025)
von: Cheng, Haoran, et al.
Veröffentlicht: (2025)
VideoGigaGAN: Towards Detail-rich Video Super-Resolution
von: Xu, Yiran, et al.
Veröffentlicht: (2024)
von: Xu, Yiran, et al.
Veröffentlicht: (2024)
AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion
von: Huang, Yangyi, et al.
Veröffentlicht: (2025)
von: Huang, Yangyi, et al.
Veröffentlicht: (2025)
Lie Flow: Video Dynamic Fields Modeling and Predicting with Lie Algebra as Geometric Physics Principle
von: Qiao, Weidong, et al.
Veröffentlicht: (2026)
von: Qiao, Weidong, et al.
Veröffentlicht: (2026)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
IMWA: Iterative Model Weight Averaging Benefits Class-Imbalanced Learning Tasks
von: Huang, Zitong, et al.
Veröffentlicht: (2024)
von: Huang, Zitong, et al.
Veröffentlicht: (2024)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
von: Liu, Huaize, et al.
Veröffentlicht: (2025)
von: Liu, Huaize, et al.
Veröffentlicht: (2025)
Detail-Preserving Latent Diffusion for Stable Shadow Removal
von: Xu, Jiamin, et al.
Veröffentlicht: (2024)
von: Xu, Jiamin, et al.
Veröffentlicht: (2024)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
GALA: Geometry-Aware Local Adaptive Grids for Detailed 3D Generation
von: Yang, Dingdong, et al.
Veröffentlicht: (2024)
von: Yang, Dingdong, et al.
Veröffentlicht: (2024)
Rethinking Direct Preference Optimization in Diffusion Models
von: Kang, Junyong, et al.
Veröffentlicht: (2025)
von: Kang, Junyong, et al.
Veröffentlicht: (2025)
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
von: Zhu, Jingyuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jingyuan, et al.
Veröffentlicht: (2026)
Enhancing Perceptual Quality in Video Super-Resolution through Temporally-Consistent Detail Synthesis using Diffusion Models
von: Rota, Claudio, et al.
Veröffentlicht: (2023)
von: Rota, Claudio, et al.
Veröffentlicht: (2023)
Real-Time Animatable 2DGS-Avatars with Detail Enhancement from Monocular Videos
von: Yuan, Xia, et al.
Veröffentlicht: (2025)
von: Yuan, Xia, et al.
Veröffentlicht: (2025)
See In Detail: Enhancing Sparse-view 3D Gaussian Splatting with Local Depth and Semantic Regularization
von: He, Zongqi, et al.
Veröffentlicht: (2025)
von: He, Zongqi, et al.
Veröffentlicht: (2025)
LOD-GS: Level-of-Detail-Sensitive 3D Gaussian Splatting for Detail Conserved Anti-Aliasing
von: Yang, Zhenya, et al.
Veröffentlicht: (2025)
von: Yang, Zhenya, et al.
Veröffentlicht: (2025)
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
von: Zheng, Mi, et al.
Veröffentlicht: (2025)
von: Zheng, Mi, et al.
Veröffentlicht: (2025)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
von: Sun, Yujing, et al.
Veröffentlicht: (2025)
von: Sun, Yujing, et al.
Veröffentlicht: (2025)
ShoeModel: Learning to Wear on the User-specified Shoes via Diffusion Model
von: Chen, Binghui, et al.
Veröffentlicht: (2024)
von: Chen, Binghui, et al.
Veröffentlicht: (2024)
Controllable Generative Video Compression
von: Ding, Ding, et al.
Veröffentlicht: (2026)
von: Ding, Ding, et al.
Veröffentlicht: (2026)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
von: Jiao, Qirui, et al.
Veröffentlicht: (2025)
von: Jiao, Qirui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
von: Qu, Yunpeng, et al.
Veröffentlicht: (2026) -
Gaze into the Details: Locality-Sensitive Enhancement for OCTA Retinal Vessel Segmentation
von: Huang, Tuopusen, et al.
Veröffentlicht: (2026) -
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
von: Dang, Jisheng, et al.
Veröffentlicht: (2025) -
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025) -
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)