YaART: Yet Another ART Rendering Technology
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917633792671744 |
|---|---|
| author | Kastryulin, Sergey Konev, Artem Shishenya, Alexander Lyapustin, Eugene Khurshudov, Artem Tselousov, Alexander Vinokurov, Nikita Kuznedelev, Denis Markovich, Alexander Livshits, Grigoriy Kirillov, Alexey Tabisheva, Anastasiia Chubarova, Liubov Kaminskaia, Marina Ustyuzhanin, Alexander Shvetsov, Artemii Shlenskii, Daniil Startsev, Valerii Kornilov, Dmitrii Romanov, Mikhail Babenko, Artem Ovcharenko, Sergei Khrulkov, Valentin |
| author_facet | Kastryulin, Sergey Konev, Artem Shishenya, Alexander Lyapustin, Eugene Khurshudov, Artem Tselousov, Alexander Vinokurov, Nikita Kuznedelev, Denis Markovich, Alexander Livshits, Grigoriy Kirillov, Alexey Tabisheva, Anastasiia Chubarova, Liubov Kaminskaia, Marina Ustyuzhanin, Alexander Shvetsov, Artemii Shlenskii, Daniil Startsev, Valerii Kornilov, Dmitrii Romanov, Mikhail Babenko, Artem Ovcharenko, Sergei Khrulkov, Valentin |
| contents | In the rapidly progressing field of generative models, the development of efficient and high-fidelity text-to-image diffusion systems represents a significant frontier. This study introduces YaART, a novel production-grade text-to-image cascaded diffusion model aligned to human preferences using Reinforcement Learning from Human Feedback (RLHF). During the development of YaART, we especially focus on the choices of the model and training dataset sizes, the aspects that were not systematically investigated for text-to-image cascaded diffusion models before. In particular, we comprehensively analyze how these choices affect both the efficiency of the training process and the quality of the generated images, which are highly important in practice. Furthermore, we demonstrate that models trained on smaller datasets of higher-quality images can successfully compete with those trained on larger datasets, establishing a more efficient scenario of diffusion models training. From the quality perspective, YaART is consistently preferred by users over many existing state-of-the-art models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_05666 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | YaART: Yet Another ART Rendering Technology Kastryulin, Sergey Konev, Artem Shishenya, Alexander Lyapustin, Eugene Khurshudov, Artem Tselousov, Alexander Vinokurov, Nikita Kuznedelev, Denis Markovich, Alexander Livshits, Grigoriy Kirillov, Alexey Tabisheva, Anastasiia Chubarova, Liubov Kaminskaia, Marina Ustyuzhanin, Alexander Shvetsov, Artemii Shlenskii, Daniil Startsev, Valerii Kornilov, Dmitrii Romanov, Mikhail Babenko, Artem Ovcharenko, Sergei Khrulkov, Valentin Computer Vision and Pattern Recognition In the rapidly progressing field of generative models, the development of efficient and high-fidelity text-to-image diffusion systems represents a significant frontier. This study introduces YaART, a novel production-grade text-to-image cascaded diffusion model aligned to human preferences using Reinforcement Learning from Human Feedback (RLHF). During the development of YaART, we especially focus on the choices of the model and training dataset sizes, the aspects that were not systematically investigated for text-to-image cascaded diffusion models before. In particular, we comprehensively analyze how these choices affect both the efficiency of the training process and the quality of the generated images, which are highly important in practice. Furthermore, we demonstrate that models trained on smaller datasets of higher-quality images can successfully compete with those trained on larger datasets, establishing a more efficient scenario of diffusion models training. From the quality perspective, YaART is consistently preferred by users over many existing state-of-the-art models. |
| title | YaART: Yet Another ART Rendering Technology |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2404.05666 |