Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Arkhipkin, Vladimir, Vasilev, Viacheslav, Filatov, Andrei, Pavlov, Igor, Agafonova, Julia, Gerasimenko, Nikolai, Averchenkova, Anna, Mironova, Evelina, Bukashkin, Anton, Kulikov, Konstantin, Kuznetsov, Andrey, Dimitrov, Denis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kandinsky 3.0 Technical Report
by: Arkhipkin, Vladimir, et al.
Published: (2023)
by: Arkhipkin, Vladimir, et al.
Published: (2023)
CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
by: Vasilev, Viacheslav, et al.
Published: (2025)
by: Vasilev, Viacheslav, et al.
Published: (2025)
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
by: Novitskiy, Lev, et al.
Published: (2025)
by: Novitskiy, Lev, et al.
Published: (2025)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
by: Vasilev, Viacheslav, et al.
Published: (2025)
by: Vasilev, Viacheslav, et al.
Published: (2025)
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
by: Arkhipkin, Vladimir, et al.
Published: (2025)
by: Arkhipkin, Vladimir, et al.
Published: (2025)
ESQA: Event Sequences Question Answering
by: Abdullaeva, Irina, et al.
Published: (2024)
by: Abdullaeva, Irina, et al.
Published: (2024)
$\nabla$NABLA: Neighborhood Adaptive Block-Level Attention
by: Mikhailov, Dmitrii, et al.
Published: (2025)
by: Mikhailov, Dmitrii, et al.
Published: (2025)
Human Aesthetic Preference-Based Large Text-to-Image Model Personalization: Kandinsky Generation as an Example
by: Zhou, Aven-Le, et al.
Published: (2024)
by: Zhou, Aven-Le, et al.
Published: (2024)
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024)
by: Razzhigaev, Anton, et al.
Published: (2024)
FPGA‐Based Deep Neural Network Implementation for Handwritten Digit Recognition
by: Matej Štajnbrikner, et al.
Published: (2025)
by: Matej Štajnbrikner, et al.
Published: (2025)
Stickers on Facebook: Multifunctionality and face-enhancing politeness in everyday social interaction
by: Porrino-Moscoso, Laura M.
Published: (2026)
by: Porrino-Moscoso, Laura M.
Published: (2026)
SyMuPe: Affective and Controllable Symbolic Music Performance
by: Borovik, Ilya, et al.
Published: (2025)
by: Borovik, Ilya, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Towards Railways Remote Driving: Analysis of Video Streaming Latency and Adaptive Rate Control
by: Mejias, Daniel, et al.
Published: (2024)
by: Mejias, Daniel, et al.
Published: (2024)
Inferencias causales durante la comprensión de textos expositivos en formato multimedia
by: Gastón Saux
Published: (2012)
by: Gastón Saux
Published: (2012)
Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise
by: Xu, Zijing, et al.
Published: (2025)
by: Xu, Zijing, et al.
Published: (2025)
CAFE: Channel-Autoregressive Factorized Encoding for Robust Biosignal Spatial Super-Resolution
by: Liu, Hongjun, et al.
Published: (2026)
by: Liu, Hongjun, et al.
Published: (2026)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026)
by: Zhou, Qianrui, et al.
Published: (2026)
TPIFM: A Task-Aware Model for Evaluating Perceptual Interaction Fluency in Remote AR Collaboration
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
Learning Quality from Complexity and Structure: A Feature-Fused XGBoost Model for Video Quality Assessment
by: Premkumar, Amritha, et al.
Published: (2025)
by: Premkumar, Amritha, et al.
Published: (2025)
ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection
by: Yu, Zihao, et al.
Published: (2025)
by: Yu, Zihao, et al.
Published: (2025)
Merge Mode for Template-based Intra Mode Derivation (TIMD) in ECM
by: Abdoli, Mohsen, et al.
Published: (2025)
by: Abdoli, Mohsen, et al.
Published: (2025)
A 3D Framework for Improving Low-Latency Multi-Channel Live Streaming
by: Aiersilan, Aizierjiang, et al.
Published: (2024)
by: Aiersilan, Aizierjiang, et al.
Published: (2024)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
by: Liu, Shuhang, et al.
Published: (2025)
by: Liu, Shuhang, et al.
Published: (2025)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
Evaluation of Objective Image Quality Metrics for High-Fidelity Image Compression
by: Mohammadi, Shima, et al.
Published: (2025)
by: Mohammadi, Shima, et al.
Published: (2025)
Reply with Sticker: New Dataset and Model for Sticker Retrieval
by: Liang, Bin, et al.
Published: (2024)
by: Liang, Bin, et al.
Published: (2024)
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
by: Zhang, Beibei, et al.
Published: (2025)
by: Zhang, Beibei, et al.
Published: (2025)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
TAROT: Towards Optimization-Driven Adaptive FEC Parameter Tuning for Video Streaming
by: Sidhu, Jashanjot Singh, et al.
Published: (2026)
by: Sidhu, Jashanjot Singh, et al.
Published: (2026)
Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation
by: Che, Xinyi, et al.
Published: (2025)
by: Che, Xinyi, et al.
Published: (2025)
SFQA: A Comprehensive Perceptual Quality Assessment Dataset for Singing Face Generation
by: Gao, Zhilin, et al.
Published: (2026)
by: Gao, Zhilin, et al.
Published: (2026)
Block Erasure-Aware Semantic Multimedia Compression via JSCC Autoencoder
by: Esfahanizadeh, Homa, et al.
Published: (2026)
by: Esfahanizadeh, Homa, et al.
Published: (2026)
Subjective Evaluation of Frame Rate in Bitrate-Constrained Live Streaming
by: He, Jiaqi, et al.
Published: (2026)
by: He, Jiaqi, et al.
Published: (2026)
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024)
by: Zhao, Cairong, et al.
Published: (2024)
Feedback-Driven Rate Control for Learned Video Compression
by: Xu, Zhiheng, et al.
Published: (2026)
by: Xu, Zhiheng, et al.
Published: (2026)
Contextual Wireless Video Semantic Communication in MIMO-OFDM Systems
by: Xie, Bingyan, et al.
Published: (2026)
by: Xie, Bingyan, et al.
Published: (2026)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
CLIPRerank: An Extremely Simple Method for Improving Ad-hoc Video Search
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Editing on the Generative Manifold: A Theoretical and Empirical Study of General Diffusion-Based Image Editing Trade-offs
by: Hu, Yi, et al.
Published: (2026)
by: Hu, Yi, et al.
Published: (2026)
Similar Items
-
Kandinsky 3.0 Technical Report
by: Arkhipkin, Vladimir, et al.
Published: (2023) -
CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
by: Vasilev, Viacheslav, et al.
Published: (2025) -
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
by: Novitskiy, Lev, et al.
Published: (2025) -
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
by: Vasilev, Viacheslav, et al.
Published: (2025) -
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
by: Arkhipkin, Vladimir, et al.
Published: (2025)