Kandinsky 3.0 Technical Report
Fuente:
arXiv
Salvato in:
| Autori principali: | Arkhipkin, Vladimir, Filatov, Andrei, Vasilev, Viacheslav, Maltseva, Anastasia, Azizov, Said, Pavlov, Igor, Agafonova, Julia, Kuznetsov, Andrey, Dimitrov, Denis |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2024)
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2024)
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
di: Novitskiy, Lev, et al.
Pubblicazione: (2025)
di: Novitskiy, Lev, et al.
Pubblicazione: (2025)
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2025)
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2025)
CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
di: Vasilev, Viacheslav, et al.
Pubblicazione: (2025)
di: Vasilev, Viacheslav, et al.
Pubblicazione: (2025)
ESQA: Event Sequences Question Answering
di: Abdullaeva, Irina, et al.
Pubblicazione: (2024)
di: Abdullaeva, Irina, et al.
Pubblicazione: (2024)
Human Aesthetic Preference-Based Large Text-to-Image Model Personalization: Kandinsky Generation as an Example
di: Zhou, Aven-Le, et al.
Pubblicazione: (2024)
di: Zhou, Aven-Le, et al.
Pubblicazione: (2024)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
di: Vasilev, Viacheslav, et al.
Pubblicazione: (2025)
di: Vasilev, Viacheslav, et al.
Pubblicazione: (2025)
Mano Technical Report
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
$\nabla$NABLA: Neighborhood Adaptive Block-Level Attention
di: Mikhailov, Dmitrii, et al.
Pubblicazione: (2025)
di: Mikhailov, Dmitrii, et al.
Pubblicazione: (2025)
Pegasus-v1 Technical Report
di: Jung, Raehyuk, et al.
Pubblicazione: (2024)
di: Jung, Raehyuk, et al.
Pubblicazione: (2024)
MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
di: Zhao, Haochen, et al.
Pubblicazione: (2025)
di: Zhao, Haochen, et al.
Pubblicazione: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
di: Gong, Jingyao
Pubblicazione: (2026)
di: Gong, Jingyao
Pubblicazione: (2026)
Less is More - diveXplore 5.0 at VBS 2021
di: Leibetseder, Andreas, et al.
Pubblicazione: (2025)
di: Leibetseder, Andreas, et al.
Pubblicazione: (2025)
Automated Radiology Report Generation Based on Topic-Keyword Semantic Guidance
di: Xiao, Jing, et al.
Pubblicazione: (2025)
di: Xiao, Jing, et al.
Pubblicazione: (2025)
diveXplore 6.0: ITEC's Interactive Video Exploration System at VBS 2022
di: Leibetseder, Andreas, et al.
Pubblicazione: (2025)
di: Leibetseder, Andreas, et al.
Pubblicazione: (2025)
FPGA‐Based Deep Neural Network Implementation for Handwritten Digit Recognition
di: Matej Štajnbrikner, et al.
Pubblicazione: (2025)
di: Matej Štajnbrikner, et al.
Pubblicazione: (2025)
Kimi-Audio Technical Report
di: KimiTeam, et al.
Pubblicazione: (2025)
di: KimiTeam, et al.
Pubblicazione: (2025)
MCPNS: A Macropixel Collocated Position and Its Neighbors Search for Plenoptic 2.0 Video Coding
di: Van Duong, Vinh, et al.
Pubblicazione: (2023)
di: Van Duong, Vinh, et al.
Pubblicazione: (2023)
The Sketchfab 3D Creative Commons Collection (S3D3C)
di: Spiess, Florian, et al.
Pubblicazione: (2024)
di: Spiess, Florian, et al.
Pubblicazione: (2024)
Leum-VL Technical Report
di: He, Yuxuan, et al.
Pubblicazione: (2026)
di: He, Yuxuan, et al.
Pubblicazione: (2026)
Smiling Regulates Emotion During Traumatic Recollection
di: Ma, Marcus, et al.
Pubblicazione: (2026)
di: Ma, Marcus, et al.
Pubblicazione: (2026)
Adaptive 3D Mesh Steganography Based on Feature-Preserving Distortion
di: Zhang, Yushu, et al.
Pubblicazione: (2022)
di: Zhang, Yushu, et al.
Pubblicazione: (2022)
Efficient Geometry Compression and Communication for 3D Gaussian Splatting Point Clouds
di: Xie, Liang, et al.
Pubblicazione: (2025)
di: Xie, Liang, et al.
Pubblicazione: (2025)
Subjective Quality Assessment of Dynamic 3D Meshes in Virtual Reality Environment
di: Nguyen, Duc V., et al.
Pubblicazione: (2026)
di: Nguyen, Duc V., et al.
Pubblicazione: (2026)
Perceptual Quality Assessment of Octree-RAHT Encoded 3D Point Clouds
di: Duan, Dongshuai, et al.
Pubblicazione: (2024)
di: Duan, Dongshuai, et al.
Pubblicazione: (2024)
A 3D Framework for Improving Low-Latency Multi-Channel Live Streaming
di: Aiersilan, Aizierjiang, et al.
Pubblicazione: (2024)
di: Aiersilan, Aizierjiang, et al.
Pubblicazione: (2024)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
di: Zhao, Yuan, et al.
Pubblicazione: (2026)
di: Zhao, Yuan, et al.
Pubblicazione: (2026)
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
di: Nguyen, Duc, et al.
Pubblicazione: (2024)
di: Nguyen, Duc, et al.
Pubblicazione: (2024)
High Capacity Reversible Data Hiding for Encrypted 3D Mesh Models Based on Topology
di: Tang, Yun, et al.
Pubblicazione: (2022)
di: Tang, Yun, et al.
Pubblicazione: (2022)
Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture Generation
di: Gao, Zhilin, et al.
Pubblicazione: (2025)
di: Gao, Zhilin, et al.
Pubblicazione: (2025)
Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering
di: Fei, Ben, et al.
Pubblicazione: (2024)
di: Fei, Ben, et al.
Pubblicazione: (2024)
High-Fidelity 3D Gaussian Human Reconstruction via Region-Aware Initialization and Geometric Priors
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
3DMambaIPF: A State Space Model for Iterative Point Cloud Filtering via Differentiable Rendering
di: Zhou, Qingyuan, et al.
Pubblicazione: (2024)
di: Zhou, Qingyuan, et al.
Pubblicazione: (2024)
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
di: Yang, An, et al.
Pubblicazione: (2025)
di: Yang, An, et al.
Pubblicazione: (2025)
TSC-PCAC: Voxel Transformer and Sparse Convolution Based Point Cloud Attribute Compression for 3D Broadcasting
di: Guo, Zixi, et al.
Pubblicazione: (2024)
di: Guo, Zixi, et al.
Pubblicazione: (2024)
Structure-Aware Residual-Center Representation for Self-Supervised Open-Set 3D Cross-Modal Retrieval
di: Xu, Yang, et al.
Pubblicazione: (2024)
di: Xu, Yang, et al.
Pubblicazione: (2024)
A 3D-Cascading Crossing Coupling Framework for Hyperchaotic Map Construction and Its Application to Color Image Encryption
di: Sun, Jilei, et al.
Pubblicazione: (2025)
di: Sun, Jilei, et al.
Pubblicazione: (2025)
LongCat-Flash-Omni Technical Report
di: Meituan LongCat Team, et al.
Pubblicazione: (2025)
di: Meituan LongCat Team, et al.
Pubblicazione: (2025)
SyMuPe: Affective and Controllable Symbolic Music Performance
di: Borovik, Ilya, et al.
Pubblicazione: (2025)
di: Borovik, Ilya, et al.
Pubblicazione: (2025)
HiLight: Technical Report on the Motern AI Video Language Model
di: Wang, Zhiting, et al.
Pubblicazione: (2024)
di: Wang, Zhiting, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2024) -
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
di: Novitskiy, Lev, et al.
Pubblicazione: (2025) -
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
di: Arkhipkin, Vladimir, et al.
Pubblicazione: (2025) -
CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
di: Vasilev, Viacheslav, et al.
Pubblicazione: (2025) -
ESQA: Event Sequences Question Answering
di: Abdullaeva, Irina, et al.
Pubblicazione: (2024)