Improving face generation quality and prompt following with synthetic captions
Fuente:
arXiv
Saved in:
| Main Authors: | Tarasiou, Michail, Moschoglou, Stylianos, Deng, Jiankang, Zafeiriou, Stefanos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context-self contrastive pretraining for crop type semantic segmentation
by: Tarasiou, Michail, et al.
Published: (2021)
by: Tarasiou, Michail, et al.
Published: (2021)
ShapeFusion: A 3D diffusion model for localized shape editing
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
Arc2Face: A Foundation Model for ID-Consistent Human Faces
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024)
FitDiff: Robust monocular 3D facial shape and reflectance estimation using Diffusion Models
by: Galanakis, Stathis, et al.
Published: (2023)
by: Galanakis, Stathis, et al.
Published: (2023)
Locally Adaptive Neural 3D Morphable Models
by: Tarasiou, Michail, et al.
Published: (2024)
by: Tarasiou, Michail, et al.
Published: (2024)
SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models
by: Galanakis, Stathis, et al.
Published: (2025)
by: Galanakis, Stathis, et al.
Published: (2025)
Spatio-temporal Prompting Network for Robust Video Feature Extraction
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
AnimateMe: 4D Facial Expressions via Diffusion Models
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
UV-free Texture Generation with Denoising and Geodesic Heat Diffusions
by: Foti, Simone, et al.
Published: (2024)
by: Foti, Simone, et al.
Published: (2024)
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling
by: Babiloni, Francesca, et al.
Published: (2024)
by: Babiloni, Francesca, et al.
Published: (2024)
Text To 3D Object Generation For Scalable Room Assembly
by: Laguna, Sonia, et al.
Published: (2025)
by: Laguna, Sonia, et al.
Published: (2025)
Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
by: Barsbey, Melih, et al.
Published: (2025)
by: Barsbey, Melih, et al.
Published: (2025)
SAGS: Structure-Aware 3D Gaussian Splatting
by: Ververas, Evangelos, et al.
Published: (2024)
by: Ververas, Evangelos, et al.
Published: (2024)
ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation
by: Pratikaki, Chrysa, et al.
Published: (2026)
by: Pratikaki, Chrysa, et al.
Published: (2026)
Deep Face Restoration: A Survey
by: Wang, Tao, et al.
Published: (2022)
by: Wang, Tao, et al.
Published: (2022)
Evaluating authenticity and quality of image captions via sentiment and semantic analyses
by: Krotov, Aleksei, et al.
Published: (2024)
by: Krotov, Aleksei, et al.
Published: (2024)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
Parallelised Differentiable Straightest Geodesics for 3D Meshes
by: Verninas, Hippolyte, et al.
Published: (2026)
by: Verninas, Hippolyte, et al.
Published: (2026)
Neural Sign Actors: A diffusion model for 3D sign language production from text
by: Baltatzis, Vasileios, et al.
Published: (2023)
by: Baltatzis, Vasileios, et al.
Published: (2023)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics
by: Dirik, Alara, et al.
Published: (2026)
by: Dirik, Alara, et al.
Published: (2026)
Reverse Stable Diffusion: What prompt was used to generate this image?
by: Croitoru, Florinel-Alin, et al.
Published: (2023)
by: Croitoru, Florinel-Alin, et al.
Published: (2023)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
Semantic search for 100M+ galaxy images using AI-generated captions
by: Koblischke, Nolan, et al.
Published: (2025)
by: Koblischke, Nolan, et al.
Published: (2025)
Adaptive Parametric Activation: Unifying and Generalising Activation Functions Across Tasks
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
Deep Active Learning: A Reality Check
by: Gashi, Edrina, et al.
Published: (2024)
by: Gashi, Edrina, et al.
Published: (2024)
Importance of realism in procedurally-generated synthetic images for deep learning: case studies in maize and canola
by: Khan, Nazifa Azam, et al.
Published: (2024)
by: Khan, Nazifa Azam, et al.
Published: (2024)
Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera
by: Yu, Zhengdi, et al.
Published: (2024)
by: Yu, Zhengdi, et al.
Published: (2024)
Distribution Matching for Multi-Task Learning of Classification Tasks: a Large-Scale Study on Faces & Beyond
by: Kollias, Dimitrios, et al.
Published: (2024)
by: Kollias, Dimitrios, et al.
Published: (2024)
DermaFlux: Synthetic Skin Lesion Generation with Rectified Flows for Enhanced Image Classification
by: Galanakis, Stathis, et al.
Published: (2026)
by: Galanakis, Stathis, et al.
Published: (2026)
FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution Detector
by: Chen, Jiankang, et al.
Published: (2024)
by: Chen, Jiankang, et al.
Published: (2024)
Visual prompting reimagined: The power of the Activation Prompts
by: Zhang, Yihua, et al.
Published: (2026)
by: Zhang, Yihua, et al.
Published: (2026)
BECLR: Batch Enhanced Contrastive Few-Shot Learning
by: Poulakakis-Daktylidis, Stylianos, et al.
Published: (2024)
by: Poulakakis-Daktylidis, Stylianos, et al.
Published: (2024)
EEG-D3: A Solution to the Hidden Overfitting Problem of Deep Learning Models
by: Ludwig, Siegfried, et al.
Published: (2025)
by: Ludwig, Siegfried, et al.
Published: (2025)
Design2Cloth: 3D Cloth Generation from 2D Masks
by: Zheng, Jiali, et al.
Published: (2024)
by: Zheng, Jiali, et al.
Published: (2024)
Similar Items
-
Context-self contrastive pretraining for crop type semantic segmentation
by: Tarasiou, Michail, et al.
Published: (2021) -
ShapeFusion: A 3D diffusion model for localized shape editing
by: Potamias, Rolandos Alexandros, et al.
Published: (2024) -
Arc2Face: A Foundation Model for ID-Consistent Human Faces
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024) -
FitDiff: Robust monocular 3D facial shape and reflectance estimation using Diffusion Models
by: Galanakis, Stathis, et al.
Published: (2023) -
Locally Adaptive Neural 3D Morphable Models
by: Tarasiou, Michail, et al.
Published: (2024)