Autoregressive Styled Text Image Generation, but Make it Reliable
Fuente:
arXiv
Saved in:
| Main Authors: | Zaccagnino, Carmine, Quattrini, Fabio, Pippi, Vittorio, Cascianelli, Silvia, Tonioni, Alessio, Cucchiara, Rita |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Binarizing Documents by Leveraging both Space and Frequency
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026)
by: Zaccagnino, Carmine, et al.
Published: (2026)
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
VATr++: Choose Your Words Wisely for Handwritten Text Generation
by: Vanherle, Bram, et al.
Published: (2024)
by: Vanherle, Bram, et al.
Published: (2024)
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
by: Nikolaidou, Konstantina, et al.
Published: (2025)
by: Nikolaidou, Konstantina, et al.
Published: (2025)
Zero-Shot Synthetic-to-Real Handwritten Text Recognition via Task Analogies
by: Garrido-Munoz, Carlos, et al.
Published: (2026)
by: Garrido-Munoz, Carlos, et al.
Published: (2026)
Embodied Agents for Efficient Exploration and Smart Scene Description
by: Bigazzi, Roberto, et al.
Published: (2023)
by: Bigazzi, Roberto, et al.
Published: (2023)
Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
by: Landi, Federico, et al.
Published: (2022)
by: Landi, Federico, et al.
Published: (2022)
Embodied Navigation at the Art Gallery
by: Bigazzi, Roberto, et al.
Published: (2022)
by: Bigazzi, Roberto, et al.
Published: (2022)
Explore and Explain: Self-supervised Navigation and Recounting
by: Bigazzi, Roberto, et al.
Published: (2020)
by: Bigazzi, Roberto, et al.
Published: (2020)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026)
by: Bill, Eric Tillmann, et al.
Published: (2026)
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
by: Lobba, Davide, et al.
Published: (2025)
by: Lobba, Davide, et al.
Published: (2025)
Unveiling the Truth: Exploring Human Gaze Patterns in Fake Images
by: Cartella, Giuseppe, et al.
Published: (2024)
by: Cartella, Giuseppe, et al.
Published: (2024)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
by: Wang, Haofan, et al.
Published: (2024)
by: Wang, Haofan, et al.
Published: (2024)
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image Generation
by: Wang, Haofan, et al.
Published: (2024)
by: Wang, Haofan, et al.
Published: (2024)
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Generating Multimodal Images with GAN: Integrating Text, Image, and Style
by: Tan, Chaoyi, et al.
Published: (2025)
by: Tan, Chaoyi, et al.
Published: (2025)
A Training-Free Style-aligned Image Generation with Scale-wise Autoregressive Model
by: Park, Jihun, et al.
Published: (2025)
by: Park, Jihun, et al.
Published: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025)
by: Sanguigni, Fulvio, et al.
Published: (2025)
CSGO: Content-Style Composition in Text-to-Image Generation
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
Training-free Online Video Step Grounding
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
by: Shahbazi, Mohamad, et al.
Published: (2024)
by: Shahbazi, Mohamad, et al.
Published: (2024)
Similar Items
-
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
by: Pippi, Vittorio, et al.
Published: (2025) -
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024) -
Binarizing Documents by Leveraging both Space and Frequency
by: Quattrini, Fabio, et al.
Published: (2024) -
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
by: Quattrini, Fabio, et al.
Published: (2024) -
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026)