Zero-Shot Styled Text Image Generation, but Make It Autoregressive
Fuente:
arXiv
Saved in:
| Main Authors: | Pippi, Vittorio, Quattrini, Fabio, Cascianelli, Silvia, Tonioni, Alessio, Cucchiara, Rita |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autoregressive Styled Text Image Generation, but Make it Reliable
by: Zaccagnino, Carmine, et al.
Published: (2025)
by: Zaccagnino, Carmine, et al.
Published: (2025)
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Binarizing Documents by Leveraging both Space and Frequency
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026)
by: Zaccagnino, Carmine, et al.
Published: (2026)
VATr++: Choose Your Words Wisely for Handwritten Text Generation
by: Vanherle, Bram, et al.
Published: (2024)
by: Vanherle, Bram, et al.
Published: (2024)
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Zero-Shot Synthetic-to-Real Handwritten Text Recognition via Task Analogies
by: Garrido-Munoz, Carlos, et al.
Published: (2026)
by: Garrido-Munoz, Carlos, et al.
Published: (2026)
Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
by: Nikolaidou, Konstantina, et al.
Published: (2025)
by: Nikolaidou, Konstantina, et al.
Published: (2025)
Embodied Agents for Efficient Exploration and Smart Scene Description
by: Bigazzi, Roberto, et al.
Published: (2023)
by: Bigazzi, Roberto, et al.
Published: (2023)
Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
by: Landi, Federico, et al.
Published: (2022)
by: Landi, Federico, et al.
Published: (2022)
Embodied Navigation at the Art Gallery
by: Bigazzi, Roberto, et al.
Published: (2022)
by: Bigazzi, Roberto, et al.
Published: (2022)
Explore and Explain: Self-supervised Navigation and Recounting
by: Bigazzi, Roberto, et al.
Published: (2020)
by: Bigazzi, Roberto, et al.
Published: (2020)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026)
by: Bill, Eric Tillmann, et al.
Published: (2026)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
by: Lobba, Davide, et al.
Published: (2025)
by: Lobba, Davide, et al.
Published: (2025)
Unveiling the Truth: Exploring Human Gaze Patterns in Fake Images
by: Cartella, Giuseppe, et al.
Published: (2024)
by: Cartella, Giuseppe, et al.
Published: (2024)
Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation
by: Borgi, Alessio, et al.
Published: (2025)
by: Borgi, Alessio, et al.
Published: (2025)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
Fluent and Accurate Image Captioning with a Self-Trained Reward Model
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
StyleSculptor: Zero-Shot Style-Controllable 3D Asset Generation with Texture-Geometry Dual Guidance
by: Qu, Zefan, et al.
Published: (2025)
by: Qu, Zefan, et al.
Published: (2025)
HAC: Parameter-Efficient Hyperbolic Adaptation of CLIP for Zero-Shot VQA
by: Dibitonto, Francesco, et al.
Published: (2026)
by: Dibitonto, Francesco, et al.
Published: (2026)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
by: Shin, Chaehun, et al.
Published: (2024)
by: Shin, Chaehun, et al.
Published: (2024)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
Next-Scale Autoregressive Models are Zero-Shot Single-Image Object View Synthesizers
by: Yuan, Shiran, et al.
Published: (2025)
by: Yuan, Shiran, et al.
Published: (2025)
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
by: Wang, Haofan, et al.
Published: (2024)
by: Wang, Haofan, et al.
Published: (2024)
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image Generation
by: Wang, Haofan, et al.
Published: (2024)
by: Wang, Haofan, et al.
Published: (2024)
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Generating Multimodal Images with GAN: Integrating Text, Image, and Style
by: Tan, Chaoyi, et al.
Published: (2025)
by: Tan, Chaoyi, et al.
Published: (2025)
Zero-Shot Detection of AI-Generated Images
by: Cozzolino, Davide, et al.
Published: (2024)
by: Cozzolino, Davide, et al.
Published: (2024)
A Training-Free Style-aligned Image Generation with Scale-wise Autoregressive Model
by: Park, Jihun, et al.
Published: (2025)
by: Park, Jihun, et al.
Published: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025)
by: Sanguigni, Fulvio, et al.
Published: (2025)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
by: Zhang, Wentao, et al.
Published: (2024)
by: Zhang, Wentao, et al.
Published: (2024)
CSGO: Content-Style Composition in Text-to-Image Generation
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
Similar Items
-
Autoregressive Styled Text Image Generation, but Make it Reliable
by: Zaccagnino, Carmine, et al.
Published: (2025) -
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024) -
Binarizing Documents by Leveraging both Space and Frequency
by: Quattrini, Fabio, et al.
Published: (2024) -
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
by: Quattrini, Fabio, et al.
Published: (2024) -
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026)