Alfie: Democratising RGBA Image Generation With No $$$
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Quattrini, Fabio, Pippi, Vittorio, Cascianelli, Silvia, Cucchiara, Rita |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Binarizing Documents by Leveraging both Space and Frequency
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
Autoregressive Styled Text Image Generation, but Make it Reliable
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2025)
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2025)
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
VATr++: Choose Your Words Wisely for Handwritten Text Generation
von: Vanherle, Bram, et al.
Veröffentlicht: (2024)
von: Vanherle, Bram, et al.
Veröffentlicht: (2024)
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024)
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
von: Sanguigni, Fulvio, et al.
Veröffentlicht: (2025)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images
von: Amoroso, Roberto, et al.
Veröffentlicht: (2023)
von: Amoroso, Roberto, et al.
Veröffentlicht: (2023)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2024)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2026)
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2026)
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
Tiny Inference-Time Scaling with Latent Verifiers
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2026)
von: Bucciarelli, Davide, et al.
Veröffentlicht: (2026)
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2025)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
Causal Graphical Models for Vision-Language Compositional Understanding
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Recurrence Meets Transformers for Universal Multimodal Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Hyperbolic Safety-Aware Vision-Language Models
von: Poppi, Tobia, et al.
Veröffentlicht: (2025)
von: Poppi, Tobia, et al.
Veröffentlicht: (2025)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
von: Poppi, Samuele, et al.
Veröffentlicht: (2023)
Semantically Consistent Person Image Generation
von: Roy, Prasun, et al.
Veröffentlicht: (2023)
von: Roy, Prasun, et al.
Veröffentlicht: (2023)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
G-Refine: A General Quality Refiner for Text-to-Image Generation
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
von: Poppi, Tobia, et al.
Veröffentlicht: (2026)
von: Poppi, Tobia, et al.
Veröffentlicht: (2026)
Bringing Textual Prompt to AI-Generated Image Quality Assessment
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
Scene Aware Person Image Generation through Global Contextual Conditioning
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
von: Deng, Yuchen, et al.
Veröffentlicht: (2025)
von: Deng, Yuchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Binarizing Documents by Leveraging both Space and Frequency
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024) -
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
von: Quattrini, Fabio, et al.
Veröffentlicht: (2024) -
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025) -
Autoregressive Styled Text Image Generation, but Make it Reliable
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2025) -
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
von: Pippi, Vittorio, et al.
Veröffentlicht: (2025)