Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Zaccagnino, Carmine, Quattrini, Fabio, Simsar, Enis, Gazulla, Marta Tintoré, Cucchiara, Rita, Tonioni, Alessio, Cascianelli, Silvia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autoregressive Styled Text Image Generation, but Make it Reliable
by: Zaccagnino, Carmine, et al.
Published: (2025)
by: Zaccagnino, Carmine, et al.
Published: (2025)
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
Binarizing Documents by Leveraging both Space and Frequency
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026)
by: Bill, Eric Tillmann, et al.
Published: (2026)
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)
by: Segu, Mattia, et al.
Published: (2025)
ToFu: Visual Tokens Reduction via Fusion for Multi-modal, Multi-patch, Multi-image Task
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
by: Bill, Eric Tillmann, et al.
Published: (2025)
by: Bill, Eric Tillmann, et al.
Published: (2025)
Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
by: Nikolaidou, Konstantina, et al.
Published: (2025)
by: Nikolaidou, Konstantina, et al.
Published: (2025)
Embodied Agents for Efficient Exploration and Smart Scene Description
by: Bigazzi, Roberto, et al.
Published: (2023)
by: Bigazzi, Roberto, et al.
Published: (2023)
Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
by: Landi, Federico, et al.
Published: (2022)
by: Landi, Federico, et al.
Published: (2022)
Embodied Navigation at the Art Gallery
by: Bigazzi, Roberto, et al.
Published: (2022)
by: Bigazzi, Roberto, et al.
Published: (2022)
IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
VATr++: Choose Your Words Wisely for Handwritten Text Generation
by: Vanherle, Bram, et al.
Published: (2024)
by: Vanherle, Bram, et al.
Published: (2024)
Explore and Explain: Self-supervised Navigation and Recounting
by: Bigazzi, Roberto, et al.
Published: (2020)
by: Bigazzi, Roberto, et al.
Published: (2020)
JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
by: Bill, Eric Tillmann, et al.
Published: (2025)
by: Bill, Eric Tillmann, et al.
Published: (2025)
LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
Zero-Shot Synthetic-to-Real Handwritten Text Recognition via Task Analogies
by: Garrido-Munoz, Carlos, et al.
Published: (2026)
by: Garrido-Munoz, Carlos, et al.
Published: (2026)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
by: An, Zhaochong, et al.
Published: (2026)
by: An, Zhaochong, et al.
Published: (2026)
PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM
by: Stefanache, Stefan, et al.
Published: (2024)
by: Stefanache, Stefan, et al.
Published: (2024)
Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
by: Meral, Tuna Han Salih, et al.
Published: (2024)
by: Meral, Tuna Han Salih, et al.
Published: (2024)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models
by: Zheng, Matthew, et al.
Published: (2024)
by: Zheng, Matthew, et al.
Published: (2024)
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
by: Xia, Tianxiang, et al.
Published: (2025)
by: Xia, Tianxiang, et al.
Published: (2025)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025)
by: Sanguigni, Fulvio, et al.
Published: (2025)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
by: Sanguigni, Fulvio, et al.
Published: (2026)
by: Sanguigni, Fulvio, et al.
Published: (2026)
3D Focusing-and-Matching Network for Multi-Instance Point Cloud Registration
by: Zhang, Liyuan, et al.
Published: (2024)
by: Zhang, Liyuan, et al.
Published: (2024)
MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
by: Huang, Jiahui, et al.
Published: (2026)
by: Huang, Jiahui, et al.
Published: (2026)
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing
by: Liu, Ziqian, et al.
Published: (2026)
by: Liu, Ziqian, et al.
Published: (2026)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
Similar Items
-
Autoregressive Styled Text Image Generation, but Make it Reliable
by: Zaccagnino, Carmine, et al.
Published: (2025) -
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
by: Quattrini, Fabio, et al.
Published: (2024) -
Zero-Shot Styled Text Image Generation, but Make It Autoregressive
by: Pippi, Vittorio, et al.
Published: (2025) -
Binarizing Documents by Leveraging both Space and Frequency
by: Quattrini, Fabio, et al.
Published: (2024) -
Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
by: Quattrini, Fabio, et al.
Published: (2024)