FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Bill, Eric Tillmann, Simsar, Enis, Hofmann, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026)
by: Bill, Eric Tillmann, et al.
Published: (2026)
JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
by: Bill, Eric Tillmann, et al.
Published: (2025)
by: Bill, Eric Tillmann, et al.
Published: (2025)
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
by: Xia, Tianxiang, et al.
Published: (2025)
by: Xia, Tianxiang, et al.
Published: (2025)
LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models
by: Zheng, Matthew, et al.
Published: (2024)
by: Zheng, Matthew, et al.
Published: (2024)
Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
by: Meral, Tuna Han Salih, et al.
Published: (2024)
by: Meral, Tuna Han Salih, et al.
Published: (2024)
IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM
by: Stefanache, Stefan, et al.
Published: (2024)
by: Stefanache, Stefan, et al.
Published: (2024)
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026)
by: Zaccagnino, Carmine, et al.
Published: (2026)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
by: Yang, Mingyue, et al.
Published: (2025)
by: Yang, Mingyue, et al.
Published: (2025)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes
by: Hamamci, Ibrahim Ethem, et al.
Published: (2023)
by: Hamamci, Ibrahim Ethem, et al.
Published: (2023)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
Compositional Image-Text Matching and Retrieval by Grounding Entities
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
by: Vongala, Madhukar Reddy, et al.
Published: (2025)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
FOCUS: Bridging Fine-Grained Recognition and Open-World Discovery across Domains
by: Rathore, Vaibhav, et al.
Published: (2026)
by: Rathore, Vaibhav, et al.
Published: (2026)
EliGen: Entity-Level Controlled Image Generation with Regional Attention
by: Zhang, Hong, et al.
Published: (2025)
by: Zhang, Hong, et al.
Published: (2025)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
by: Zhu, Shangwen, et al.
Published: (2026)
by: Zhu, Shangwen, et al.
Published: (2026)
OpenAI ChatGPT interprets Radiological Images: GPT-4 as a Medical Doctor for a Fast Check-Up
by: Aydin, Omer, et al.
Published: (2025)
by: Aydin, Omer, et al.
Published: (2025)
A Review on Generative AI For Text-To-Image and Image-To-Image Generation and Implications To Scientific Images
by: Sordo, Zineb, et al.
Published: (2025)
by: Sordo, Zineb, et al.
Published: (2025)
Yume-1.5: A Text-Controlled Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
by: Cao, Pu, et al.
Published: (2024)
by: Cao, Pu, et al.
Published: (2024)
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
by: Tudosiu, Petru-Daniel, et al.
Published: (2024)
by: Tudosiu, Petru-Daniel, et al.
Published: (2024)
OptiWorld: Optimal Control for Video World Generation under Physical Constraints
by: Yuan, Yu, et al.
Published: (2026)
by: Yuan, Yu, et al.
Published: (2026)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
DynGhost: Temporally-Modelled Transformer for Dynamic Ghost Imaging with Quantum Detectors
by: Palladino, Vittorio, et al.
Published: (2026)
by: Palladino, Vittorio, et al.
Published: (2026)
FOCUS -- Multi-View Foot Reconstruction From Synthetically Trained Dense Correspondences
by: Boyne, Oliver, et al.
Published: (2025)
by: Boyne, Oliver, et al.
Published: (2025)
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
by: Sun, Haoze, et al.
Published: (2024)
by: Sun, Haoze, et al.
Published: (2024)
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
by: Son, Moo Hyun, et al.
Published: (2025)
by: Son, Moo Hyun, et al.
Published: (2025)
Boosting Diffusion-Based Text Image Super-Resolution Model Towards Generalized Real-World Scenarios
by: Pan, Chenglu, et al.
Published: (2025)
by: Pan, Chenglu, et al.
Published: (2025)
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
FlowDet: Unifying Object Detection and Generative Transport Flows
by: Baty, Enis, et al.
Published: (2025)
by: Baty, Enis, et al.
Published: (2025)
Similar Items
-
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
by: Bill, Eric Tillmann, et al.
Published: (2026) -
JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
by: Bill, Eric Tillmann, et al.
Published: (2025) -
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
by: Yang, Han, et al.
Published: (2024) -
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
by: Xia, Tianxiang, et al.
Published: (2025) -
LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
by: Simsar, Enis, et al.
Published: (2024)