AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zarei, Arman, Pan, Jiacheng, Gwilliam, Matthew, Feizi, Soheil, Yang, Zhenheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
Localizing Knowledge in Diffusion Transformers
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
DREW : Towards Robust Data Provenance by Leveraging Error-Controlled Watermarking
by: Saberi, Mehrdad, et al.
Published: (2024)
by: Saberi, Mehrdad, et al.
Published: (2024)
Implicit Neural Representation Facilitates Unified Universal Vision Encoding
by: Gwilliam, Matthew, et al.
Published: (2026)
by: Gwilliam, Matthew, et al.
Published: (2026)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
by: Balasubramanian, Sriram, et al.
Published: (2024)
by: Balasubramanian, Sriram, et al.
Published: (2024)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
Understanding the Effect of using Semantically Meaningful Tokens for Visual Representation Learning
by: Kalibhat, Neha, et al.
Published: (2024)
by: Kalibhat, Neha, et al.
Published: (2024)
IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation
by: Zhang, Xinchen, et al.
Published: (2024)
by: Zhang, Xinchen, et al.
Published: (2024)
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
CompAgent: An Agentic Framework for Visual Compliance Verification
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
by: Jiang, Kaixun, et al.
Published: (2026)
by: Jiang, Kaixun, et al.
Published: (2026)
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
by: Aggarwal, Anirud, et al.
Published: (2025)
by: Aggarwal, Anirud, et al.
Published: (2025)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023)
by: Rezaei, Keivan, et al.
Published: (2023)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)
by: Jayatilaka, Gihan, et al.
Published: (2025)
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
by: Padmanabhan, Namitha, et al.
Published: (2026)
by: Padmanabhan, Namitha, et al.
Published: (2026)
Understanding Information Storage and Transfer in Multi-modal Large Language Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
by: Saberi, Mehrdad, et al.
Published: (2023)
by: Saberi, Mehrdad, et al.
Published: (2023)
How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric Encodings
by: Audia, Samuel, et al.
Published: (2025)
by: Audia, Samuel, et al.
Published: (2025)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
by: Sameti, Mohammad Hossein, et al.
Published: (2025)
by: Sameti, Mohammad Hossein, et al.
Published: (2025)
Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling
by: Ye, Ruijie, et al.
Published: (2026)
by: Ye, Ruijie, et al.
Published: (2026)
CRAFT: Continuous Reasoning and Agentic Feedback Tuning for Multimodal Text-to-Image Generation
by: Kovalev, V., et al.
Published: (2025)
by: Kovalev, V., et al.
Published: (2025)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023)
by: Huang, Kaiyi, et al.
Published: (2023)
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
by: Sun, Kaiyue, et al.
Published: (2024)
by: Sun, Kaiyue, et al.
Published: (2024)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
by: Zhu, Zixin, et al.
Published: (2025)
by: Zhu, Zixin, et al.
Published: (2025)
Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
by: Gao, Yixin, et al.
Published: (2025)
by: Gao, Yixin, et al.
Published: (2025)
Accelerate High-Quality Diffusion Models with Inner Loop Feedback
by: Gwilliam, Matthew, et al.
Published: (2025)
by: Gwilliam, Matthew, et al.
Published: (2025)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025)
by: Guo, Yifu, et al.
Published: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussians
by: Ge, Chongjian, et al.
Published: (2024)
by: Ge, Chongjian, et al.
Published: (2024)
ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion
by: Zhang, Zitian, et al.
Published: (2024)
by: Zhang, Zitian, et al.
Published: (2024)
Show-o2: Improved Native Unified Multimodal Models
by: Xie, Jinheng, et al.
Published: (2025)
by: Xie, Jinheng, et al.
Published: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
by: Li, Mingcheng, et al.
Published: (2025)
by: Li, Mingcheng, et al.
Published: (2025)
What do we learn from inverting CLIP models?
by: Kazemi, Hamid, et al.
Published: (2024)
by: Kazemi, Hamid, et al.
Published: (2024)
Similar Items
-
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024) -
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
by: Zarei, Arman, et al.
Published: (2025) -
Localizing Knowledge in Diffusion Transformers
by: Zarei, Arman, et al.
Published: (2025) -
DREW : Towards Robust Data Provenance by Leveraging Error-Controlled Watermarking
by: Saberi, Mehrdad, et al.
Published: (2024) -
Implicit Neural Representation Facilitates Unified Universal Vision Encoding
by: Gwilliam, Matthew, et al.
Published: (2026)