Compositional Image Synthesis with Inference-Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Minsuk, Lee, Sanghyeok, Ahn, Namhyuk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
by: Kim, Sungnyun, et al.
Published: (2023)
by: Kim, Sungnyun, et al.
Published: (2023)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
BrainDecoder: Style-Based Visual Decoding of EEG Signals
by: Choi, Minsuk, et al.
Published: (2024)
by: Choi, Minsuk, et al.
Published: (2024)
Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
by: Phuntsho, Karma, et al.
Published: (2025)
by: Phuntsho, Karma, et al.
Published: (2025)
Inference-Time Diffusion Model Distillation
by: Park, Geon Yeong, et al.
Published: (2024)
by: Park, Geon Yeong, et al.
Published: (2024)
Training-free Composite Scene Generation for Layout-to-Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2024)
by: Liu, Jiaqi, et al.
Published: (2024)
Tiny Inference-Time Scaling with Latent Verifiers
by: Bucciarelli, Davide, et al.
Published: (2026)
by: Bucciarelli, Davide, et al.
Published: (2026)
GreenStableYolo: Optimizing Inference Time and Image Quality of Text-to-Image Generation
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
by: Rastegar, Sarah, et al.
Published: (2025)
by: Rastegar, Sarah, et al.
Published: (2025)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
by: Bang, Junseo, et al.
Published: (2025)
by: Bang, Junseo, et al.
Published: (2025)
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model
by: Lee, Donghwna, et al.
Published: (2024)
by: Lee, Donghwna, et al.
Published: (2024)
Cooperative Inference for Real-Time 3D Human Pose Estimation in Multi-Device Edge Networks
by: Choi, Hyun-Ho, et al.
Published: (2025)
by: Choi, Hyun-Ho, et al.
Published: (2025)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
by: Byun, Ji Young, et al.
Published: (2025)
by: Byun, Ji Young, et al.
Published: (2025)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
by: Park, Jihwan, et al.
Published: (2025)
by: Park, Jihwan, et al.
Published: (2025)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
by: Zhan, Wengyi, et al.
Published: (2024)
by: Zhan, Wengyi, et al.
Published: (2024)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
by: Kwon, Soyeong, et al.
Published: (2024)
by: Kwon, Soyeong, et al.
Published: (2024)
Real-Time Person Image Synthesis Using a Flow Matching Model
by: Jeong, Jiwoo, et al.
Published: (2025)
by: Jeong, Jiwoo, et al.
Published: (2025)
Scaling View Synthesis Transformers
by: Kim, Evan, et al.
Published: (2026)
by: Kim, Evan, et al.
Published: (2026)
Melon Fruit Detection and Quality Assessment Using Generative AI-Based Image Data Augmentation
by: Yoon, Seungri, et al.
Published: (2024)
by: Yoon, Seungri, et al.
Published: (2024)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
by: Lee, Youngwan, et al.
Published: (2023)
by: Lee, Youngwan, et al.
Published: (2023)
Inference-Time Scaling of Diffusion Models for Infrared Data Generation
by: Horstmann, Kai A., et al.
Published: (2025)
by: Horstmann, Kai A., et al.
Published: (2025)
AnyTrans: Translate AnyText in the Image with Large Scale Models
by: Qian, Zhipeng, et al.
Published: (2024)
by: Qian, Zhipeng, et al.
Published: (2024)
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
by: Shivika, et al.
Published: (2026)
by: Shivika, et al.
Published: (2026)
ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects
by: Lee, Woojin, et al.
Published: (2025)
by: Lee, Woojin, et al.
Published: (2025)
Development of Image Collection Method Using YOLO and Siamese Network
by: Shin, Chan Young, et al.
Published: (2024)
by: Shin, Chan Young, et al.
Published: (2024)
Test-Time Mixup Augmentation for Data and Class-Specific Uncertainty Estimation in Deep Learning Image Classification
by: Lee, Hansang, et al.
Published: (2022)
by: Lee, Hansang, et al.
Published: (2022)
Ctrl-VI: Controllable Video Synthesis via Variational Inference
by: Duan, Haoyi, et al.
Published: (2025)
by: Duan, Haoyi, et al.
Published: (2025)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
Towards Deconfounded Image-Text Matching with Causal Inference
by: Li, Wenhui, et al.
Published: (2024)
by: Li, Wenhui, et al.
Published: (2024)
FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
Imperceptible Protection against Style Imitation from Diffusion Models
by: Ahn, Namhyuk, et al.
Published: (2024)
by: Ahn, Namhyuk, et al.
Published: (2024)
Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion Models
by: Ahn, Namhyuk, et al.
Published: (2024)
by: Ahn, Namhyuk, et al.
Published: (2024)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
by: Huang, Nisha, et al.
Published: (2024)
by: Huang, Nisha, et al.
Published: (2024)
Token-Level Inference-Time Alignment for Vision-Language Models
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
by: Lu, Guansong, et al.
Published: (2023)
by: Lu, Guansong, et al.
Published: (2023)
Similar Items
-
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
by: Kim, Sungnyun, et al.
Published: (2023) -
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
by: Choi, Dasol, et al.
Published: (2025) -
BrainDecoder: Style-Based Visual Decoding of EEG Signals
by: Choi, Minsuk, et al.
Published: (2024) -
Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
by: Phuntsho, Karma, et al.
Published: (2025) -
Inference-Time Diffusion Model Distillation
by: Park, Geon Yeong, et al.
Published: (2024)