APE: Agentic Prompt Enhancer for Image Generation and Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zijian, Wu, Jay Zhangjie, Wang, Zian, Cao, Tianshi, Chen, Jiasi, Fidler, Sanja, Ling, Huan, Ren, Xuanchi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
by: Lu, Yifan, et al.
Published: (2026)
by: Lu, Yifan, et al.
Published: (2026)
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
by: Ren, Xuanchi, et al.
Published: (2024)
by: Ren, Xuanchi, et al.
Published: (2024)
Lyra 2.0: Explorable Generative 3D Worlds
by: Shen, Tianchang, et al.
Published: (2026)
by: Shen, Tianchang, et al.
Published: (2026)
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
by: Ren, Xuanchi, et al.
Published: (2025)
by: Ren, Xuanchi, et al.
Published: (2025)
XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
by: Ren, Xuanchi, et al.
Published: (2023)
by: Ren, Xuanchi, et al.
Published: (2023)
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
by: Ren, Xuanchi, et al.
Published: (2025)
by: Ren, Xuanchi, et al.
Published: (2025)
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
by: Lu, Yifan, et al.
Published: (2024)
by: Lu, Yifan, et al.
Published: (2024)
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026)
by: Liu, Shaowei, et al.
Published: (2026)
DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
by: Bahmani, Sherwin, et al.
Published: (2025)
by: Bahmani, Sherwin, et al.
Published: (2025)
Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models
by: Ling, Huan, et al.
Published: (2023)
by: Ling, Huan, et al.
Published: (2023)
Augmented Reality based Simulated Data (ARSim) with multi-view consistency for AV perception networks
by: Anwar, Aqeel, et al.
Published: (2024)
by: Anwar, Aqeel, et al.
Published: (2024)
Controllable Weather Synthesis and Removal with Video Diffusion Models
by: Lin, Chih-Hao, et al.
Published: (2025)
by: Lin, Chih-Hao, et al.
Published: (2025)
DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
by: Sabour, Amirmojtaba, et al.
Published: (2025)
by: Sabour, Amirmojtaba, et al.
Published: (2025)
Align Your Steps: Optimizing Sampling Schedules in Diffusion Models
by: Sabour, Amirmojtaba, et al.
Published: (2024)
by: Sabour, Amirmojtaba, et al.
Published: (2024)
Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions
by: Zhao, Bo, et al.
Published: (2026)
by: Zhao, Bo, et al.
Published: (2026)
PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
by: Wang, Linqing, et al.
Published: (2025)
by: Wang, Linqing, et al.
Published: (2025)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
RadarGen: Automotive Radar Point Cloud Generation from Cameras
by: Borreda, Tomer, et al.
Published: (2025)
by: Borreda, Tomer, et al.
Published: (2025)
Photorealistic Object Insertion with Diffusion-Guided Inverse Rendering
by: Liang, Ruofan, et al.
Published: (2024)
by: Liang, Ruofan, et al.
Published: (2024)
Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?
by: Liao, Yuan-Hong, et al.
Published: (2024)
by: Liao, Yuan-Hong, et al.
Published: (2024)
CPA-Enhancer: Chain-of-Thought Prompted Adaptive Enhancer for Object Detection under Unknown Degradations
by: Zhang, Yuwei, et al.
Published: (2024)
by: Zhang, Yuwei, et al.
Published: (2024)
UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting
by: He, Kai, et al.
Published: (2025)
by: He, Kai, et al.
Published: (2025)
ViPE: Video Pose Engine for 3D Geometric Perception
by: Huang, Jiahui, et al.
Published: (2025)
by: Huang, Jiahui, et al.
Published: (2025)
fVDB: A Deep-Learning Framework for Sparse, Large-Scale, and High-Performance Spatial Intelligence
by: Williams, Francis, et al.
Published: (2024)
by: Williams, Francis, et al.
Published: (2024)
Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata
by: Zhang, Dongsu, et al.
Published: (2024)
by: Zhang, Dongsu, et al.
Published: (2024)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
by: Liao, Yuan-Hong, et al.
Published: (2024)
by: Liao, Yuan-Hong, et al.
Published: (2024)
EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models
by: Namekata, Koichi, et al.
Published: (2024)
by: Namekata, Koichi, et al.
Published: (2024)
Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
by: Liang, Hanxue, et al.
Published: (2024)
by: Liang, Hanxue, et al.
Published: (2024)
UNICE: Training A Universal Image Contrast Enhancer
by: Cui, Ruodai, et al.
Published: (2025)
by: Cui, Ruodai, et al.
Published: (2025)
L4GM: Large 4D Gaussian Reconstruction Model
by: Ren, Jiawei, et al.
Published: (2024)
by: Ren, Jiawei, et al.
Published: (2024)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
by: Xu, Zitong, et al.
Published: (2026)
by: Xu, Zitong, et al.
Published: (2026)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
by: Xu, Pengcheng, et al.
Published: (2025)
by: Xu, Pengcheng, et al.
Published: (2025)
PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
by: Ruan, Bo-Kai, et al.
Published: (2025)
by: Ruan, Bo-Kai, et al.
Published: (2025)
WildFusion: Learning 3D-Aware Latent Diffusion Models in View Space
by: Schwarz, Katja, et al.
Published: (2023)
by: Schwarz, Katja, et al.
Published: (2023)
PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding
by: Hou, Chenshu, et al.
Published: (2024)
by: Hou, Chenshu, et al.
Published: (2024)
Similar Items
-
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
by: Lu, Yifan, et al.
Published: (2026) -
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
by: Wu, Jay Zhangjie, et al.
Published: (2025) -
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
by: Wu, Jay Zhangjie, et al.
Published: (2025) -
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026) -
SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
by: Ren, Xuanchi, et al.
Published: (2024)