VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Meng, Yang, Senqiao, Che, Haoxuan, Zhang, Suiyun, Zhang, Xichen, Yu, Shaozuo, Gui, Haokun, Rao, Zhefan, Tu, Dandan, Liu, Rui, Jia, Jiaya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
by: Zhang, Xichen, et al.
Published: (2026)
by: Zhang, Xichen, et al.
Published: (2026)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
by: Yang, Senqiao, et al.
Published: (2025)
by: Yang, Senqiao, et al.
Published: (2025)
SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
by: Yang, Senqiao, et al.
Published: (2024)
by: Yang, Senqiao, et al.
Published: (2024)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
by: Qu, Tianyuan, et al.
Published: (2025)
by: Qu, Tianyuan, et al.
Published: (2025)
Unified Language-driven Zero-shot Domain Adaptation
by: Yang, Senqiao, et al.
Published: (2024)
by: Yang, Senqiao, et al.
Published: (2024)
MOODv2: Masked Image Modeling for Out-of-Distribution Detection
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
by: Wang, Chengyao, et al.
Published: (2025)
by: Wang, Chengyao, et al.
Published: (2025)
MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
by: Wang, Yingying, et al.
Published: (2025)
by: Wang, Yingying, et al.
Published: (2025)
InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation
by: Rao, Zhefan, et al.
Published: (2026)
by: Rao, Zhefan, et al.
Published: (2026)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language
by: Wu, Qiwei, et al.
Published: (2026)
by: Wu, Qiwei, et al.
Published: (2026)
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
by: Chen, Jinpeng, et al.
Published: (2026)
by: Chen, Jinpeng, et al.
Published: (2026)
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
by: Chu, Meng, et al.
Published: (2026)
by: Chu, Meng, et al.
Published: (2026)
VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving
by: Xu, Zhefan, et al.
Published: (2026)
by: Xu, Zhefan, et al.
Published: (2026)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
by: Zhang, Shaozuo, et al.
Published: (2025)
by: Zhang, Shaozuo, et al.
Published: (2025)
Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models
by: Liu, Xiao, et al.
Published: (2026)
by: Liu, Xiao, et al.
Published: (2026)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
by: Liu, Ziru, et al.
Published: (2025)
by: Liu, Ziru, et al.
Published: (2025)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
by: Wu, Junfei, et al.
Published: (2024)
by: Wu, Junfei, et al.
Published: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
Closed-Loop Vision-Language Planning for Multi-Agent Coordination
by: Li, Zhiyuan, et al.
Published: (2025)
by: Li, Zhiyuan, et al.
Published: (2025)
Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving
by: Zhang, Tianyuan, et al.
Published: (2025)
by: Zhang, Tianyuan, et al.
Published: (2025)
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2025)
by: Chen, Kecheng, et al.
Published: (2025)
CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
by: Zhong, Zhisheng, et al.
Published: (2024)
by: Zhong, Zhisheng, et al.
Published: (2024)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
by: Babu, Abhijith, et al.
Published: (2026)
by: Babu, Abhijith, et al.
Published: (2026)
Bionic Vision as Neuroadaptive XR: Closed-Loop Perceptual Interfaces for Neurotechnology
by: Beyeler, Michael
Published: (2025)
by: Beyeler, Michael
Published: (2025)
Vision-Guided Iterative Refinement for Frontend Code Generation
by: Sansford, Hannah, et al.
Published: (2026)
by: Sansford, Hannah, et al.
Published: (2026)
Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
by: Shen, Boyang, et al.
Published: (2026)
by: Shen, Boyang, et al.
Published: (2026)
EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
ENTED: Enhanced Neural Texture Extraction and Distribution for Reference-based Blind Face Restoration
by: Lau, Yuen-Fui, et al.
Published: (2024)
by: Lau, Yuen-Fui, et al.
Published: (2024)
From Literature to Lab: Closed-Loop Advancement of Perovskite Solar Cells via Domain Knowledge Guided LLM
by: Sun, Penglei, et al.
Published: (2026)
by: Sun, Penglei, et al.
Published: (2026)
Discovering Closed-Loop Failures of Vision-Based Controllers via Reachability Analysis
by: Chakraborty, Kaustav, et al.
Published: (2022)
by: Chakraborty, Kaustav, et al.
Published: (2022)
Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models
by: Uzun, Tolgay Atinc, et al.
Published: (2026)
by: Uzun, Tolgay Atinc, et al.
Published: (2026)
Similar Items
-
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
by: Chu, Meng, et al.
Published: (2025) -
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
by: Zhang, Xichen, et al.
Published: (2026) -
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
by: Yang, Senqiao, et al.
Published: (2025) -
SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
by: Zhang, Xichen, et al.
Published: (2025) -
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)