Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Junrong, Fang, Shancheng, Qu, Yadong, Xie, Hongtao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
by: Fu, Fengyi, et al.
Published: (2024)
by: Fu, Fengyi, et al.
Published: (2024)
IGD: Instructional Graphic Design with Multimodal Layer Generation
by: Qu, Yadong, et al.
Published: (2025)
by: Qu, Yadong, et al.
Published: (2025)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025)
by: Xu, Chengzhi, et al.
Published: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
Iterative Feedback Network for Unsupervised Point Cloud Registration
by: Xie, Yifan, et al.
Published: (2024)
by: Xie, Yifan, et al.
Published: (2024)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
How Control Information Influences Multilingual Text Image Generation and Editing?
by: Zhang, Boqiang, et al.
Published: (2024)
by: Zhang, Boqiang, et al.
Published: (2024)
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026)
by: Jaiswal, Shantanu, et al.
Published: (2026)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
by: Li, Aiden Yiliu, et al.
Published: (2025)
by: Li, Aiden Yiliu, et al.
Published: (2025)
SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
SeeingSounds: Learning Audio-to-Visual Alignment via Text
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
by: Liu, Guimeng, et al.
Published: (2025)
by: Liu, Guimeng, et al.
Published: (2025)
IRNet: Iterative Refinement Network for Noisy Partial Label Learning
by: Lian, Zheng, et al.
Published: (2022)
by: Lian, Zheng, et al.
Published: (2022)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
by: Zheng, Zirui, et al.
Published: (2025)
by: Zheng, Zirui, et al.
Published: (2025)
InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding
by: Zhang, Huaxiang, et al.
Published: (2024)
by: Zhang, Huaxiang, et al.
Published: (2024)
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
by: Choi, Minkyu, et al.
Published: (2025)
by: Choi, Minkyu, et al.
Published: (2025)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
by: Xiao, Xin, et al.
Published: (2024)
by: Xiao, Xin, et al.
Published: (2024)
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
by: Song, Shezheng, et al.
Published: (2026)
by: Song, Shezheng, et al.
Published: (2026)
VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
by: Huy, Ta Duc, et al.
Published: (2025)
by: Huy, Ta Duc, et al.
Published: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
by: Brioschi, Riccardo, et al.
Published: (2025)
by: Brioschi, Riccardo, et al.
Published: (2025)
Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion
by: Tang, Zhenggang, et al.
Published: (2026)
by: Tang, Zhenggang, et al.
Published: (2026)
LTOS: Layout-controllable Text-Object Synthesis via Adaptive Cross-attention Fusions
by: Zhao, Xiaoran, et al.
Published: (2024)
by: Zhao, Xiaoran, et al.
Published: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
by: Fhima, Jonathan, et al.
Published: (2024)
by: Fhima, Jonathan, et al.
Published: (2024)
Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
by: Wan, Hanwen, et al.
Published: (2025)
by: Wan, Hanwen, et al.
Published: (2025)
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
by: Lou, Ange, et al.
Published: (2026)
by: Lou, Ange, et al.
Published: (2026)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
See Through the Noise: Improving Domain Generalization in Gaze Estimation
by: Peng, Yanming, et al.
Published: (2026)
by: Peng, Yanming, et al.
Published: (2026)
Layout-and-Retouch: A Dual-stage Framework for Improving Diversity in Personalized Image Generation
by: Kim, Kangyeol, et al.
Published: (2024)
by: Kim, Kangyeol, et al.
Published: (2024)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
by: Xu, Ziqiang, et al.
Published: (2025)
by: Xu, Ziqiang, et al.
Published: (2025)
IVLMap: Instance-Aware Visual Language Grounding for Consumer Robot Navigation
by: Huang, Jiacui, et al.
Published: (2024)
by: Huang, Jiacui, et al.
Published: (2024)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
by: Cho, Jaemin, et al.
Published: (2023)
by: Cho, Jaemin, et al.
Published: (2023)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
Similar Items
-
Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
by: Fu, Fengyi, et al.
Published: (2024) -
IGD: Instructional Graphic Design with Multimodal Layer Generation
by: Qu, Yadong, et al.
Published: (2025) -
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025) -
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025) -
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)