ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Xingyu, Liu, Minqian, Yang, Zhengyuan, Corring, John, Lu, Yijuan, Yang, Jianwei, Roth, Dan, Florencio, Dinei, Zhang, Cha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RCEM: Embedder Equipped with Query Rewriting Skill for Robust Conversational Search in Distributional Shift
by: Son, Kilho, et al.
Published: (2026)
by: Son, Kilho, et al.
Published: (2026)
ReFocus
by: Ostrowska, Elzbieta
Published: (2025)
by: Ostrowska, Elzbieta
Published: (2025)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
by: Hu, Yushi, et al.
Published: (2024)
by: Hu, Yushi, et al.
Published: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
by: Fu, Xingyu, et al.
Published: (2024)
by: Fu, Xingyu, et al.
Published: (2024)
Generative Visual Chain-of-Thought for Image Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
by: Zhao, Kesen, et al.
Published: (2026)
by: Zhao, Kesen, et al.
Published: (2026)
Chain-of-Thought Re-ranking for Image Retrieval Tasks
by: Wu, Shangrong, et al.
Published: (2025)
by: Wu, Shangrong, et al.
Published: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
by: Cheng, Zihui, et al.
Published: (2025)
by: Cheng, Zihui, et al.
Published: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
by: Liao, Jiaqi, et al.
Published: (2025)
by: Liao, Jiaqi, et al.
Published: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
Reinforcing Structured Chain-of-Thought for Video Understanding
by: Wang, Peiyao, et al.
Published: (2026)
by: Wang, Peiyao, et al.
Published: (2026)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
Knowledge Editing through Chain-of-Thought
by: Wang, Changyue, et al.
Published: (2024)
by: Wang, Changyue, et al.
Published: (2024)
S-Chain: Structured Visual Chain-of-Thought For Medicine
by: Le-Duc, Khai, et al.
Published: (2025)
by: Le-Duc, Khai, et al.
Published: (2025)
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
by: Rose, Daniel, et al.
Published: (2023)
by: Rose, Daniel, et al.
Published: (2023)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
by: Guo, Guangfu, et al.
Published: (2026)
by: Guo, Guangfu, et al.
Published: (2026)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
by: Li, Bangzheng, et al.
Published: (2023)
by: Li, Bangzheng, et al.
Published: (2023)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
by: Struppek, Lukas, et al.
Published: (2025)
by: Struppek, Lukas, et al.
Published: (2025)
DataProphet: Demystifying Supervision Data Generalization in Multimodal LLMs
by: Qi, Xuan, et al.
Published: (2026)
by: Qi, Xuan, et al.
Published: (2026)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
ReFocus: Reinforcing Mid-Frequency and Key-Frequency Modeling for Multivariate Time Series Forecasting
by: Yu, Guoqi, et al.
Published: (2025)
by: Yu, Guoqi, et al.
Published: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
by: Shi, Weikang, et al.
Published: (2025)
by: Shi, Weikang, et al.
Published: (2025)
Why Chain of Thought Fails in Clinical Text Understanding
by: Wu, Jiageng, et al.
Published: (2025)
by: Wu, Jiageng, et al.
Published: (2025)
Understanding Chain-of-Thought in LLMs through Information Theory
by: Ton, Jean-Francois, et al.
Published: (2024)
by: Ton, Jean-Francois, et al.
Published: (2024)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
by: Cha, Junuk, et al.
Published: (2025)
by: Cha, Junuk, et al.
Published: (2025)
Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought
by: Huo, Yu, et al.
Published: (2026)
by: Huo, Yu, et al.
Published: (2026)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
by: Ni, Minheng, et al.
Published: (2026)
by: Ni, Minheng, et al.
Published: (2026)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
by: Chen, Qiguang, et al.
Published: (2026)
by: Chen, Qiguang, et al.
Published: (2026)
RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
by: Yang, Junyao, et al.
Published: (2025)
by: Yang, Junyao, et al.
Published: (2025)
Self-Harmonized Chain of Thought
by: Jin, Ziqi, et al.
Published: (2024)
by: Jin, Ziqi, et al.
Published: (2024)
Psy-Copilot: Visual Chain of Thought for Counseling
by: Chen, Keqi, et al.
Published: (2025)
by: Chen, Keqi, et al.
Published: (2025)
Latent Chain-of-Thought for Visual Reasoning
by: Sun, Guohao, et al.
Published: (2025)
by: Sun, Guohao, et al.
Published: (2025)
Spatial Chain-of-Thought: Bridging Understanding and Generation Models for Spatial Reasoning Generation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
by: He, Runze, et al.
Published: (2026)
by: He, Runze, et al.
Published: (2026)
Understanding Hidden Computations in Chain-of-Thought Reasoning
by: Bharadwaj, Aryasomayajula Ram
Published: (2024)
by: Bharadwaj, Aryasomayajula Ram
Published: (2024)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
ERA-CoT: Improving Chain-of-Thought through Entity Relationship Analysis
by: Liu, Yanming, et al.
Published: (2024)
by: Liu, Yanming, et al.
Published: (2024)
Similar Items
-
RCEM: Embedder Equipped with Query Rewriting Skill for Robust Conversational Search in Distributional Shift
by: Son, Kilho, et al.
Published: (2026) -
ReFocus
by: Ostrowska, Elzbieta
Published: (2025) -
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
by: Hu, Yushi, et al.
Published: (2024) -
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
by: Fu, Xingyu, et al.
Published: (2024) -
Generative Visual Chain-of-Thought for Image Editing
by: Yin, Zijin, et al.
Published: (2026)