Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
Fuente:
arXiv
Saved in:
| Main Authors: | Gordon, Brian, Bitton, Yonatan, Shafir, Yonatan, Garg, Roopal, Chen, Xi, Lischinski, Dani, Cohen-Or, Daniel, Szpektor, Idan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
by: Gordon, Brian, et al.
Published: (2025)
by: Gordon, Brian, et al.
Published: (2025)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
by: Slobodkin, Aviv, et al.
Published: (2025)
by: Slobodkin, Aviv, et al.
Published: (2025)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
by: Yosef, Ron, et al.
Published: (2025)
by: Yosef, Ron, et al.
Published: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
by: Ramos, Vasco, et al.
Published: (2024)
by: Ramos, Vasco, et al.
Published: (2024)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
by: Garg, Roopal, et al.
Published: (2024)
by: Garg, Roopal, et al.
Published: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Distinguishing Ignorance from Error in LLM Hallucinations
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
by: Bordalo, João, et al.
Published: (2024)
by: Bordalo, João, et al.
Published: (2024)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
DOCCI: Descriptions of Connected and Contrasting Images
by: Onoe, Yasumasa, et al.
Published: (2024)
by: Onoe, Yasumasa, et al.
Published: (2024)
NL-Eye: Abductive NLI for Images
by: Ventura, Mor, et al.
Published: (2024)
by: Ventura, Mor, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
by: Ventura, Mor, et al.
Published: (2026)
by: Ventura, Mor, et al.
Published: (2026)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
by: Kaplan, Guy, et al.
Published: (2025)
by: Kaplan, Guy, et al.
Published: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
by: Collins, Katherine M., et al.
Published: (2024)
by: Collins, Katherine M., et al.
Published: (2024)
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
by: Ventura, Mor, et al.
Published: (2025)
by: Ventura, Mor, et al.
Published: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
Seed-to-Seed: Image Translation in Diffusion Seed Space
by: Greenberg, Or, et al.
Published: (2024)
by: Greenberg, Or, et al.
Published: (2024)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
by: Sultan, Oren, et al.
Published: (2024)
by: Sultan, Oren, et al.
Published: (2024)
Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
by: Gordon, Ori, et al.
Published: (2023)
by: Gordon, Ori, et al.
Published: (2023)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
Latent Beam Diffusion Models for Generating Visual Sequences
by: Fernandes, Guilherme, et al.
Published: (2025)
by: Fernandes, Guilherme, et al.
Published: (2025)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
by: Cho, Jaemin, et al.
Published: (2023)
by: Cho, Jaemin, et al.
Published: (2023)
Cycle-Consistent Tuning for Layered Image Decomposition
by: Gu, Zheng, et al.
Published: (2026)
by: Gu, Zheng, et al.
Published: (2026)
EmoEdit: Evoking Emotions through Image Manipulation
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
by: Ramos, Vasco, et al.
Published: (2025)
by: Ramos, Vasco, et al.
Published: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
by: Orgad, Hadas, et al.
Published: (2024)
by: Orgad, Hadas, et al.
Published: (2024)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
by: Avrahami, Omri, et al.
Published: (2023)
by: Avrahami, Omri, et al.
Published: (2023)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
Palette Aligned Image Diffusion
by: Aharoni, Elad, et al.
Published: (2025)
by: Aharoni, Elad, et al.
Published: (2025)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Similar Items
-
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
by: Gordon, Brian, et al.
Published: (2025) -
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
by: Slobodkin, Aviv, et al.
Published: (2025) -
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024) -
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
by: Yosef, Ron, et al.
Published: (2025) -
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)