Vision-Language Binding in In-Context Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Chris, Gandikota, Rohit, Torralba, Antonio, Shaham, Tamar Rott |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
by: Cui, Kelly, et al.
Published: (2026)
by: Cui, Kelly, et al.
Published: (2026)
SketchAgent: Language-Driven Sequential Sketch Generation
by: Vinker, Yael, et al.
Published: (2024)
by: Vinker, Yael, et al.
Published: (2024)
A Vision Check-up for Language Models
by: Sharma, Pratyusha, et al.
Published: (2024)
by: Sharma, Pratyusha, et al.
Published: (2024)
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
by: Golbari, Yuval, et al.
Published: (2026)
by: Golbari, Yuval, et al.
Published: (2026)
Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
by: Li, Christy, et al.
Published: (2025)
by: Li, Christy, et al.
Published: (2025)
Opt-In Art: Learning Art Styles Only from Few Examples
by: Ren, Hui, et al.
Published: (2024)
by: Ren, Hui, et al.
Published: (2024)
A Multimodal Automated Interpretability Agent
by: Shaham, Tamar Rott, et al.
Published: (2024)
by: Shaham, Tamar Rott, et al.
Published: (2024)
Distilling Diversity and Control in Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2025)
by: Gandikota, Rohit, et al.
Published: (2025)
MultiModal Action Conditioned Video Generation
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Text-guided Explorable Image Super-resolution
by: Gandikota, Kanchana Vaishnavi, et al.
Published: (2024)
by: Gandikota, Kanchana Vaishnavi, et al.
Published: (2024)
Investigating Mechanisms for In-Context Vision Language Binding
by: Saravanan, Darshana, et al.
Published: (2025)
by: Saravanan, Darshana, et al.
Published: (2025)
LDEdit: Towards Generalized Text Guided Image Manipulation via Latent Diffusion Models
by: Chandramouli, Paramanand, et al.
Published: (2022)
by: Chandramouli, Paramanand, et al.
Published: (2022)
Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning
by: Liu, Shih-Wen, et al.
Published: (2025)
by: Liu, Shih-Wen, et al.
Published: (2025)
RealStats: A Rigorous Real-Only Statistical Framework for Fake Image Detection
by: Zisman, Haim, et al.
Published: (2026)
by: Zisman, Haim, et al.
Published: (2026)
Seeing Through the PRISM: Compound & Controllable Restoration of Scientific Images
by: Kurinchi-Vendhan, Rupa, et al.
Published: (2026)
by: Kurinchi-Vendhan, Rupa, et al.
Published: (2026)
Unified Concept Editing in Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2023)
by: Gandikota, Rohit, et al.
Published: (2023)
Context-Aware Decoding for Faithful Vision-Language Generation
by: Fazli, Mehrdad, et al.
Published: (2026)
by: Fazli, Mehrdad, et al.
Published: (2026)
Generalized Dynamics Generation towards Scannable Physical World Model
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
Cropper: Vision-Language Model for Image Cropping through In-Context Learning
by: Lee, Seung Hyun, et al.
Published: (2024)
by: Lee, Seung Hyun, et al.
Published: (2024)
Navigating Text-to-Image Generative Bias across Indic Languages
by: Mittal, Surbhi, et al.
Published: (2024)
by: Mittal, Surbhi, et al.
Published: (2024)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
by: Wang, Zhaowei, et al.
Published: (2026)
by: Wang, Zhaowei, et al.
Published: (2026)
RAWDet-7: A Multi-Scenario Benchmark for Object Detection and Description on Quantized RAW Images
by: Fatima, Mishal, et al.
Published: (2026)
by: Fatima, Mishal, et al.
Published: (2026)
Efficient 3D Instance Mapping and Localization with Neural Fields
by: Tang, George, et al.
Published: (2024)
by: Tang, George, et al.
Published: (2024)
Evaluating Vision-Language Models on Bistable Images
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
by: Wang, Zhaowei, et al.
Published: (2025)
by: Wang, Zhaowei, et al.
Published: (2025)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Enhancing VICReg: Random-Walk Pairing for Improved Generalization and Better Global Semantics Capturing
by: Simai, Idan, et al.
Published: (2025)
by: Simai, Idan, et al.
Published: (2025)
Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
Long Context Transfer from Language to Vision
by: Zhang, Peiyuan, et al.
Published: (2024)
by: Zhang, Peiyuan, et al.
Published: (2024)
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
Context Diffusion: In-Context Aware Image Generation
by: Najdenkoska, Ivona, et al.
Published: (2023)
by: Najdenkoska, Ivona, et al.
Published: (2023)
Enhancing Vision-Language Models Generalization via Diversity-Driven Novel Feature Synthesis
by: Yan, Siyuan, et al.
Published: (2024)
by: Yan, Siyuan, et al.
Published: (2024)
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
by: Chavan, Jitesh, et al.
Published: (2025)
by: Chavan, Jitesh, et al.
Published: (2025)
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2025)
by: Gandikota, Rohit, et al.
Published: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
Similar Items
-
The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
by: Cui, Kelly, et al.
Published: (2026) -
SketchAgent: Language-Driven Sequential Sketch Generation
by: Vinker, Yael, et al.
Published: (2024) -
A Vision Check-up for Language Models
by: Sharma, Pratyusha, et al.
Published: (2024) -
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
by: Wasserman, Navve, et al.
Published: (2025) -
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
by: Golbari, Yuval, et al.
Published: (2026)