The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
Fuente:
arXiv
Guardado en:
| Autores principales: | Kodathala, Sai Varun, Vunnam, Rakesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
por: Kodathala, Sai Varun, et al.
Publicado: (2025)
por: Kodathala, Sai Varun, et al.
Publicado: (2025)
LLMs can Compress LLMs: Adaptive Pruning by Agents
por: Kodathala, Sai Varun, et al.
Publicado: (2026)
por: Kodathala, Sai Varun, et al.
Publicado: (2026)
SV3.3B: A Sports Video Understanding Model for Action Recognition
por: Kodathala, Sai Varun, et al.
Publicado: (2025)
por: Kodathala, Sai Varun, et al.
Publicado: (2025)
Fast OTSU Thresholding Using Bisection Method
por: Kodathala, Sai Varun
Publicado: (2025)
por: Kodathala, Sai Varun
Publicado: (2025)
Six Sigma For Neural Networks: Taguchi-based optimization
por: Kodathala, Sai Varun
Publicado: (2025)
por: Kodathala, Sai Varun
Publicado: (2025)
Can Large Language Models Solve Engineering Equations? A Systematic Comparison of Direct Prediction and Solver-Assisted Approaches
por: Kodathala, Sai Varun, et al.
Publicado: (2026)
por: Kodathala, Sai Varun, et al.
Publicado: (2026)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
por: He, Chiyuan, et al.
Publicado: (2025)
por: He, Chiyuan, et al.
Publicado: (2025)
PSA-VLM: Enhancing Vision-Language Model Safety through Progressive Concept-Bottleneck-Driven Alignment
por: Liu, Zhendong, et al.
Publicado: (2024)
por: Liu, Zhendong, et al.
Publicado: (2024)
Describe Anything: Detailed Localized Image and Video Captioning
por: Lian, Long, et al.
Publicado: (2025)
por: Lian, Long, et al.
Publicado: (2025)
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
por: Huang, Peng, et al.
Publicado: (2025)
por: Huang, Peng, et al.
Publicado: (2025)
CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches
por: Wu, Sifan, et al.
Publicado: (2024)
por: Wu, Sifan, et al.
Publicado: (2024)
VLM6D: VLM based 6Dof Pose Estimation based on RGB-D Images
por: Sarowar, Md Selim, et al.
Publicado: (2025)
por: Sarowar, Md Selim, et al.
Publicado: (2025)
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
por: Feiding, et al.
Publicado: (2026)
por: Feiding, et al.
Publicado: (2026)
CoBELa: Steering Transparent Generation via Concept Bottlenecks on Energy Landscapes
por: Kim, Sangwon, et al.
Publicado: (2025)
por: Kim, Sangwon, et al.
Publicado: (2025)
Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
por: Agarwal, Lakshita, et al.
Publicado: (2025)
por: Agarwal, Lakshita, et al.
Publicado: (2025)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
por: An, Na Min, et al.
Publicado: (2025)
por: An, Na Min, et al.
Publicado: (2025)
DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editing
por: Das, Gyanendra, et al.
Publicado: (2026)
por: Das, Gyanendra, et al.
Publicado: (2026)
Look, Recite, Then Answer: Enhancing VLM Performance via Self-Generated Knowledge Hints
por: Feng, Xisheng
Publicado: (2025)
por: Feng, Xisheng
Publicado: (2025)
FLAIR: VLM with Fine-grained Language-informed Image Representations
por: Xiao, Rui, et al.
Publicado: (2024)
por: Xiao, Rui, et al.
Publicado: (2024)
Concept Complement Bottleneck Model for Interpretable Medical Image Diagnosis
por: Wang, Hongmei, et al.
Publicado: (2024)
por: Wang, Hongmei, et al.
Publicado: (2024)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
por: Fu, Yongjie, et al.
Publicado: (2024)
por: Fu, Yongjie, et al.
Publicado: (2024)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
por: Park, Jun-Hyung, et al.
Publicado: (2024)
por: Park, Jun-Hyung, et al.
Publicado: (2024)
Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability
por: Zhu, Zhiyu, et al.
Publicado: (2025)
por: Zhu, Zhiyu, et al.
Publicado: (2025)
DRCT: Saving Image Super-resolution away from Information Bottleneck
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
por: Chen, Haoyu, et al.
Publicado: (2026)
por: Chen, Haoyu, et al.
Publicado: (2026)
Real-Time Drowsiness Detection Using Eye Aspect Ratio and Facial Landmark Detection
por: Rupani, Varun Shiva Krishna, et al.
Publicado: (2024)
por: Rupani, Varun Shiva Krishna, et al.
Publicado: (2024)
Contact-aware Human Motion Generation from Textual Descriptions
por: Ma, Sihan, et al.
Publicado: (2024)
por: Ma, Sihan, et al.
Publicado: (2024)
Guess the Unified Model: How Much Can We Recover from Generated Images?
por: Cekinmez, Jasin, et al.
Publicado: (2026)
por: Cekinmez, Jasin, et al.
Publicado: (2026)
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
por: Qian, Chen, et al.
Publicado: (2025)
por: Qian, Chen, et al.
Publicado: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
por: Singh, Aditya Kumar, et al.
Publicado: (2026)
por: Singh, Aditya Kumar, et al.
Publicado: (2026)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
por: Vu, Tuan-Anh, et al.
Publicado: (2023)
por: Vu, Tuan-Anh, et al.
Publicado: (2023)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
por: Quang, Ngoc Bui Lam, et al.
Publicado: (2025)
por: Quang, Ngoc Bui Lam, et al.
Publicado: (2025)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
por: Shukla, Shreya, et al.
Publicado: (2025)
por: Shukla, Shreya, et al.
Publicado: (2025)
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
por: Patel, Shivansh, et al.
Publicado: (2025)
por: Patel, Shivansh, et al.
Publicado: (2025)
A Recipe for Improving Remote Sensing VLM Zero Shot Generalization
por: Barzilai, Aviad, et al.
Publicado: (2025)
por: Barzilai, Aviad, et al.
Publicado: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
por: Wang, Xudong, et al.
Publicado: (2024)
por: Wang, Xudong, et al.
Publicado: (2024)
Describe Anything Anywhere At Any Moment
por: Gorlo, Nicolas, et al.
Publicado: (2025)
por: Gorlo, Nicolas, et al.
Publicado: (2025)
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
por: Bandyopadhyay, Hmrishav, et al.
Publicado: (2025)
por: Bandyopadhyay, Hmrishav, et al.
Publicado: (2025)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
por: Wang, Zhenting, et al.
Publicado: (2024)
por: Wang, Zhenting, et al.
Publicado: (2024)
Overcoming the Curvature Bottleneck in MeanFlow
por: Zhang, Xinxi, et al.
Publicado: (2025)
por: Zhang, Xinxi, et al.
Publicado: (2025)
Ejemplares similares
-
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
por: Kodathala, Sai Varun, et al.
Publicado: (2025) -
LLMs can Compress LLMs: Adaptive Pruning by Agents
por: Kodathala, Sai Varun, et al.
Publicado: (2026) -
SV3.3B: A Sports Video Understanding Model for Action Recognition
por: Kodathala, Sai Varun, et al.
Publicado: (2025) -
Fast OTSU Thresholding Using Bisection Method
por: Kodathala, Sai Varun
Publicado: (2025) -
Six Sigma For Neural Networks: Taguchi-based optimization
por: Kodathala, Sai Varun
Publicado: (2025)