When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
Fuente:
arXiv
Saved in:
| Main Author: | Gupta, Hitesh Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extreme Blind Image Restoration via Prompt-Conditioned Information Bottleneck
by: Kim, Hongeun, et al.
Published: (2025)
by: Kim, Hongeun, et al.
Published: (2025)
ScribFormer: Transformer Makes CNN Work Better for Scribble-based Medical Image Segmentation
by: Li, Zihan, et al.
Published: (2024)
by: Li, Zihan, et al.
Published: (2024)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
ASTM :Autonomous Smart Traffic Management System Using Artificial Intelligence CNN and LSTM
by: Goenawan, Christofel Rio
Published: (2024)
by: Goenawan, Christofel Rio
Published: (2024)
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
Generalizable Geometric Image Caption Synthesis
by: Xin, Yue, et al.
Published: (2025)
by: Xin, Yue, et al.
Published: (2025)
Explanation Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
Evaluating BiLSTM and CNN+GRU Approaches for Human Activity Recognition Using WiFi CSI Data
by: Wakili, Almustapha A., et al.
Published: (2025)
by: Wakili, Almustapha A., et al.
Published: (2025)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
by: Dalvi, Abhishek, et al.
Published: (2026)
by: Dalvi, Abhishek, et al.
Published: (2026)
Editable Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
When Generative Augmentation Hurts: A Benchmark Study of GAN and Diffusion Models for Bias Correction in AI Classification Systems
by: Gupta, Shesh Narayan, et al.
Published: (2026)
by: Gupta, Shesh Narayan, et al.
Published: (2026)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024)
by: Lai, Zhengfeng, et al.
Published: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
by: Merchant, Nicholas, et al.
Published: (2025)
by: Merchant, Nicholas, et al.
Published: (2025)
Zero-shot Concept Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Process-Guided Concept Bottleneck Model
by: Asiyabi, Reza M., et al.
Published: (2026)
by: Asiyabi, Reza M., et al.
Published: (2026)
Semi-supervised Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
by: Yang, Chenglin, et al.
Published: (2023)
by: Yang, Chenglin, et al.
Published: (2023)
Semi-Supervised Image Captioning Considering Wasserstein Graph Matching
by: Yang, Yang
Published: (2024)
by: Yang, Yang
Published: (2024)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026)
by: Huang, Tzu-Heng, et al.
Published: (2026)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
Building Efficient Lightweight CNN Models
by: Isong, Nathan
Published: (2025)
by: Isong, Nathan
Published: (2025)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
This Looks Better than That: Better Interpretable Models with ProtoPNeXt
by: Willard, Frank, et al.
Published: (2024)
by: Willard, Frank, et al.
Published: (2024)
Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts
by: Choi, Jihye, et al.
Published: (2024)
by: Choi, Jihye, et al.
Published: (2024)
InstantIR: Blind Image Restoration with Instant Generative Reference
by: Huang, Jen-Yuan, et al.
Published: (2024)
by: Huang, Jen-Yuan, et al.
Published: (2024)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
by: Das, Swadhin, et al.
Published: (2025)
by: Das, Swadhin, et al.
Published: (2025)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
ChartEye: A Deep Learning Framework for Chart Information Extraction
by: Mustafa, Osama, et al.
Published: (2024)
by: Mustafa, Osama, et al.
Published: (2024)
Blind Inverse Problem Solving Made Easy by Text-to-Image Latent Diffusion
by: Dontas, Michail, et al.
Published: (2024)
by: Dontas, Michail, et al.
Published: (2024)
Optimal Eye Surgeon: Finding Image Priors through Sparse Generators at Initialization
by: Ghosh, Avrajit, et al.
Published: (2024)
by: Ghosh, Avrajit, et al.
Published: (2024)
Similar Items
-
Extreme Blind Image Restoration via Prompt-Conditioned Information Bottleneck
by: Kim, Hongeun, et al.
Published: (2025) -
ScribFormer: Transformer Makes CNN Work Better for Scribble-based Medical Image Segmentation
by: Li, Zihan, et al.
Published: (2024) -
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023) -
ASTM :Autonomous Smart Traffic Management System Using Artificial Intelligence CNN and LSTM
by: Goenawan, Christofel Rio
Published: (2024) -
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
by: Zhu, Qinfeng, et al.
Published: (2024)