Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Zhicheng, Xie, Zhaotian, Ying, Huazhang, Zhang, Likun, Qin, Peiwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Modal interpretable automatic video captioning
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024)
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024)
Object-oriented backdoor attack against image captioning
von: Li, Meiling, et al.
Veröffentlicht: (2024)
von: Li, Meiling, et al.
Veröffentlicht: (2024)
MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
von: Du, Zhicheng, et al.
Veröffentlicht: (2025)
von: Du, Zhicheng, et al.
Veröffentlicht: (2025)
LAMPER: LanguAge Model and Prompt EngineeRing for zero-shot time series classification
von: Du, Zhicheng, et al.
Veröffentlicht: (2024)
von: Du, Zhicheng, et al.
Veröffentlicht: (2024)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
Image captioning for Brazilian Portuguese using GRIT model
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
Evaluating authenticity and quality of image captions via sentiment and semantic analyses
von: Krotov, Aleksei, et al.
Veröffentlicht: (2024)
von: Krotov, Aleksei, et al.
Veröffentlicht: (2024)
Stochastic positional embeddings improve masked image modeling
von: Bar, Amir, et al.
Veröffentlicht: (2023)
von: Bar, Amir, et al.
Veröffentlicht: (2023)
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
von: Kim, Yearim, et al.
Veröffentlicht: (2024)
von: Kim, Yearim, et al.
Veröffentlicht: (2024)
Semantic search for 100M+ galaxy images using AI-generated captions
von: Koblischke, Nolan, et al.
Veröffentlicht: (2025)
von: Koblischke, Nolan, et al.
Veröffentlicht: (2025)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
von: Liang, Yingshan, et al.
Veröffentlicht: (2025)
Explaining generative diffusion models via visual analysis for interpretable decision-making process
von: Park, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Park, Ji-Hoon, et al.
Veröffentlicht: (2024)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhui, et al.
Veröffentlicht: (2024)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
von: Shi, Danli, et al.
Veröffentlicht: (2024)
von: Shi, Danli, et al.
Veröffentlicht: (2024)
Symmetric masking strategy enhances the performance of Masked Image Modeling
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
An interpretable framework using foundation models for fish sex identification
von: Miao, Zheng, et al.
Veröffentlicht: (2026)
von: Miao, Zheng, et al.
Veröffentlicht: (2026)
CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation
von: Sharma, Sonali, et al.
Veröffentlicht: (2026)
von: Sharma, Sonali, et al.
Veröffentlicht: (2026)
Weakly Supervised Teacher-Student Framework with Progressive Pseudo-mask Refinement for Gland Segmentation
von: Khan, Hikmat, et al.
Veröffentlicht: (2026)
von: Khan, Hikmat, et al.
Veröffentlicht: (2026)
MedSAM-based lung masking for multi-label chest X-ray classification
von: Miao, Brayden, et al.
Veröffentlicht: (2025)
von: Miao, Brayden, et al.
Veröffentlicht: (2025)
Block-wise LoRA: Revisiting Fine-grained LoRA for Effective Personalization and Stylization in Text-to-Image Generation
von: Li, Likun, et al.
Veröffentlicht: (2024)
von: Li, Likun, et al.
Veröffentlicht: (2024)
AutoMR: A Universal Time Series Motion Recognition Pipeline
von: Zhang, Likun, et al.
Veröffentlicht: (2025)
von: Zhang, Likun, et al.
Veröffentlicht: (2025)
Quantifying the human visual exposome with vision language models
von: Rominger, Christian, et al.
Veröffentlicht: (2026)
von: Rominger, Christian, et al.
Veröffentlicht: (2026)
L-MAE: Longitudinal masked auto-encoder with time and severity-aware encoding for diabetic retinopathy progression prediction
von: Zeghlache, Rachid, et al.
Veröffentlicht: (2024)
von: Zeghlache, Rachid, et al.
Veröffentlicht: (2024)
Eye image segmentation using visual and concept prompts with Segment Anything Model 3 (SAM3)
von: Niehorster, Diederick C., et al.
Veröffentlicht: (2026)
von: Niehorster, Diederick C., et al.
Veröffentlicht: (2026)
ActiveMark: on watermarking of visual foundation models via massive activations
von: Chistyakova, Anna, et al.
Veröffentlicht: (2025)
von: Chistyakova, Anna, et al.
Veröffentlicht: (2025)
Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma
von: Pastor-Naranjo, Alvaro, et al.
Veröffentlicht: (2025)
von: Pastor-Naranjo, Alvaro, et al.
Veröffentlicht: (2025)
Vision language models are blind: Failing to translate detailed visual features into words
von: Rahmanzadehgervi, Pooyan, et al.
Veröffentlicht: (2024)
von: Rahmanzadehgervi, Pooyan, et al.
Veröffentlicht: (2024)
Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective
von: Weng, Zhaotian, et al.
Veröffentlicht: (2024)
von: Weng, Zhaotian, et al.
Veröffentlicht: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)
von: Fan, Zhaoyu, et al.
Veröffentlicht: (2025)
Leveraging image captions for selective whole slide image annotation
von: Qiu, Jingna, et al.
Veröffentlicht: (2024)
von: Qiu, Jingna, et al.
Veröffentlicht: (2024)
Wetland mapping from sparse annotations with satellite image time series and temporal-aware segment anything model
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
Memory augment is All You Need for image restoration
von: Zhang, Xiao Feng, et al.
Veröffentlicht: (2023)
von: Zhang, Xiao Feng, et al.
Veröffentlicht: (2023)
Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation
von: Liu, Kang, et al.
Veröffentlicht: (2025)
von: Liu, Kang, et al.
Veröffentlicht: (2025)
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
von: Jelaca, Aleksa, et al.
Veröffentlicht: (2025)
von: Jelaca, Aleksa, et al.
Veröffentlicht: (2025)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
Dilated Convolution with Learnable Spacings makes visual models more aligned with humans: a Grad-CAM study
von: Chamas, Rabih, et al.
Veröffentlicht: (2024)
von: Chamas, Rabih, et al.
Veröffentlicht: (2024)
Integrated feature analysis for deep learning interpretation and class activation maps
von: Li, Yanli, et al.
Veröffentlicht: (2024)
von: Li, Yanli, et al.
Veröffentlicht: (2024)
Enhancing Interpretability of Vertebrae Fracture Grading using Human-interpretable Prototypes
von: Sinhamahapatra, Poulami, et al.
Veröffentlicht: (2024)
von: Sinhamahapatra, Poulami, et al.
Veröffentlicht: (2024)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Modal interpretable automatic video captioning
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024) -
Object-oriented backdoor attack against image captioning
von: Li, Meiling, et al.
Veröffentlicht: (2024) -
MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
von: Du, Zhicheng, et al.
Veröffentlicht: (2025) -
LAMPER: LanguAge Model and Prompt EngineeRing for zero-shot time series classification
von: Du, Zhicheng, et al.
Veröffentlicht: (2024) -
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)