Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Albadarneh, Israa A., Hammo, Bassam H., Al-Kadi, Omar S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Ensemble Model with Attention Based Mechanism for Image Captioning
by: Badarneh, Israa Al, et al.
Published: (2025)
by: Badarneh, Israa Al, et al.
Published: (2025)
Image captioning in different languages
by: van Miltenburg, Emiel
Published: (2024)
by: van Miltenburg, Emiel
Published: (2024)
Image captioning for Brazilian Portuguese using GRIT model
by: de Alencar, Rafael Silva, et al.
Published: (2024)
by: de Alencar, Rafael Silva, et al.
Published: (2024)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
by: Zhang, Chenhui, et al.
Published: (2024)
by: Zhang, Chenhui, et al.
Published: (2024)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)
by: Padlewski, Piotr, et al.
Published: (2024)
Object-oriented backdoor attack against image captioning
by: Li, Meiling, et al.
Published: (2024)
by: Li, Meiling, et al.
Published: (2024)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024)
by: Du, Zhicheng, et al.
Published: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
PathAlign: A vision-language model for whole slide images in histopathology
by: Ahmed, Faruk, et al.
Published: (2024)
by: Ahmed, Faruk, et al.
Published: (2024)
Deep learning in computed tomography pulmonary angiography imaging: a dual-pronged approach for pulmonary embolism detection
by: Bushra, Fabiha, et al.
Published: (2023)
by: Bushra, Fabiha, et al.
Published: (2023)
Multi-Modal interpretable automatic video captioning
by: Hanna-Asaad, Antoine, et al.
Published: (2024)
by: Hanna-Asaad, Antoine, et al.
Published: (2024)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models
by: Aman, Tabinda, et al.
Published: (2025)
by: Aman, Tabinda, et al.
Published: (2025)
AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging
by: Chowdhury, Saniah Kayenat, et al.
Published: (2025)
by: Chowdhury, Saniah Kayenat, et al.
Published: (2025)
Comparative Analysis of Deep Convolutional Neural Networks for Detecting Medical Image Deepfakes
by: Alsabbagh, Abdel Rahman, et al.
Published: (2024)
by: Alsabbagh, Abdel Rahman, et al.
Published: (2024)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
by: Tu, Rong-Cheng, et al.
Published: (2024)
by: Tu, Rong-Cheng, et al.
Published: (2024)
Evaluating authenticity and quality of image captions via sentiment and semantic analyses
by: Krotov, Aleksei, et al.
Published: (2024)
by: Krotov, Aleksei, et al.
Published: (2024)
Leveraging image captions for selective whole slide image annotation
by: Qiu, Jingna, et al.
Published: (2024)
by: Qiu, Jingna, et al.
Published: (2024)
Multimodal Evaluation of Russian-language Architectures
by: Chervyakov, Artem, et al.
Published: (2025)
by: Chervyakov, Artem, et al.
Published: (2025)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Tracing 3D Anatomy in 2D Strokes: A Multi-Stage Projection Driven Approach to Cervical Spine Fracture Identification
by: Madhurja, Fabi Nahian, et al.
Published: (2026)
by: Madhurja, Fabi Nahian, et al.
Published: (2026)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
by: Khan, Shaharukh, et al.
Published: (2026)
by: Khan, Shaharukh, et al.
Published: (2026)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
by: Zhao, Haozhe, et al.
Published: (2023)
by: Zhao, Haozhe, et al.
Published: (2023)
Learning Style Identification Using Semi-Supervised Self-Taught Labeling
by: Ayyoub, Hani Y., et al.
Published: (2024)
by: Ayyoub, Hani Y., et al.
Published: (2024)
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
by: Xue, Yufei, et al.
Published: (2025)
by: Xue, Yufei, et al.
Published: (2025)
Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models
by: AlKhalifah, Khaloud S., et al.
Published: (2025)
by: AlKhalifah, Khaloud S., et al.
Published: (2025)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
Advancing vision-language models in front-end development via data synthesis
by: Ge, Tong, et al.
Published: (2025)
by: Ge, Tong, et al.
Published: (2025)
Sparser Block-Sparse Attention via Token Permutation
by: Wang, Xinghao, et al.
Published: (2025)
by: Wang, Xinghao, et al.
Published: (2025)
Towards Deployable OCR models for Indic languages
by: Mathew, Minesh, et al.
Published: (2022)
by: Mathew, Minesh, et al.
Published: (2022)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
by: Zhou, Ziwei, et al.
Published: (2025)
by: Zhou, Ziwei, et al.
Published: (2025)
BrainChat: Decoding Semantic Information from fMRI using Vision-language Pretrained Models
by: Huang, Wanaiu
Published: (2024)
by: Huang, Wanaiu
Published: (2024)
Faithful Attention Explainer: Verbalizing Decisions Based on Discriminative Features
by: Rong, Yao, et al.
Published: (2024)
by: Rong, Yao, et al.
Published: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning
by: Fauzulhaq, Alfirsa Damasyifa, et al.
Published: (2024)
by: Fauzulhaq, Alfirsa Damasyifa, et al.
Published: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
Similar Items
-
An Ensemble Model with Attention Based Mechanism for Image Captioning
by: Badarneh, Israa Al, et al.
Published: (2025) -
Image captioning in different languages
by: van Miltenburg, Emiel
Published: (2024) -
Image captioning for Brazilian Portuguese using GRIT model
by: de Alencar, Rafael Silva, et al.
Published: (2024) -
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
by: Zhang, Chenhui, et al.
Published: (2024) -
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)