An Ensemble Model with Attention Based Mechanism for Image Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Badarneh, Israa Al, Hammo, Bassam, Al-Kadi, Omar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging
von: Chowdhury, Saniah Kayenat, et al.
Veröffentlicht: (2025)
von: Chowdhury, Saniah Kayenat, et al.
Veröffentlicht: (2025)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025)
Cross Modification Attention Based Deliberation Model for Image Captioning
von: Lian, Zheng, et al.
Veröffentlicht: (2021)
von: Lian, Zheng, et al.
Veröffentlicht: (2021)
Lesion-Aware Visual-Language Fusion for Automated Image Captioning of Ulcerative Colitis Endoscopic Examinations
von: Escamilla, Alexis Ivan Lopez, et al.
Veröffentlicht: (2025)
von: Escamilla, Alexis Ivan Lopez, et al.
Veröffentlicht: (2025)
CaptionFool: Universal Image Captioning Model Attacks
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
ResAF-Net: An Anchor-Free Attention-Based Network for Tree Detection and Agricultural Mapping in Palestine
von: Al-Qasem, Rabee
Veröffentlicht: (2026)
von: Al-Qasem, Rabee
Veröffentlicht: (2026)
Tracing 3D Anatomy in 2D Strokes: A Multi-Stage Projection Driven Approach to Cervical Spine Fracture Identification
von: Madhurja, Fabi Nahian, et al.
Veröffentlicht: (2026)
von: Madhurja, Fabi Nahian, et al.
Veröffentlicht: (2026)
Comparative Analysis of Deep Convolutional Neural Networks for Detecting Medical Image Deepfakes
von: Alsabbagh, Abdel Rahman, et al.
Veröffentlicht: (2024)
von: Alsabbagh, Abdel Rahman, et al.
Veröffentlicht: (2024)
Deep learning in computed tomography pulmonary angiography imaging: a dual-pronged approach for pulmonary embolism detection
von: Bushra, Fabiha, et al.
Veröffentlicht: (2023)
von: Bushra, Fabiha, et al.
Veröffentlicht: (2023)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning
von: Yang, Pu, et al.
Veröffentlicht: (2025)
von: Yang, Pu, et al.
Veröffentlicht: (2025)
Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
von: Dufera, Amanuel Tafese
Veröffentlicht: (2025)
von: Dufera, Amanuel Tafese
Veröffentlicht: (2025)
Is Your Text-to-Image Model Robust to Caption Noise?
von: Yu, Weichen, et al.
Veröffentlicht: (2024)
von: Yu, Weichen, et al.
Veröffentlicht: (2024)
IGAN: A New Inception-based Model for Stable and High-Fidelity Image Synthesis Using Generative Adversarial Networks
von: Hashim, Ahmed A., et al.
Veröffentlicht: (2026)
von: Hashim, Ahmed A., et al.
Veröffentlicht: (2026)
Inserting Faces inside Captions: Image Captioning with Attention Guided Merging
von: Tevissen, Yannis, et al.
Veröffentlicht: (2024)
von: Tevissen, Yannis, et al.
Veröffentlicht: (2024)
Image Captioning in news report scenario
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
Automated Image Captioning with CNNs and Transformers
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
von: Cahyono, Joshua Adrian, et al.
Veröffentlicht: (2024)
Deep Learning-Driven Segmentation of Ischemic Stroke Lesions Using Multi-Channel MRI
von: Rahman, Ashiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Ashiqur, et al.
Veröffentlicht: (2025)
Gender Stereotypes in Professional Roles Among Saudis: An Analytical Study of AI-Generated Images Using Language Models
von: AlKhalifah, Khaloud S., et al.
Veröffentlicht: (2025)
von: AlKhalifah, Khaloud S., et al.
Veröffentlicht: (2025)
Image Embedding Sampling Method for Diverse Captioning
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
von: Waheed, Sania, et al.
Veröffentlicht: (2025)
Top-Down Semantic Refinement for Image Captioning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
von: Aggarwal, Sajal, et al.
Veröffentlicht: (2024)
von: Aggarwal, Sajal, et al.
Veröffentlicht: (2024)
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
von: Ye, Shaokai, et al.
Veröffentlicht: (2026)
von: Ye, Shaokai, et al.
Veröffentlicht: (2026)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
von: Sharifzadeh, Sahand, et al.
Veröffentlicht: (2024)
von: Sharifzadeh, Sahand, et al.
Veröffentlicht: (2024)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
ReflectCAP: Detailed Image Captioning with Reflective Memory
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
Image Captions are Natural Prompts for Text-to-Image Models
von: Lei, Shiye, et al.
Veröffentlicht: (2023)
von: Lei, Shiye, et al.
Veröffentlicht: (2023)
Prompt-to-Prompt: Text-Based Image Editing Via Cross-Attention Mechanisms -- The Research of Hyperparameters and Novel Mechanisms to Enhance Existing Frameworks
von: Bieske, Linn, et al.
Veröffentlicht: (2025)
von: Bieske, Linn, et al.
Veröffentlicht: (2025)
Heat Diffusion Models -- Interpixel Attention Mechanism
von: Zhang, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhang, Pengfei, et al.
Veröffentlicht: (2025)
Provenance-Driven Reliable Semantic Medical Image Vector Reconstruction via Lightweight Blockchain-Verified Latent Fingerprints
von: Rasheed, Mohsin, et al.
Veröffentlicht: (2025)
von: Rasheed, Mohsin, et al.
Veröffentlicht: (2025)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
The Role of Data Curation in Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Uterine Ultrasound Image Captioning Using Deep Learning Techniques
von: Boulesnane, Abdennour, et al.
Veröffentlicht: (2024)
von: Boulesnane, Abdennour, et al.
Veröffentlicht: (2024)
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
von: Jiang, Yanbei, et al.
Veröffentlicht: (2024)
von: Jiang, Yanbei, et al.
Veröffentlicht: (2024)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
von: Pham, Anh-Cuong, et al.
Veröffentlicht: (2024)
von: Pham, Anh-Cuong, et al.
Veröffentlicht: (2024)
Intelligent Healthcare Imaging Platform: A VLM-Based Framework for Automated Medical Image Analysis and Clinical Report Generation
von: Al-Hamadani, Samer
Veröffentlicht: (2025)
von: Al-Hamadani, Samer
Veröffentlicht: (2025)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025) -
AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging
von: Chowdhury, Saniah Kayenat, et al.
Veröffentlicht: (2025) -
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
von: Teja, L. D. M. S. Sai, et al.
Veröffentlicht: (2025) -
Cross Modification Attention Based Deliberation Model for Image Captioning
von: Lian, Zheng, et al.
Veröffentlicht: (2021) -
Lesion-Aware Visual-Language Fusion for Automated Image Captioning of Ulcerative Colitis Endoscopic Examinations
von: Escamilla, Alexis Ivan Lopez, et al.
Veröffentlicht: (2025)