Automated Image Captioning with CNNs and Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Cahyono, Joshua Adrian, Jusuf, Jeremy Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
by: Kolouju, Pranavi, et al.
Published: (2025)
by: Kolouju, Pranavi, et al.
Published: (2025)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
The Effective Depth Paradox: Evaluating the Relationship between Architectural Topology and Trainability in Deep CNNs
by: Fischer, Manfred M., et al.
Published: (2026)
by: Fischer, Manfred M., et al.
Published: (2026)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
by: Dufera, Amanuel Tafese
Published: (2025)
by: Dufera, Amanuel Tafese
Published: (2025)
Lesion-Aware Visual-Language Fusion for Automated Image Captioning of Ulcerative Colitis Endoscopic Examinations
by: Escamilla, Alexis Ivan Lopez, et al.
Published: (2025)
by: Escamilla, Alexis Ivan Lopez, et al.
Published: (2025)
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
by: Truong, Khang, et al.
Published: (2025)
by: Truong, Khang, et al.
Published: (2025)
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge
by: Violos, John, et al.
Published: (2024)
by: Violos, John, et al.
Published: (2024)
Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers
by: Sun, Bohang, et al.
Published: (2025)
by: Sun, Bohang, et al.
Published: (2025)
Image Captioning in news report scenario
by: Liu, Tianrui, et al.
Published: (2024)
by: Liu, Tianrui, et al.
Published: (2024)
When CNNs Outperform Transformers and Mambas: Revisiting Deep Architectures for Dental Caries Segmentation
by: Ghimire, Aashish, et al.
Published: (2025)
by: Ghimire, Aashish, et al.
Published: (2025)
Image Embedding Sampling Method for Diverse Captioning
by: Waheed, Sania, et al.
Published: (2025)
by: Waheed, Sania, et al.
Published: (2025)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
by: Papadimitriou, Christos, et al.
Published: (2024)
by: Papadimitriou, Christos, et al.
Published: (2024)
Is Your Text-to-Image Model Robust to Caption Noise?
by: Yu, Weichen, et al.
Published: (2024)
by: Yu, Weichen, et al.
Published: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
by: Min, Kyungmin, et al.
Published: (2026)
by: Min, Kyungmin, et al.
Published: (2026)
An Ensemble Model with Attention Based Mechanism for Image Captioning
by: Badarneh, Israa Al, et al.
Published: (2025)
by: Badarneh, Israa Al, et al.
Published: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
by: Lee, Hankyeol, et al.
Published: (2025)
by: Lee, Hankyeol, et al.
Published: (2025)
Less is More: The Influence of Pruning on the Explainability of CNNs
by: Merkle, Florian, et al.
Published: (2023)
by: Merkle, Florian, et al.
Published: (2023)
TraNCE: Transformative Non-linear Concept Explainer for CNNs
by: Akpudo, Ugochukwu Ejike, et al.
Published: (2025)
by: Akpudo, Ugochukwu Ejike, et al.
Published: (2025)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Uterine Ultrasound Image Captioning Using Deep Learning Techniques
by: Boulesnane, Abdennour, et al.
Published: (2024)
by: Boulesnane, Abdennour, et al.
Published: (2024)
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
by: Jiang, Yanbei, et al.
Published: (2024)
by: Jiang, Yanbei, et al.
Published: (2024)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
by: Pham, Anh-Cuong, et al.
Published: (2024)
by: Pham, Anh-Cuong, et al.
Published: (2024)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
Investigating Calibration and Corruption Robustness of Post-hoc Pruned Perception CNNs: An Image Classification Benchmark Study
by: Mitra, Pallavi, et al.
Published: (2024)
by: Mitra, Pallavi, et al.
Published: (2024)
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
Enhancing Image Caption Generation Using Reinforcement Learning with Human Feedback
by: L, Adarsh N, et al.
Published: (2024)
by: L, Adarsh N, et al.
Published: (2024)
Reframing Image Difference Captioning with BLIP2IDC and Synthetic Augmentation
by: Evennou, Gautier, et al.
Published: (2024)
by: Evennou, Gautier, et al.
Published: (2024)
PixLore: A Dataset-driven Approach to Rich Image Captioning
by: Bonilla-Salvador, Diego, et al.
Published: (2023)
by: Bonilla-Salvador, Diego, et al.
Published: (2023)
CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
by: Tang, Zhijiang, et al.
Published: (2026)
by: Tang, Zhijiang, et al.
Published: (2026)
Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm
by: Li, Sixing, et al.
Published: (2026)
by: Li, Sixing, et al.
Published: (2026)
RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
by: Long, Xiaosheng, et al.
Published: (2025)
by: Long, Xiaosheng, et al.
Published: (2025)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
Integrative CAM: Adaptive Layer Fusion for Comprehensive Interpretation of CNNs
by: Singh, Aniket K., et al.
Published: (2024)
by: Singh, Aniket K., et al.
Published: (2024)
Enhancing CNNs robustness to occlusions with bioinspired filters for border completion
by: Coutinho, Catarina P., et al.
Published: (2025)
by: Coutinho, Catarina P., et al.
Published: (2025)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
by: Sharifzadeh, Sahand, et al.
Published: (2024)
by: Sharifzadeh, Sahand, et al.
Published: (2024)
Similar Items
-
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
by: Kolouju, Pranavi, et al.
Published: (2025) -
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026) -
The Effective Depth Paradox: Evaluating the Relationship between Architectural Topology and Trainability in Deep CNNs
by: Fischer, Manfred M., et al.
Published: (2026) -
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025) -
Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
by: Dufera, Amanuel Tafese
Published: (2025)