Compressed Image Captioning using CNN-based Encoder-Decoder Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Ridoy, Md Alif Rahman, Hasan, M Mahmud, Bhowmick, Shovon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
di: Rahman, Md Mostafijur, et al.
Pubblicazione: (2025)
di: Rahman, Md Mostafijur, et al.
Pubblicazione: (2025)
Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)
di: Parajuli, Kabita, et al.
Pubblicazione: (2023)
di: Parajuli, Kabita, et al.
Pubblicazione: (2023)
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning
di: Adewale, Sikiru, et al.
Pubblicazione: (2023)
di: Adewale, Sikiru, et al.
Pubblicazione: (2023)
A Deep Learning-based Multimodal Depth-Aware Dynamic Hand Gesture Recognition System
di: Mahmud, Hasan, et al.
Pubblicazione: (2021)
di: Mahmud, Hasan, et al.
Pubblicazione: (2021)
Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)
Enhancing Diabetic Retinopathy Diagnosis: A Lightweight CNN Architecture for Efficient Exudate Detection in Retinal Fundus Images
di: Alif, Mujadded Al Rabbani
Pubblicazione: (2024)
di: Alif, Mujadded Al Rabbani
Pubblicazione: (2024)
DE-KAN: A Kolmogorov Arnold Network with Dual Encoder for accurate 2D Teeth Segmentation
di: Mustakim, Md Mizanur Rahman, et al.
Pubblicazione: (2025)
di: Mustakim, Md Mizanur Rahman, et al.
Pubblicazione: (2025)
Blind Image Deblurring with FFT-ReLU Sparsity Prior
di: Radi, Abdul Mohaimen Al, et al.
Pubblicazione: (2024)
di: Radi, Abdul Mohaimen Al, et al.
Pubblicazione: (2024)
BdSLW60: A Word-Level Bangla Sign Language Dataset
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2024)
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2024)
A Comparison Study of Deep CNN Architecture in Detecting of Pneumonia
di: Porag, Al Mohidur Rahman, et al.
Pubblicazione: (2022)
di: Porag, Al Mohidur Rahman, et al.
Pubblicazione: (2022)
GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer
di: Hasan, Md. Mahmudul, et al.
Pubblicazione: (2025)
di: Hasan, Md. Mahmudul, et al.
Pubblicazione: (2025)
Explainable Image Captioning using CNN- CNN architecture and Hierarchical Attention
di: Mohan, Rishi Kesav, et al.
Pubblicazione: (2024)
di: Mohan, Rishi Kesav, et al.
Pubblicazione: (2024)
New Encoder Learning for Captioning Heavy Rain Images via Semantic Visual Feature Matching
di: Son, Chang-Hwan, et al.
Pubblicazione: (2021)
di: Son, Chang-Hwan, et al.
Pubblicazione: (2021)
Fine-Tuned CNN-Based Approach for Multi-Class Mango Leaf Disease Detection
di: Ahmmed, Jalal, et al.
Pubblicazione: (2025)
di: Ahmmed, Jalal, et al.
Pubblicazione: (2025)
ParaTransCNN: Parallelized TransCNN Encoder for Medical Image Segmentation
di: Sun, Hongkun, et al.
Pubblicazione: (2024)
di: Sun, Hongkun, et al.
Pubblicazione: (2024)
An empirical study for the early detection of Mpox from skin lesion images using pretrained CNN models leveraging XAI technique
di: Rahim, Mohammad Asifur, et al.
Pubblicazione: (2025)
di: Rahim, Mohammad Asifur, et al.
Pubblicazione: (2025)
Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
di: Zhang, Tianyu, et al.
Pubblicazione: (2025)
di: Zhang, Tianyu, et al.
Pubblicazione: (2025)
Deblurring in the Wild: A Real-World Image Deblurring Dataset from Smartphone High-Speed Videos
di: Mahmud, Syed Mumtahin, et al.
Pubblicazione: (2025)
di: Mahmud, Syed Mumtahin, et al.
Pubblicazione: (2025)
From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
di: Ishmam, Md Farhan, et al.
Pubblicazione: (2023)
di: Ishmam, Md Farhan, et al.
Pubblicazione: (2023)
Comparative Analysis of Custom CNN Architectures versus Pre-trained Models and Transfer Learning: A Study on Five Bangladesh Datasets
di: Tanvir, Ibrahim, et al.
Pubblicazione: (2026)
di: Tanvir, Ibrahim, et al.
Pubblicazione: (2026)
Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation
di: Tariquzzaman, Md, et al.
Pubblicazione: (2025)
di: Tariquzzaman, Md, et al.
Pubblicazione: (2025)
KDC-Diff: A Latent-Aware Diffusion Model with Knowledge Retention for Memory-Efficient Image Generation
di: Borno, Md. Naimur Asif, et al.
Pubblicazione: (2025)
di: Borno, Md. Naimur Asif, et al.
Pubblicazione: (2025)
MedPrompt: LLM-CNN Fusion with Weight Routing for Medical Image Segmentation and Classification
di: Sobhan, Shadman, et al.
Pubblicazione: (2025)
di: Sobhan, Shadman, et al.
Pubblicazione: (2025)
Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
di: Dufera, Amanuel Tafese
Pubblicazione: (2025)
di: Dufera, Amanuel Tafese
Pubblicazione: (2025)
From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring
di: Mahmud, Syed Mumtahin, et al.
Pubblicazione: (2025)
di: Mahmud, Syed Mumtahin, et al.
Pubblicazione: (2025)
WiFi based Human Fall and Activity Recognition using Transformer based Encoder Decoder and Graph Neural Networks
di: Cho, Younggeol, et al.
Pubblicazione: (2025)
di: Cho, Younggeol, et al.
Pubblicazione: (2025)
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
di: Xu, Zitong, et al.
Pubblicazione: (2026)
di: Xu, Zitong, et al.
Pubblicazione: (2026)
DeCoDrift: Stabilizing Decoder Coupling in Closed-Loop Foundation Segmentation
di: Tabib, H. M. Shadman, et al.
Pubblicazione: (2026)
di: Tabib, H. M. Shadman, et al.
Pubblicazione: (2026)
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2026)
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2026)
MsEdF: A Multi-stream Encoder-decoder Framework for Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025)
di: Das, Swadhin, et al.
Pubblicazione: (2025)
Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks
di: Shawon, Jubayer Ahmed Bhuiyan, et al.
Pubblicazione: (2025)
di: Shawon, Jubayer Ahmed Bhuiyan, et al.
Pubblicazione: (2025)
Road Traffic Sign Recognition method using Siamese network Combining Efficient-CNN based Encoder
di: Xi, Zhenghao, et al.
Pubblicazione: (2025)
di: Xi, Zhenghao, et al.
Pubblicazione: (2025)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
di: Cai, Chen, et al.
Pubblicazione: (2023)
di: Cai, Chen, et al.
Pubblicazione: (2023)
A Computer Vision Based Approach for Stalking Detection Using a CNN-LSTM-MLP Hybrid Fusion Model
di: Hasan, Murad, et al.
Pubblicazione: (2024)
di: Hasan, Murad, et al.
Pubblicazione: (2024)
HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
di: Rahman, Mohammad Amanour
Pubblicazione: (2025)
di: Rahman, Mohammad Amanour
Pubblicazione: (2025)
Decodable and Sample Invariant Continuous Object Encoder
di: Yuan, Dehao, et al.
Pubblicazione: (2023)
di: Yuan, Dehao, et al.
Pubblicazione: (2023)
Diabetic Retinopathy Classification from Retinal Images using Machine Learning Approaches
di: Bhattacharjee, Indronil, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Indronil, et al.
Pubblicazione: (2024)
From Explanations to Architecture: Explainability-Driven CNN Refinement for Brain Tumor Classification in MRI
di: Gupta, Rajan Das, et al.
Pubblicazione: (2025)
di: Gupta, Rajan Das, et al.
Pubblicazione: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
di: Yang, Shijia, et al.
Pubblicazione: (2025)
di: Yang, Shijia, et al.
Pubblicazione: (2025)
EMCAD: Efficient Multi-scale Convolutional Attention Decoding for Medical Image Segmentation
di: Rahman, Md Mostafijur, et al.
Pubblicazione: (2024)
di: Rahman, Md Mostafijur, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
di: Rahman, Md Mostafijur, et al.
Pubblicazione: (2025) -
Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)
di: Parajuli, Kabita, et al.
Pubblicazione: (2023) -
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning
di: Adewale, Sikiru, et al.
Pubblicazione: (2023) -
A Deep Learning-based Multimodal Depth-Aware Dynamic Hand Gesture Recognition System
di: Mahmud, Hasan, et al.
Pubblicazione: (2021) -
Bangla Sign Language Translation: Dataset Creation Challenges, Benchmarking and Prospects
di: Rubaiyeat, Husne Ara, et al.
Pubblicazione: (2025)