MsEdF: A Multi-stream Encoder-decoder Framework for Remote Sensing Image Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Das, Swadhin, Sharma, Raksha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025)
di: Das, Swadhin, et al.
Pubblicazione: (2025)
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025)
di: Das, Swadhin, et al.
Pubblicazione: (2025)
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2026)
di: Das, Swadhin, et al.
Pubblicazione: (2026)
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
di: Si, Yonghao, et al.
Pubblicazione: (2026)
di: Si, Yonghao, et al.
Pubblicazione: (2026)
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
di: Sharma, Pavan Kumar, et al.
Pubblicazione: (2023)
di: Sharma, Pavan Kumar, et al.
Pubblicazione: (2023)
CHiQPM: Calibrated Hierarchical Interpretable Image Classification
di: Norrenbrock, Thomas, et al.
Pubblicazione: (2025)
di: Norrenbrock, Thomas, et al.
Pubblicazione: (2025)
Evaluating the Utility of Conformal Prediction Sets for AI-Advised Image Labeling
di: Zhang, Dongping, et al.
Pubblicazione: (2024)
di: Zhang, Dongping, et al.
Pubblicazione: (2024)
FreeDrag: Feature Dragging for Reliable Point-based Image Editing
di: Ling, Pengyang, et al.
Pubblicazione: (2023)
di: Ling, Pengyang, et al.
Pubblicazione: (2023)
Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness
di: Lee, Soohyun, et al.
Pubblicazione: (2024)
di: Lee, Soohyun, et al.
Pubblicazione: (2024)
Safeguarding Generative AI Applications in Preclinical Imaging through Hybrid Anomaly Detection
di: Binda, Jakub, et al.
Pubblicazione: (2025)
di: Binda, Jakub, et al.
Pubblicazione: (2025)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
di: Zimmermann, Robert, et al.
Pubblicazione: (2026)
di: Zimmermann, Robert, et al.
Pubblicazione: (2026)
MiMICRI: Towards Domain-centered Counterfactual Explanations of Cardiovascular Image Classification Models
di: Guo, Grace, et al.
Pubblicazione: (2024)
di: Guo, Grace, et al.
Pubblicazione: (2024)
Accessible, At-Home Detection of Parkinson's Disease via Multi-task Video Analysis
di: Islam, Md Saiful, et al.
Pubblicazione: (2024)
di: Islam, Md Saiful, et al.
Pubblicazione: (2024)
ArtCognition: A Multimodal AI Framework for Affective State Sensing from Visual and Kinematic Drawing Cues
di: Binaei-Haghighi, Behrad, et al.
Pubblicazione: (2026)
di: Binaei-Haghighi, Behrad, et al.
Pubblicazione: (2026)
A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2024)
di: Das, Swadhin, et al.
Pubblicazione: (2024)
VRMN-bD: A Multi-modal Natural Behavior Dataset of Immersive Human Fear Responses in VR Stand-up Interactive Games
di: Zhang, He, et al.
Pubblicazione: (2024)
di: Zhang, He, et al.
Pubblicazione: (2024)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
di: Zhang, He, et al.
Pubblicazione: (2025)
di: Zhang, He, et al.
Pubblicazione: (2025)
Text-to-Image Generation for Vocabulary Learning Using the Keyword Method
di: Attygalle, Nuwan T., et al.
Pubblicazione: (2025)
di: Attygalle, Nuwan T., et al.
Pubblicazione: (2025)
Human-in-the-Loop Segmentation of Multi-species Coral Imagery
di: Raine, Scarlett, et al.
Pubblicazione: (2024)
di: Raine, Scarlett, et al.
Pubblicazione: (2024)
Real Time Captioning of Sign Language Gestures in Video Meetings
di: Mukherjee, Sharanya, et al.
Pubblicazione: (2025)
di: Mukherjee, Sharanya, et al.
Pubblicazione: (2025)
Enhanced Automated Quality Assessment Network for Interactive Building Segmentation in High-Resolution Remote Sensing Imagery
di: Zhang, Zhili, et al.
Pubblicazione: (2024)
di: Zhang, Zhili, et al.
Pubblicazione: (2024)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
di: Garg, Kapil, et al.
Pubblicazione: (2025)
di: Garg, Kapil, et al.
Pubblicazione: (2025)
Assessing Intersectional Bias in Representations of Pre-Trained Image Recognition Models
di: Krug, Valerie, et al.
Pubblicazione: (2025)
di: Krug, Valerie, et al.
Pubblicazione: (2025)
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
di: Tan, Taoliang, et al.
Pubblicazione: (2025)
di: Tan, Taoliang, et al.
Pubblicazione: (2025)
Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition
di: Zhou, Zishu, et al.
Pubblicazione: (2026)
di: Zhou, Zishu, et al.
Pubblicazione: (2026)
No Need to Sacrifice Data Quality for Quantity: Crowd-Informed Machine Annotation for Cost-Effective Understanding of Visual Data
di: Klugmann, Christopher, et al.
Pubblicazione: (2024)
di: Klugmann, Christopher, et al.
Pubblicazione: (2024)
Helios: An extremely low power event-based gesture recognition for always-on smart eyewear
di: Bhattacharyya, Prarthana, et al.
Pubblicazione: (2024)
di: Bhattacharyya, Prarthana, et al.
Pubblicazione: (2024)
SLIMBRAIN: Augmented Reality Real-Time Acquisition and Processing System For Hyperspectral Classification Mapping with Depth Information for In-Vivo Surgical Procedures
di: Sancho, Jaime, et al.
Pubblicazione: (2024)
di: Sancho, Jaime, et al.
Pubblicazione: (2024)
From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review
di: Sbeyti, Moussa Kassem, et al.
Pubblicazione: (2026)
di: Sbeyti, Moussa Kassem, et al.
Pubblicazione: (2026)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
di: Gutiérrez, Juan, et al.
Pubblicazione: (2025)
di: Gutiérrez, Juan, et al.
Pubblicazione: (2025)
VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection
di: Chavan, Kunal, et al.
Pubblicazione: (2025)
di: Chavan, Kunal, et al.
Pubblicazione: (2025)
Exploring Emotion Expression Recognition in Older Adults Interacting with a Virtual Coach
di: Palmero, Cristina, et al.
Pubblicazione: (2023)
di: Palmero, Cristina, et al.
Pubblicazione: (2023)
Improve accessibility for Low Vision and Blind people using Machine Learning and Computer Vision
di: Shukurov, Jasur
Pubblicazione: (2024)
di: Shukurov, Jasur
Pubblicazione: (2024)
ExeChecker: Where Did I Go Wrong?
di: Gu, Yiwen, et al.
Pubblicazione: (2024)
di: Gu, Yiwen, et al.
Pubblicazione: (2024)
A General Model for Detecting Learner Engagement: Implementation and Evaluation
di: Malekshahi, Somayeh, et al.
Pubblicazione: (2024)
di: Malekshahi, Somayeh, et al.
Pubblicazione: (2024)
Learning Confident Classifiers in the Presence of Label Noise
di: Hashmi, Asma Ahmed, et al.
Pubblicazione: (2023)
di: Hashmi, Asma Ahmed, et al.
Pubblicazione: (2023)
Using a CNN Model to Assess Paintings' Creativity
di: Zhang, Zhehan, et al.
Pubblicazione: (2024)
di: Zhang, Zhehan, et al.
Pubblicazione: (2024)
Methodology to Deploy CNN-Based Computer Vision Models on Immersive Wearable Devices
di: Malek, Kaveh, et al.
Pubblicazione: (2024)
di: Malek, Kaveh, et al.
Pubblicazione: (2024)
Evaluating the Evaluators: Towards Human-aligned Metrics for Missing Markers Reconstruction
di: Kucherenko, Taras, et al.
Pubblicazione: (2024)
di: Kucherenko, Taras, et al.
Pubblicazione: (2024)
emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation
di: Salter, Sasha, et al.
Pubblicazione: (2024)
di: Salter, Sasha, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025) -
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025) -
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2026) -
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
di: Si, Yonghao, et al.
Pubblicazione: (2026) -
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
di: Sharma, Pavan Kumar, et al.
Pubblicazione: (2023)