MsEdF: A Multi-stream Encoder-decoder Framework for Remote Sensing Image Captioning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Das, Swadhin, Sharma, Raksha |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2025)
par: Das, Swadhin, et autres
Publié: (2025)
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2025)
par: Das, Swadhin, et autres
Publié: (2025)
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2026)
par: Das, Swadhin, et autres
Publié: (2026)
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
par: Si, Yonghao, et autres
Publié: (2026)
par: Si, Yonghao, et autres
Publié: (2026)
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
par: Sharma, Pavan Kumar, et autres
Publié: (2023)
par: Sharma, Pavan Kumar, et autres
Publié: (2023)
CHiQPM: Calibrated Hierarchical Interpretable Image Classification
par: Norrenbrock, Thomas, et autres
Publié: (2025)
par: Norrenbrock, Thomas, et autres
Publié: (2025)
Evaluating the Utility of Conformal Prediction Sets for AI-Advised Image Labeling
par: Zhang, Dongping, et autres
Publié: (2024)
par: Zhang, Dongping, et autres
Publié: (2024)
FreeDrag: Feature Dragging for Reliable Point-based Image Editing
par: Ling, Pengyang, et autres
Publié: (2023)
par: Ling, Pengyang, et autres
Publié: (2023)
Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness
par: Lee, Soohyun, et autres
Publié: (2024)
par: Lee, Soohyun, et autres
Publié: (2024)
Safeguarding Generative AI Applications in Preclinical Imaging through Hybrid Anomaly Detection
par: Binda, Jakub, et autres
Publié: (2025)
par: Binda, Jakub, et autres
Publié: (2025)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
par: Zimmermann, Robert, et autres
Publié: (2026)
par: Zimmermann, Robert, et autres
Publié: (2026)
MiMICRI: Towards Domain-centered Counterfactual Explanations of Cardiovascular Image Classification Models
par: Guo, Grace, et autres
Publié: (2024)
par: Guo, Grace, et autres
Publié: (2024)
Accessible, At-Home Detection of Parkinson's Disease via Multi-task Video Analysis
par: Islam, Md Saiful, et autres
Publié: (2024)
par: Islam, Md Saiful, et autres
Publié: (2024)
ArtCognition: A Multimodal AI Framework for Affective State Sensing from Visual and Kinematic Drawing Cues
par: Binaei-Haghighi, Behrad, et autres
Publié: (2026)
par: Binaei-Haghighi, Behrad, et autres
Publié: (2026)
A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2024)
par: Das, Swadhin, et autres
Publié: (2024)
VRMN-bD: A Multi-modal Natural Behavior Dataset of Immersive Human Fear Responses in VR Stand-up Interactive Games
par: Zhang, He, et autres
Publié: (2024)
par: Zhang, He, et autres
Publié: (2024)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
par: Zhang, He, et autres
Publié: (2025)
par: Zhang, He, et autres
Publié: (2025)
Text-to-Image Generation for Vocabulary Learning Using the Keyword Method
par: Attygalle, Nuwan T., et autres
Publié: (2025)
par: Attygalle, Nuwan T., et autres
Publié: (2025)
Human-in-the-Loop Segmentation of Multi-species Coral Imagery
par: Raine, Scarlett, et autres
Publié: (2024)
par: Raine, Scarlett, et autres
Publié: (2024)
Real Time Captioning of Sign Language Gestures in Video Meetings
par: Mukherjee, Sharanya, et autres
Publié: (2025)
par: Mukherjee, Sharanya, et autres
Publié: (2025)
Enhanced Automated Quality Assessment Network for Interactive Building Segmentation in High-Resolution Remote Sensing Imagery
par: Zhang, Zhili, et autres
Publié: (2024)
par: Zhang, Zhili, et autres
Publié: (2024)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
par: Garg, Kapil, et autres
Publié: (2025)
par: Garg, Kapil, et autres
Publié: (2025)
Assessing Intersectional Bias in Representations of Pre-Trained Image Recognition Models
par: Krug, Valerie, et autres
Publié: (2025)
par: Krug, Valerie, et autres
Publié: (2025)
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
par: Tan, Taoliang, et autres
Publié: (2025)
par: Tan, Taoliang, et autres
Publié: (2025)
Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition
par: Zhou, Zishu, et autres
Publié: (2026)
par: Zhou, Zishu, et autres
Publié: (2026)
No Need to Sacrifice Data Quality for Quantity: Crowd-Informed Machine Annotation for Cost-Effective Understanding of Visual Data
par: Klugmann, Christopher, et autres
Publié: (2024)
par: Klugmann, Christopher, et autres
Publié: (2024)
Helios: An extremely low power event-based gesture recognition for always-on smart eyewear
par: Bhattacharyya, Prarthana, et autres
Publié: (2024)
par: Bhattacharyya, Prarthana, et autres
Publié: (2024)
SLIMBRAIN: Augmented Reality Real-Time Acquisition and Processing System For Hyperspectral Classification Mapping with Depth Information for In-Vivo Surgical Procedures
par: Sancho, Jaime, et autres
Publié: (2024)
par: Sancho, Jaime, et autres
Publié: (2024)
From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review
par: Sbeyti, Moussa Kassem, et autres
Publié: (2026)
par: Sbeyti, Moussa Kassem, et autres
Publié: (2026)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
par: Gutiérrez, Juan, et autres
Publié: (2025)
par: Gutiérrez, Juan, et autres
Publié: (2025)
VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection
par: Chavan, Kunal, et autres
Publié: (2025)
par: Chavan, Kunal, et autres
Publié: (2025)
Exploring Emotion Expression Recognition in Older Adults Interacting with a Virtual Coach
par: Palmero, Cristina, et autres
Publié: (2023)
par: Palmero, Cristina, et autres
Publié: (2023)
Improve accessibility for Low Vision and Blind people using Machine Learning and Computer Vision
par: Shukurov, Jasur
Publié: (2024)
par: Shukurov, Jasur
Publié: (2024)
ExeChecker: Where Did I Go Wrong?
par: Gu, Yiwen, et autres
Publié: (2024)
par: Gu, Yiwen, et autres
Publié: (2024)
A General Model for Detecting Learner Engagement: Implementation and Evaluation
par: Malekshahi, Somayeh, et autres
Publié: (2024)
par: Malekshahi, Somayeh, et autres
Publié: (2024)
Learning Confident Classifiers in the Presence of Label Noise
par: Hashmi, Asma Ahmed, et autres
Publié: (2023)
par: Hashmi, Asma Ahmed, et autres
Publié: (2023)
Using a CNN Model to Assess Paintings' Creativity
par: Zhang, Zhehan, et autres
Publié: (2024)
par: Zhang, Zhehan, et autres
Publié: (2024)
Methodology to Deploy CNN-Based Computer Vision Models on Immersive Wearable Devices
par: Malek, Kaveh, et autres
Publié: (2024)
par: Malek, Kaveh, et autres
Publié: (2024)
Evaluating the Evaluators: Towards Human-aligned Metrics for Missing Markers Reconstruction
par: Kucherenko, Taras, et autres
Publié: (2024)
par: Kucherenko, Taras, et autres
Publié: (2024)
emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation
par: Salter, Sasha, et autres
Publié: (2024)
par: Salter, Sasha, et autres
Publié: (2024)
Documents similaires
-
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2025) -
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2025) -
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
par: Das, Swadhin, et autres
Publié: (2026) -
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
par: Si, Yonghao, et autres
Publié: (2026) -
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
par: Sharma, Pavan Kumar, et autres
Publié: (2023)