MsEdF: A Multi-stream Encoder-decoder Framework for Remote Sensing Image Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Das, Swadhin, Sharma, Raksha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2025)
por: Das, Swadhin, et al.
Publicado: (2025)
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2025)
por: Das, Swadhin, et al.
Publicado: (2025)
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2026)
por: Das, Swadhin, et al.
Publicado: (2026)
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
por: Si, Yonghao, et al.
Publicado: (2026)
por: Si, Yonghao, et al.
Publicado: (2026)
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
por: Sharma, Pavan Kumar, et al.
Publicado: (2023)
por: Sharma, Pavan Kumar, et al.
Publicado: (2023)
CHiQPM: Calibrated Hierarchical Interpretable Image Classification
por: Norrenbrock, Thomas, et al.
Publicado: (2025)
por: Norrenbrock, Thomas, et al.
Publicado: (2025)
Evaluating the Utility of Conformal Prediction Sets for AI-Advised Image Labeling
por: Zhang, Dongping, et al.
Publicado: (2024)
por: Zhang, Dongping, et al.
Publicado: (2024)
FreeDrag: Feature Dragging for Reliable Point-based Image Editing
por: Ling, Pengyang, et al.
Publicado: (2023)
por: Ling, Pengyang, et al.
Publicado: (2023)
Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness
por: Lee, Soohyun, et al.
Publicado: (2024)
por: Lee, Soohyun, et al.
Publicado: (2024)
Safeguarding Generative AI Applications in Preclinical Imaging through Hybrid Anomaly Detection
por: Binda, Jakub, et al.
Publicado: (2025)
por: Binda, Jakub, et al.
Publicado: (2025)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
por: Zimmermann, Robert, et al.
Publicado: (2026)
por: Zimmermann, Robert, et al.
Publicado: (2026)
MiMICRI: Towards Domain-centered Counterfactual Explanations of Cardiovascular Image Classification Models
por: Guo, Grace, et al.
Publicado: (2024)
por: Guo, Grace, et al.
Publicado: (2024)
Accessible, At-Home Detection of Parkinson's Disease via Multi-task Video Analysis
por: Islam, Md Saiful, et al.
Publicado: (2024)
por: Islam, Md Saiful, et al.
Publicado: (2024)
ArtCognition: A Multimodal AI Framework for Affective State Sensing from Visual and Kinematic Drawing Cues
por: Binaei-Haghighi, Behrad, et al.
Publicado: (2026)
por: Binaei-Haghighi, Behrad, et al.
Publicado: (2026)
A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2024)
por: Das, Swadhin, et al.
Publicado: (2024)
VRMN-bD: A Multi-modal Natural Behavior Dataset of Immersive Human Fear Responses in VR Stand-up Interactive Games
por: Zhang, He, et al.
Publicado: (2024)
por: Zhang, He, et al.
Publicado: (2024)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
por: Zhang, He, et al.
Publicado: (2025)
por: Zhang, He, et al.
Publicado: (2025)
Text-to-Image Generation for Vocabulary Learning Using the Keyword Method
por: Attygalle, Nuwan T., et al.
Publicado: (2025)
por: Attygalle, Nuwan T., et al.
Publicado: (2025)
Human-in-the-Loop Segmentation of Multi-species Coral Imagery
por: Raine, Scarlett, et al.
Publicado: (2024)
por: Raine, Scarlett, et al.
Publicado: (2024)
Real Time Captioning of Sign Language Gestures in Video Meetings
por: Mukherjee, Sharanya, et al.
Publicado: (2025)
por: Mukherjee, Sharanya, et al.
Publicado: (2025)
Enhanced Automated Quality Assessment Network for Interactive Building Segmentation in High-Resolution Remote Sensing Imagery
por: Zhang, Zhili, et al.
Publicado: (2024)
por: Zhang, Zhili, et al.
Publicado: (2024)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
por: Garg, Kapil, et al.
Publicado: (2025)
por: Garg, Kapil, et al.
Publicado: (2025)
Assessing Intersectional Bias in Representations of Pre-Trained Image Recognition Models
por: Krug, Valerie, et al.
Publicado: (2025)
por: Krug, Valerie, et al.
Publicado: (2025)
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
por: Tan, Taoliang, et al.
Publicado: (2025)
por: Tan, Taoliang, et al.
Publicado: (2025)
Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition
por: Zhou, Zishu, et al.
Publicado: (2026)
por: Zhou, Zishu, et al.
Publicado: (2026)
No Need to Sacrifice Data Quality for Quantity: Crowd-Informed Machine Annotation for Cost-Effective Understanding of Visual Data
por: Klugmann, Christopher, et al.
Publicado: (2024)
por: Klugmann, Christopher, et al.
Publicado: (2024)
Helios: An extremely low power event-based gesture recognition for always-on smart eyewear
por: Bhattacharyya, Prarthana, et al.
Publicado: (2024)
por: Bhattacharyya, Prarthana, et al.
Publicado: (2024)
SLIMBRAIN: Augmented Reality Real-Time Acquisition and Processing System For Hyperspectral Classification Mapping with Depth Information for In-Vivo Surgical Procedures
por: Sancho, Jaime, et al.
Publicado: (2024)
por: Sancho, Jaime, et al.
Publicado: (2024)
From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review
por: Sbeyti, Moussa Kassem, et al.
Publicado: (2026)
por: Sbeyti, Moussa Kassem, et al.
Publicado: (2026)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
por: Gutiérrez, Juan, et al.
Publicado: (2025)
por: Gutiérrez, Juan, et al.
Publicado: (2025)
VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection
por: Chavan, Kunal, et al.
Publicado: (2025)
por: Chavan, Kunal, et al.
Publicado: (2025)
Exploring Emotion Expression Recognition in Older Adults Interacting with a Virtual Coach
por: Palmero, Cristina, et al.
Publicado: (2023)
por: Palmero, Cristina, et al.
Publicado: (2023)
Improve accessibility for Low Vision and Blind people using Machine Learning and Computer Vision
por: Shukurov, Jasur
Publicado: (2024)
por: Shukurov, Jasur
Publicado: (2024)
ExeChecker: Where Did I Go Wrong?
por: Gu, Yiwen, et al.
Publicado: (2024)
por: Gu, Yiwen, et al.
Publicado: (2024)
A General Model for Detecting Learner Engagement: Implementation and Evaluation
por: Malekshahi, Somayeh, et al.
Publicado: (2024)
por: Malekshahi, Somayeh, et al.
Publicado: (2024)
Learning Confident Classifiers in the Presence of Label Noise
por: Hashmi, Asma Ahmed, et al.
Publicado: (2023)
por: Hashmi, Asma Ahmed, et al.
Publicado: (2023)
Using a CNN Model to Assess Paintings' Creativity
por: Zhang, Zhehan, et al.
Publicado: (2024)
por: Zhang, Zhehan, et al.
Publicado: (2024)
Methodology to Deploy CNN-Based Computer Vision Models on Immersive Wearable Devices
por: Malek, Kaveh, et al.
Publicado: (2024)
por: Malek, Kaveh, et al.
Publicado: (2024)
Evaluating the Evaluators: Towards Human-aligned Metrics for Missing Markers Reconstruction
por: Kucherenko, Taras, et al.
Publicado: (2024)
por: Kucherenko, Taras, et al.
Publicado: (2024)
emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation
por: Salter, Sasha, et al.
Publicado: (2024)
por: Salter, Sasha, et al.
Publicado: (2024)
Ejemplares similares
-
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2025) -
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2025) -
JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2026) -
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
por: Si, Yonghao, et al.
Publicado: (2026) -
A Review of Driver Gaze Estimation and Application in Gaze Behavior Understanding
por: Sharma, Pavan Kumar, et al.
Publicado: (2023)