Research on Image Recognition Technology Based on Multimodal Deep Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jinyin, Li, Xingchen, Jin, Yixuan, Zhong, Yihao, Zhang, Keke, Zhou, Chang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
Research on Driver Facial Fatigue Detection Based on Yolov8 Model
von: Zhou, Chang, et al.
Veröffentlicht: (2024)
von: Zhou, Chang, et al.
Veröffentlicht: (2024)
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
von: Jing, Long, et al.
Veröffentlicht: (2026)
von: Jing, Long, et al.
Veröffentlicht: (2026)
MSDS: Deep Structural Similarity with Multiscale Representation
von: Kang, Danling, et al.
Veröffentlicht: (2026)
von: Kang, Danling, et al.
Veröffentlicht: (2026)
Hyperspectral Image Analysis in Single-Modal and Multimodal setting using Deep Learning Techniques
von: Pande, Shivam
Veröffentlicht: (2024)
von: Pande, Shivam
Veröffentlicht: (2024)
Deep Learning-based Text-in-Image Watermarking
von: Karki, Bishwa, et al.
Veröffentlicht: (2024)
von: Karki, Bishwa, et al.
Veröffentlicht: (2024)
Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects
von: Warner, Elisa, et al.
Veröffentlicht: (2023)
von: Warner, Elisa, et al.
Veröffentlicht: (2023)
Multimodal Deep Learning for Diabetic Foot Ulcer Staging Using Integrated RGB and Thermal Imaging
von: Mermer, Gulengul, et al.
Veröffentlicht: (2026)
von: Mermer, Gulengul, et al.
Veröffentlicht: (2026)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
von: Cheng, Jintao, et al.
Veröffentlicht: (2025)
MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Calibrating Multimodal Consensus for Emotion Recognition
von: Zhong, Guowei, et al.
Veröffentlicht: (2025)
von: Zhong, Guowei, et al.
Veröffentlicht: (2025)
CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive Learning
von: Wang, Yiping, et al.
Veröffentlicht: (2024)
von: Wang, Yiping, et al.
Veröffentlicht: (2024)
MIP: CLIP-based Image Reconstruction from PEFT Gradients
von: Zhou, Peiheng, et al.
Veröffentlicht: (2024)
von: Zhou, Peiheng, et al.
Veröffentlicht: (2024)
MetaWild: A Multimodal Dataset for Animal Re-Identification with Environmental Metadata
von: Li, Yuzhuo, et al.
Veröffentlicht: (2025)
von: Li, Yuzhuo, et al.
Veröffentlicht: (2025)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
von: Wang, Zeheng, et al.
Veröffentlicht: (2026)
von: Wang, Zeheng, et al.
Veröffentlicht: (2026)
STR-Cert: Robustness Certification for Deep Text Recognition on Deep Learning Pipelines and Vision Transformers
von: Shao, Daqian, et al.
Veröffentlicht: (2023)
von: Shao, Daqian, et al.
Veröffentlicht: (2023)
ECG-Image-Kit: A Synthetic Image Generation Toolbox to Facilitate Deep Learning-Based Electrocardiogram Digitization
von: Shivashankara, Kshama Kodthalu, et al.
Veröffentlicht: (2023)
von: Shivashankara, Kshama Kodthalu, et al.
Veröffentlicht: (2023)
Learning Multimodal Latent Generative Models with Energy-Based Prior
von: Yuan, Shiyu, et al.
Veröffentlicht: (2024)
von: Yuan, Shiyu, et al.
Veröffentlicht: (2024)
BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
ThermoPore: Predicting Part Porosity Based on Thermal Images Using Deep Learning
von: Pak, Peter Myung-Won, et al.
Veröffentlicht: (2024)
von: Pak, Peter Myung-Won, et al.
Veröffentlicht: (2024)
Deep Spatially-Regularized and Superpixel-Based Diffusion Learning for Unsupervised Hyperspectral Image Clustering
von: Buranasiri, Vutichart, et al.
Veröffentlicht: (2026)
von: Buranasiri, Vutichart, et al.
Veröffentlicht: (2026)
Continual Deep Active Learning for Medical Imaging: Replay-Base Architecture for Context Adaptation
von: Daniel, Rui, et al.
Veröffentlicht: (2025)
von: Daniel, Rui, et al.
Veröffentlicht: (2025)
Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
von: Xie, Baao, et al.
Veröffentlicht: (2024)
von: Xie, Baao, et al.
Veröffentlicht: (2024)
Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
von: Zhong, Siru, et al.
Veröffentlicht: (2025)
von: Zhong, Siru, et al.
Veröffentlicht: (2025)
Deep Learning for BioImaging: What Are We Learning?
von: Svatko, Ivan, et al.
Veröffentlicht: (2026)
von: Svatko, Ivan, et al.
Veröffentlicht: (2026)
Exploring Transferability of Self-Supervised Learning by Task Conflict Calibration
von: Guo, Huijie, et al.
Veröffentlicht: (2025)
von: Guo, Huijie, et al.
Veröffentlicht: (2025)
Image Restoration Using Deep Regulated Convolutional Networks
von: Liu, Peng, et al.
Veröffentlicht: (2019)
von: Liu, Peng, et al.
Veröffentlicht: (2019)
Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping
von: Amer, Mohammed, et al.
Veröffentlicht: (2025)
von: Amer, Mohammed, et al.
Veröffentlicht: (2025)
Geospatial Representation Learning: A Survey from Deep Learning to The LLM Era
von: Hao, Xixuan, et al.
Veröffentlicht: (2025)
von: Hao, Xixuan, et al.
Veröffentlicht: (2025)
Learning Topological Representations for Deep Image Understanding
von: Hu, Xiaoling
Veröffentlicht: (2024)
von: Hu, Xiaoling
Veröffentlicht: (2024)
Deep Learning in Image Classification: Evaluating VGG19's Performance on Complex Visual Data
von: He, Weijie, et al.
Veröffentlicht: (2024)
von: He, Weijie, et al.
Veröffentlicht: (2024)
CDFL: Efficient Federated Human Activity Recognition using Contrastive Learning and Deep Clustering
von: Khazaei, Ensieh, et al.
Veröffentlicht: (2024)
von: Khazaei, Ensieh, et al.
Veröffentlicht: (2024)
TDEC: Deep Embedded Image Clustering with Transformer and Distribution Information
von: Zhang, Ruilin, et al.
Veröffentlicht: (2026)
von: Zhang, Ruilin, et al.
Veröffentlicht: (2026)
Classes Are Not Equal: An Empirical Study on Image Recognition Fairness
von: Cui, Jiequan, et al.
Veröffentlicht: (2024)
von: Cui, Jiequan, et al.
Veröffentlicht: (2024)
Predicting Lung Disease Severity via Image-Based AQI Analysis using Deep Learning Techniques
von: Mahajan, Anvita, et al.
Veröffentlicht: (2024)
von: Mahajan, Anvita, et al.
Veröffentlicht: (2024)
Performance Evaluation of Image Enhancement Techniques on Transfer Learning for Touchless Fingerprint Recognition
von: Sreehari, S, et al.
Veröffentlicht: (2025)
von: Sreehari, S, et al.
Veröffentlicht: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Breast Cancer Image Classification Method Based on Deep Transfer Learning
von: Wang, Weimin, et al.
Veröffentlicht: (2024)
von: Wang, Weimin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
von: Wang, Jinyin, et al.
Veröffentlicht: (2024) -
Research on Driver Facial Fatigue Detection Based on Yolov8 Model
von: Zhou, Chang, et al.
Veröffentlicht: (2024) -
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
von: Li, Wenkai, et al.
Veröffentlicht: (2025) -
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
von: Jing, Long, et al.
Veröffentlicht: (2026) -
MSDS: Deep Structural Similarity with Multiscale Representation
von: Kang, Danling, et al.
Veröffentlicht: (2026)