Saved in:
| Main Authors: | Mousse, Mikael A., Atohoun, Bethel C. A. R. K., Motamed, Cina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2401.15055 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
by: Motamed, Saman, et al.
Published: (2023)
by: Motamed, Saman, et al.
Published: (2023)
VWise: A novel benchmark for evaluating scene classification for vehicular applications
by: Azevedo, Pedro, et al.
Published: (2024)
by: Azevedo, Pedro, et al.
Published: (2024)
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
by: Sanchez, Cristhian, et al.
Published: (2024)
by: Sanchez, Cristhian, et al.
Published: (2024)
Deep semi-supervised approach based on consistency regularization and similarity learning for weeds classification
by: Benchallal, Farouq, et al.
Published: (2025)
by: Benchallal, Farouq, et al.
Published: (2025)
Motion-guided small MAV detection in complex and non-planar scenes
by: Guo, Hanqing, et al.
Published: (2024)
by: Guo, Hanqing, et al.
Published: (2024)
Dynamic loss balancing and sequential enhancement for road-safety assessment and traffic scene classification
by: Kačan, Marin, et al.
Published: (2022)
by: Kačan, Marin, et al.
Published: (2022)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
A Systematic Review of Deep Learning-based Research on Radiology Report Generation
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
Enhancing the quality of gauge images captured in smoke and haze scenes through deep learning
by: Ramírez-Agudelo, Oscar H., et al.
Published: (2026)
by: Ramírez-Agudelo, Oscar H., et al.
Published: (2026)
3D scene generation from scene graphs and self-attention
by: Bonazzi, Pietro, et al.
Published: (2024)
by: Bonazzi, Pietro, et al.
Published: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
by: Basak, Debolena, et al.
Published: (2024)
by: Basak, Debolena, et al.
Published: (2024)
Developing an efficient corpus using Ensemble Data cleaning approach
by: Ahad, Md Taimur
Published: (2024)
by: Ahad, Md Taimur
Published: (2024)
Scaling medical imaging report generation with multimodal reinforcement learning
by: Liu, Qianchu, et al.
Published: (2026)
by: Liu, Qianchu, et al.
Published: (2026)
SignBart -- New approach with the skeleton sequence for Isolated Sign language Recognition
by: Nguyen, Tinh, et al.
Published: (2025)
by: Nguyen, Tinh, et al.
Published: (2025)
3D objects and scenes classification, recognition, segmentation, and reconstruction using 3D point cloud data: A review
by: Elharrouss, Omar, et al.
Published: (2023)
by: Elharrouss, Omar, et al.
Published: (2023)
Improving Applicability of Deep Learning based Token Classification models during Training
by: Mehra, Anket, et al.
Published: (2025)
by: Mehra, Anket, et al.
Published: (2025)
Using Multimodal Deep Neural Networks to Disentangle Language from Visual Aesthetics
by: Conwell, Colin, et al.
Published: (2024)
by: Conwell, Colin, et al.
Published: (2024)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
by: Liang, Yunhao, et al.
Published: (2026)
by: Liang, Yunhao, et al.
Published: (2026)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
by: Zhu, Wenxin, et al.
Published: (2025)
by: Zhu, Wenxin, et al.
Published: (2025)
Breaking through the learning plateaus of in-context learning in Transformer
by: Fu, Jingwen, et al.
Published: (2023)
by: Fu, Jingwen, et al.
Published: (2023)
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
by: Zhu, Kejian, et al.
Published: (2025)
by: Zhu, Kejian, et al.
Published: (2025)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Human-annotated label noise and their impact on ConvNets for remote sensing image scene classification
by: Peng, Longkang, et al.
Published: (2023)
by: Peng, Longkang, et al.
Published: (2023)
Robust image classification with multi-modal large language models
by: Villani, Francesco, et al.
Published: (2024)
by: Villani, Francesco, et al.
Published: (2024)
Machine learning approach to brain tumor detection and classification
by: Oh, Alice, et al.
Published: (2024)
by: Oh, Alice, et al.
Published: (2024)
Malayalam Sign Language Identification using Finetuned YOLOv8 and Computer Vision Techniques
by: K., Abhinand, et al.
Published: (2024)
by: K., Abhinand, et al.
Published: (2024)
Nucleus subtype classification using inter-modality learning
by: Remedios, Lucas W., et al.
Published: (2024)
by: Remedios, Lucas W., et al.
Published: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
by: Miao, Yongzhu, et al.
Published: (2023)
by: Miao, Yongzhu, et al.
Published: (2023)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
by: Lai, Songning, et al.
Published: (2023)
by: Lai, Songning, et al.
Published: (2023)
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
by: Loginova, Olga, et al.
Published: (2025)
by: Loginova, Olga, et al.
Published: (2025)
HyperWalker: Dynamic Hypergraph-Based Deep Diagnosis for Multi-Hop Clinical Modeling across EHR and X-Ray in Medical VLMs
by: Yang, Yuezhe, et al.
Published: (2026)
by: Yang, Yuezhe, et al.
Published: (2026)
Feature boosting with efficient attention for scene parsing
by: Singh, Vivek, et al.
Published: (2024)
by: Singh, Vivek, et al.
Published: (2024)
Functionality understanding and segmentation in 3D scenes
by: Corsetti, Jaime, et al.
Published: (2024)
by: Corsetti, Jaime, et al.
Published: (2024)
Cross-attention for State-based model RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
A Tale of Two Languages: Large-Vocabulary Continuous Sign Language Recognition from Spoken Language Supervision
by: Raude, Charles, et al.
Published: (2024)
by: Raude, Charles, et al.
Published: (2024)
Deep transfer learning for image classification: a survey
by: Plested, Jo, et al.
Published: (2022)
by: Plested, Jo, et al.
Published: (2022)
Similar Items
-
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
by: Motamed, Saman, et al.
Published: (2023) -
VWise: A novel benchmark for evaluating scene classification for vehicular applications
by: Azevedo, Pedro, et al.
Published: (2024) -
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
by: Sanchez, Cristhian, et al.
Published: (2024) -
Deep semi-supervised approach based on consistency regularization and similarity learning for weeds classification
by: Benchallal, Farouq, et al.
Published: (2025) -
Motion-guided small MAV detection in complex and non-planar scenes
by: Guo, Hanqing, et al.
Published: (2024)