Gespeichert in:
| Hauptverfasser: | Ahmed, Raihan, Omi, Shahed Chowdhury, Rahman, Md. Sadman, Bhuiyan, Niaz Rahman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.01728 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FUSED-Net: Detecting Traffic Signs with Limited Data
von: Rahman, Md. Atiqur, et al.
Veröffentlicht: (2024)
von: Rahman, Md. Atiqur, et al.
Veröffentlicht: (2024)
BeHGAN: Bengali Handwritten Word Generation from Plain Text Using Generative Adversarial Networks
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
Intelligent Systems in Neuroimaging: Pioneering AI Techniques for Brain Tumor Detection
von: Islam, Md. Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md. Mohaiminul, et al.
Veröffentlicht: (2025)
STAR: A Foundation Model-driven Framework for Robust Task Planning and Failure Recovery in Robotic Systems
von: Sakib, Md Sadman, et al.
Veröffentlicht: (2025)
von: Sakib, Md Sadman, et al.
Veröffentlicht: (2025)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
von: Chowdhury, Md Abtahi Majeed, et al.
Veröffentlicht: (2025)
von: Chowdhury, Md Abtahi Majeed, et al.
Veröffentlicht: (2025)
Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models
von: Samin, Niamul Hassan, et al.
Veröffentlicht: (2026)
von: Samin, Niamul Hassan, et al.
Veröffentlicht: (2026)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
Surgeons Are Indian Males and Speech Therapists Are White Females: Auditing Biases in Vision-Language Models for Healthcare Professionals
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
von: Gupta, Kishor Datta, et al.
Veröffentlicht: (2025)
von: Gupta, Kishor Datta, et al.
Veröffentlicht: (2025)
An Efficient Deep Learning Framework for Brain Stroke Diagnosis Using Computed Tomography Images
von: Hossen, Md. Sabbir, et al.
Veröffentlicht: (2025)
von: Hossen, Md. Sabbir, et al.
Veröffentlicht: (2025)
MF-GCN: A Multi-Frequency Graph Convolutional Network for Tri-Modal Depression Detection Using Eye-Tracking, Facial, and Acoustic Features
von: Rahman, Sejuti, et al.
Veröffentlicht: (2025)
von: Rahman, Sejuti, et al.
Veröffentlicht: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
von: Rahman, Kazi Mahathir, et al.
Veröffentlicht: (2025)
von: Rahman, Kazi Mahathir, et al.
Veröffentlicht: (2025)
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
von: Deng, Boyang, et al.
Veröffentlicht: (2025)
von: Deng, Boyang, et al.
Veröffentlicht: (2025)
Beyond Core and Penumbra: Bi-Temporal Image-Driven Stroke Evolution Analysis
von: Rahman, Md Sazidur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Sazidur, et al.
Veröffentlicht: (2026)
PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
UAVs and Birds: Enhancing Short-Range Navigation through Budgerigar Flight Studies
von: Rahman, Md. Mahmudur, et al.
Veröffentlicht: (2023)
von: Rahman, Md. Mahmudur, et al.
Veröffentlicht: (2023)
Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
von: Alim, Md. Samiul, et al.
Veröffentlicht: (2025)
von: Alim, Md. Samiul, et al.
Veröffentlicht: (2025)
RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone
von: Munir, Mustafa, et al.
Veröffentlicht: (2024)
von: Munir, Mustafa, et al.
Veröffentlicht: (2024)
Beyond Symbolic Solving: Multi Chain-of-Thought Voting for Geometric Reasoning in Large Language Models
von: Siddique, Md. Abu Bakor, et al.
Veröffentlicht: (2026)
von: Siddique, Md. Abu Bakor, et al.
Veröffentlicht: (2026)
Lightweight Model for Poultry Disease Detection from Fecal Images Using Multi-Color Space Feature Optimization and Machine Learning
von: Islam, A. K. M. Shoriful, et al.
Veröffentlicht: (2025)
von: Islam, A. K. M. Shoriful, et al.
Veröffentlicht: (2025)
Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2025)
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2025)
TeaLeafVision: An Explainable and Robust Deep Learning Framework for Tea Leaf Disease Classification
von: Ahamed, Rafi, et al.
Veröffentlicht: (2026)
von: Ahamed, Rafi, et al.
Veröffentlicht: (2026)
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
von: Tashdeed, Ishmam, et al.
Veröffentlicht: (2025)
von: Tashdeed, Ishmam, et al.
Veröffentlicht: (2025)
Skin Disease Detection and Classification of Actinic Keratosis and Psoriasis Utilizing Deep Transfer Learning
von: Ahmmed, Fahud, et al.
Veröffentlicht: (2025)
von: Ahmmed, Fahud, et al.
Veröffentlicht: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2026)
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2026)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
von: Rahman, Md Maklachur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Maklachur, et al.
Veröffentlicht: (2026)
Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2025)
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2025)
A Bidirectional Siamese Recurrent Neural Network for Accurate Gait Recognition Using Body Landmarks
von: Progga, Proma Hossain, et al.
Veröffentlicht: (2024)
von: Progga, Proma Hossain, et al.
Veröffentlicht: (2024)
An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
von: Newaz, Asif, et al.
Veröffentlicht: (2025)
von: Newaz, Asif, et al.
Veröffentlicht: (2025)
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
von: Jiang, Roy, et al.
Veröffentlicht: (2026)
von: Jiang, Roy, et al.
Veröffentlicht: (2026)
Unsupervised Deep Learning Image Verification Method
von: Solomon, Enoch, et al.
Veröffentlicht: (2023)
von: Solomon, Enoch, et al.
Veröffentlicht: (2023)
Enhancing Bidirectional Sign Language Communication: Integrating YOLOv8 and NLP for Real-Time Gesture Recognition & Translation
von: Bhuiyan, Hasnat Jamil, et al.
Veröffentlicht: (2024)
von: Bhuiyan, Hasnat Jamil, et al.
Veröffentlicht: (2024)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
von: Paul, Dhiman, et al.
Veröffentlicht: (2024)
Position: Towards Implicit Prompt For Text-To-Image Models
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models
von: Jin, Di, et al.
Veröffentlicht: (2024)
von: Jin, Di, et al.
Veröffentlicht: (2024)
Deep Learning and Hybrid Approaches for Dynamic Scene Analysis, Object Detection and Motion Tracking
von: Alve, Shahran Rahman
Veröffentlicht: (2024)
von: Alve, Shahran Rahman
Veröffentlicht: (2024)
Autoencoder Based Face Verification System
von: Solomon, Enoch, et al.
Veröffentlicht: (2023)
von: Solomon, Enoch, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
FUSED-Net: Detecting Traffic Signs with Limited Data
von: Rahman, Md. Atiqur, et al.
Veröffentlicht: (2024) -
BeHGAN: Bengali Handwritten Word Generation from Plain Text Using Generative Adversarial Networks
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025) -
Intelligent Systems in Neuroimaging: Pioneering AI Techniques for Brain Tumor Detection
von: Islam, Md. Mohaiminul, et al.
Veröffentlicht: (2025) -
STAR: A Foundation Model-driven Framework for Robust Task Planning and Failure Recovery in Robotic Systems
von: Sakib, Md Sadman, et al.
Veröffentlicht: (2025) -
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
von: Chowdhury, Md Abtahi Majeed, et al.
Veröffentlicht: (2025)