Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
Fuente:
arXiv
Saved in:
| Main Authors: | Sapkota, Ranjan, Karkee, Manoj |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Zero-Shot Automatic Annotation and Instance Segmentation using LLM-Generated Datasets: Eliminating Field Imaging and Manual Annotation for Deep Learning Model Development
by: Sapkota, Ranjan, et al.
Published: (2024)
by: Sapkota, Ranjan, et al.
Published: (2024)
The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Generative AI in Agriculture: Creating Image Datasets Using DALL.E's Advanced Large Language Model Capabilities
by: Sapkota, Ranjan, et al.
Published: (2023)
by: Sapkota, Ranjan, et al.
Published: (2023)
YOLO11 and Vision Transformers based 3D Pose Estimation of Immature Green Fruits in Commercial Apple Orchards for Robotic Thinning
by: Sapkota, Ranjan, et al.
Published: (2024)
by: Sapkota, Ranjan, et al.
Published: (2024)
A Review of 3D Object Detection with Vision-Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Comparing YOLOv11 and YOLOv8 for instance segmentation of occluded and non-occluded immature green fruits in complex orchard environment
by: Sapkota, Ranjan, et al.
Published: (2024)
by: Sapkota, Ranjan, et al.
Published: (2024)
YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation
by: Sapkota, Ranjan, et al.
Published: (2026)
by: Sapkota, Ranjan, et al.
Published: (2026)
Integrating YOLO11 and Convolution Block Attention Module for Multi-Season Segmentation of Tree Trunks and Branches in Commercial Apple Orchards
by: Sapkota, Ranjan, et al.
Published: (2024)
by: Sapkota, Ranjan, et al.
Published: (2024)
Comprehensive Performance Evaluation of YOLOv12, YOLO11, YOLOv10, YOLOv9 and YOLOv8 on Detecting and Counting Fruitlet in Complex Orchard Environments
by: Sapkota, Ranjan, et al.
Published: (2024)
by: Sapkota, Ranjan, et al.
Published: (2024)
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks
by: Roumeliotis, Konstantinos I., et al.
Published: (2025)
by: Roumeliotis, Konstantinos I., et al.
Published: (2025)
YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
by: Ashqar, Huthaifa I., et al.
Published: (2024)
by: Ashqar, Huthaifa I., et al.
Published: (2024)
Comparing YOLOv8 and Mask R-CNN for instance segmentation in complex orchard environments
by: Sapkota, Ranjan, et al.
Published: (2023)
by: Sapkota, Ranjan, et al.
Published: (2023)
RF-DETR Object Detection vs YOLOv12 : A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Immature Green Apple Detection and Sizing in Commercial Orchards using YOLOv8 and Shape Fitting Techniques
by: Sapkota, Ranjan, et al.
Published: (2023)
by: Sapkota, Ranjan, et al.
Published: (2023)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
by: Van, Minh-Hao, et al.
Published: (2025)
by: Van, Minh-Hao, et al.
Published: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
by: Zhong, Weihong, et al.
Published: (2024)
by: Zhong, Weihong, et al.
Published: (2024)
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
by: Ranjan, Mukul, et al.
Published: (2026)
by: Ranjan, Mukul, et al.
Published: (2026)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Multi-Object Hallucination in Vision-Language Models
by: Chen, Xuweiyi, et al.
Published: (2024)
by: Chen, Xuweiyi, et al.
Published: (2024)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
by: Li, Yanshu
Published: (2025)
by: Li, Yanshu
Published: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2025)
by: Woo, Sangmin, et al.
Published: (2025)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
by: Ren, Lingfeng, et al.
Published: (2026)
by: Ren, Lingfeng, et al.
Published: (2026)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
by: Wu, Jiulong, et al.
Published: (2025)
by: Wu, Jiulong, et al.
Published: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
by: Ye, Zekai, et al.
Published: (2025)
by: Ye, Zekai, et al.
Published: (2025)
First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
by: Ha, Jiwoo, et al.
Published: (2026)
by: Ha, Jiwoo, et al.
Published: (2026)
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
by: Cahyawijaya, Samuel, et al.
Published: (2026)
by: Cahyawijaya, Samuel, et al.
Published: (2026)
Model Composition for Multimodal Large Language Models
by: Chen, Chi, et al.
Published: (2024)
by: Chen, Chi, et al.
Published: (2024)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
by: Deniz, Omer Faruk, et al.
Published: (2026)
by: Deniz, Omer Faruk, et al.
Published: (2026)
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
by: Xu, Qipan, et al.
Published: (2025)
by: Xu, Qipan, et al.
Published: (2025)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Benchmarking Multimodal Large Language Models for Face Recognition
by: Shahreza, Hatef Otroshi, et al.
Published: (2025)
by: Shahreza, Hatef Otroshi, et al.
Published: (2025)
Similar Items
-
Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
by: Sapkota, Ranjan, et al.
Published: (2025) -
Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10
by: Sapkota, Ranjan, et al.
Published: (2025) -
Zero-Shot Automatic Annotation and Instance Segmentation using LLM-Generated Datasets: Eliminating Field Imaging and Manual Annotation for Deep Learning Model Development
by: Sapkota, Ranjan, et al.
Published: (2024) -
The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
by: Sapkota, Ranjan, et al.
Published: (2025) -
Generative AI in Agriculture: Creating Image Datasets Using DALL.E's Advanced Large Language Model Capabilities
by: Sapkota, Ranjan, et al.
Published: (2023)