Contextual Object Detection with Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zang, Yuhang, Li, Wei, Han, Jun, Zhou, Kaiyang, Loy, Chen Change |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023)
by: Xie, Jiahao, et al.
Published: (2023)
Can Multimodal Large Language Models Truly Understand Small Objects?
by: Han, Fujun, et al.
Published: (2026)
by: Han, Fujun, et al.
Published: (2026)
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
by: Lan, Yushi, et al.
Published: (2024)
by: Lan, Yushi, et al.
Published: (2024)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
HippoCamp: Benchmarking Contextual Agents on Personal Computers
by: Yang, Zhe, et al.
Published: (2026)
by: Yang, Zhe, et al.
Published: (2026)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
by: Jung, Woojun, et al.
Published: (2025)
by: Jung, Woojun, et al.
Published: (2025)
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
by: Chen, Junwen, et al.
Published: (2025)
by: Chen, Junwen, et al.
Published: (2025)
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models
by: Nabahirwa, Edwine, et al.
Published: (2025)
by: Nabahirwa, Edwine, et al.
Published: (2025)
MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
by: Liu, Ziyu, et al.
Published: (2024)
by: Liu, Ziyu, et al.
Published: (2024)
Do Pre-trained Vision-Language Models Encode Object States?
by: Newman, Kaleb, et al.
Published: (2024)
by: Newman, Kaleb, et al.
Published: (2024)
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
by: Ai, Wei, et al.
Published: (2026)
by: Ai, Wei, et al.
Published: (2026)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
by: Jiang, Xi, et al.
Published: (2024)
by: Jiang, Xi, et al.
Published: (2024)
SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models
by: Su, Yuhang, et al.
Published: (2026)
by: Su, Yuhang, et al.
Published: (2026)
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
by: Li, Jiansheng, et al.
Published: (2025)
by: Li, Jiansheng, et al.
Published: (2025)
Uncertainty Aware Human-machine Collaboration in Camouflaged Object Detection
by: Yang, Ziyue, et al.
Published: (2025)
by: Yang, Ziyue, et al.
Published: (2025)
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
by: Wen, Zichen, et al.
Published: (2026)
by: Wen, Zichen, et al.
Published: (2026)
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
by: Qian, Zekun, et al.
Published: (2026)
by: Qian, Zekun, et al.
Published: (2026)
F-LMM: Grounding Frozen Large Multimodal Models
by: Wu, Size, et al.
Published: (2024)
by: Wu, Size, et al.
Published: (2024)
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
by: Sun, Li, et al.
Published: (2024)
by: Sun, Li, et al.
Published: (2024)
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
by: Li, Hengzhuang, et al.
Published: (2025)
by: Li, Hengzhuang, et al.
Published: (2025)
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
by: Xia, Jiaer, et al.
Published: (2025)
by: Xia, Jiaer, et al.
Published: (2025)
Next Visual Granularity Generation
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
A Simple Aerial Detection Baseline of Multimodal Language Models
by: Li, Qingyun, et al.
Published: (2025)
by: Li, Qingyun, et al.
Published: (2025)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
by: Wang, Zhouxia, et al.
Published: (2024)
by: Wang, Zhouxia, et al.
Published: (2024)
AgroNVILA: Perception-Reasoning Decoupling for Multi-view Agricultural Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2026)
by: Zhang, Jiarui, et al.
Published: (2026)
VLANeXt: Recipes for Building Strong VLA Models
by: Wu, Xiao-Ming, et al.
Published: (2026)
by: Wu, Xiao-Ming, et al.
Published: (2026)
Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language Models
by: Li, Jiaqi, et al.
Published: (2024)
by: Li, Jiaqi, et al.
Published: (2024)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
by: Ashqar, Huthaifa I., et al.
Published: (2024)
by: Ashqar, Huthaifa I., et al.
Published: (2024)
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
by: Zhong, Jing, et al.
Published: (2025)
by: Zhong, Jing, et al.
Published: (2025)
Advancing Complex Video Object Segmentation via Progressive Concept Construction
by: Zhang, Zhixiong, et al.
Published: (2025)
by: Zhang, Zhixiong, et al.
Published: (2025)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Similar Items
-
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023) -
Can Multimodal Large Language Models Truly Understand Small Objects?
by: Han, Fujun, et al.
Published: (2026) -
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
by: Wang, Yuhan, et al.
Published: (2025) -
GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
by: Lan, Yushi, et al.
Published: (2024) -
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)