DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yixuan, Wang, Yizhou, Tang, Shixiang, Wu, Wenhao, He, Tong, Ouyang, Wanli, Torr, Philip, Wu, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hulk: A Universal Knowledge Translator for Human-Centric Tasks
von: Wang, Yizhou, et al.
Veröffentlicht: (2023)
von: Wang, Yizhou, et al.
Veröffentlicht: (2023)
OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM
von: Wang, Jinhong, et al.
Veröffentlicht: (2025)
von: Wang, Jinhong, et al.
Veröffentlicht: (2025)
UniPAD: A Universal Pre-training Paradigm for Autonomous Driving
von: Yang, Honghui, et al.
Veröffentlicht: (2023)
von: Yang, Honghui, et al.
Veröffentlicht: (2023)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
von: Wu, Yixuan, et al.
Veröffentlicht: (2025)
von: Wu, Yixuan, et al.
Veröffentlicht: (2025)
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
FreeVA: Offline MLLM as Training-Free Video Assistant
von: Wu, Wenhao
Veröffentlicht: (2024)
von: Wu, Wenhao
Veröffentlicht: (2024)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
von: Chen, Lu, et al.
Veröffentlicht: (2025)
von: Chen, Lu, et al.
Veröffentlicht: (2025)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
von: Zhang, Sha, et al.
Veröffentlicht: (2024)
von: Zhang, Sha, et al.
Veröffentlicht: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
MLLM-based Discovery of Intrinsic Coordinates and Governing Equations from High-Dimensional Data
von: Li, Ruikun, et al.
Veröffentlicht: (2025)
von: Li, Ruikun, et al.
Veröffentlicht: (2025)
SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
von: Xu, Qing, et al.
Veröffentlicht: (2025)
von: Xu, Qing, et al.
Veröffentlicht: (2025)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
von: Xu, Shilin, et al.
Veröffentlicht: (2023)
von: Xu, Shilin, et al.
Veröffentlicht: (2023)
Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-identification
von: He, Weizhen, et al.
Veröffentlicht: (2024)
von: He, Weizhen, et al.
Veröffentlicht: (2024)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
von: Zhu, Haoyi, et al.
Veröffentlicht: (2023)
von: Zhu, Haoyi, et al.
Veröffentlicht: (2023)
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
von: Huang, Chenxi, et al.
Veröffentlicht: (2024)
von: Huang, Chenxi, et al.
Veröffentlicht: (2024)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
von: Su, Haoyang, et al.
Veröffentlicht: (2024)
von: Su, Haoyang, et al.
Veröffentlicht: (2024)
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
von: Wu, Wentao, et al.
Veröffentlicht: (2024)
von: Wu, Wentao, et al.
Veröffentlicht: (2024)
AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
von: Chao, Yuhao, et al.
Veröffentlicht: (2025)
von: Chao, Yuhao, et al.
Veröffentlicht: (2025)
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
von: Cao, Guiping, et al.
Veröffentlicht: (2025)
von: Cao, Guiping, et al.
Veröffentlicht: (2025)
Minimal Interaction Separated Tuning: A New Paradigm for Visual Adaptation
von: Tang, Ningyuan, et al.
Veröffentlicht: (2024)
von: Tang, Ningyuan, et al.
Veröffentlicht: (2024)
PIP-MM: Pre-Integrating Prompt Information into Visual Encoding via Existing MLLM Structures
von: Wu, Tianxiang, et al.
Veröffentlicht: (2024)
von: Wu, Tianxiang, et al.
Veröffentlicht: (2024)
PromptDet: A Lightweight 3D Object Detection Framework with LiDAR Prompts
von: Guo, Kun, et al.
Veröffentlicht: (2024)
von: Guo, Kun, et al.
Veröffentlicht: (2024)
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
von: Li, Fanxiao, et al.
Veröffentlicht: (2025)
von: Li, Fanxiao, et al.
Veröffentlicht: (2025)
Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions
von: He, Weizhen, et al.
Veröffentlicht: (2023)
von: He, Weizhen, et al.
Veröffentlicht: (2023)
MatchDet: A Collaborative Framework for Image Matching and Object Detection
von: Lai, Jinxiang, et al.
Veröffentlicht: (2023)
von: Lai, Jinxiang, et al.
Veröffentlicht: (2023)
R2Det: Exploring Relaxed Rotation Equivariance in 2D object detection
von: Wu, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Wu, Zhiqiang, et al.
Veröffentlicht: (2024)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
Holistic-Motion2D: Scalable Whole-body Human Motion Generation in 2D Space
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
von: Dai, Guangzhao, et al.
Veröffentlicht: (2024)
von: Dai, Guangzhao, et al.
Veröffentlicht: (2024)
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
RemoteDet-Mamba: A Hybrid Mamba-CNN Network for Multi-modal Object Detection in Remote Sensing Images
von: Ren, Kejun, et al.
Veröffentlicht: (2024)
von: Ren, Kejun, et al.
Veröffentlicht: (2024)
Point Transformer V3 Extreme: 1st Place Solution for 2024 Waymo Open Dataset Challenge in Semantic Segmentation
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2024)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
Revisiting Out-of-Distribution Detection in Real-time Object Detection: From Benchmark Pitfalls to a New Mitigation Paradigm
von: Wu, Changshun, et al.
Veröffentlicht: (2025)
von: Wu, Changshun, et al.
Veröffentlicht: (2025)
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
von: Chen, Junyi, et al.
Veröffentlicht: (2024)
von: Chen, Junyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hulk: A Universal Knowledge Translator for Human-Centric Tasks
von: Wang, Yizhou, et al.
Veröffentlicht: (2023) -
OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM
von: Wang, Jinhong, et al.
Veröffentlicht: (2025) -
UniPAD: A Universal Pre-training Paradigm for Autonomous Driving
von: Yang, Honghui, et al.
Veröffentlicht: (2023) -
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
von: Wu, Yixuan, et al.
Veröffentlicht: (2025) -
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)