EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jiacheng, Jiao, Yang, Chen, Shaoxiang, Chen, Jingjing, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2026)
von: Li, Yian, et al.
Veröffentlicht: (2026)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
von: Han, Feng, et al.
Veröffentlicht: (2025)
von: Han, Feng, et al.
Veröffentlicht: (2025)
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model
von: Sun, Yiming, et al.
Veröffentlicht: (2024)
von: Sun, Yiming, et al.
Veröffentlicht: (2024)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
VRP-SAM: SAM with Visual Reference Prompt
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2025)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2025)
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
von: Chen, Jierun, et al.
Veröffentlicht: (2024)
von: Chen, Jierun, et al.
Veröffentlicht: (2024)
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
Visual Prompting in Multimodal Large Language Models: A Survey
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
von: Gong, Chao, et al.
Veröffentlicht: (2024)
von: Gong, Chao, et al.
Veröffentlicht: (2024)
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
von: Villa, Andrés, et al.
Veröffentlicht: (2025)
von: Villa, Andrés, et al.
Veröffentlicht: (2025)
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
von: Li, Bohao, et al.
Veröffentlicht: (2024)
von: Li, Bohao, et al.
Veröffentlicht: (2024)
Large Language Model for Lossless Image Compression with Visual Prompts
von: Du, Junhao, et al.
Veröffentlicht: (2025)
von: Du, Junhao, et al.
Veröffentlicht: (2025)
FedAPT: Federated Adversarial Prompt Tuning for Vision-Language Models
von: Zhai, Kun, et al.
Veröffentlicht: (2025)
von: Zhai, Kun, et al.
Veröffentlicht: (2025)
EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models
von: Peng, Xiaomeng, et al.
Veröffentlicht: (2026)
von: Peng, Xiaomeng, et al.
Veröffentlicht: (2026)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
Improving Visual Storytelling with Multimodal Large Language Models
von: Lin, Xiaochuan, et al.
Veröffentlicht: (2024)
von: Lin, Xiaochuan, et al.
Veröffentlicht: (2024)
Towards Training-free Multimodal Hate Localisation with Large Language Models
von: Sun, Yueming, et al.
Veröffentlicht: (2026)
von: Sun, Yueming, et al.
Veröffentlicht: (2026)
AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
von: Ying, Kaining, et al.
Veröffentlicht: (2024)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
von: Li, Duo, et al.
Veröffentlicht: (2025)
von: Li, Duo, et al.
Veröffentlicht: (2025)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
von: Choi, In Chong, et al.
Veröffentlicht: (2026)
von: Choi, In Chong, et al.
Veröffentlicht: (2026)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
von: Ding, Henghui, et al.
Veröffentlicht: (2026)
von: Ding, Henghui, et al.
Veröffentlicht: (2026)
Test-Time Computing for Referring Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
von: Jiao, Yang, et al.
Veröffentlicht: (2024) -
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2026) -
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
von: Jiao, Yang, et al.
Veröffentlicht: (2025) -
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
von: Han, Feng, et al.
Veröffentlicht: (2025) -
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)