PIP-MM: Pre-Integrating Prompt Information into Visual Encoding via Existing MLLM Structures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Tianxiang, Nie, Minxin, Cao, Ziqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings
von: Dai, Tianxiang, et al.
Veröffentlicht: (2026)
von: Dai, Tianxiang, et al.
Veröffentlicht: (2026)
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
Elysium: Exploring Object-level Perception in Videos via MLLM
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training
von: Nie, Yunshuang, et al.
Veröffentlicht: (2026)
von: Nie, Yunshuang, et al.
Veröffentlicht: (2026)
InstructX: Towards Unified Visual Editing with MLLM Guidance
von: Mou, Chong, et al.
Veröffentlicht: (2025)
von: Mou, Chong, et al.
Veröffentlicht: (2025)
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
von: Sun, Yueru, et al.
Veröffentlicht: (2026)
von: Sun, Yueru, et al.
Veröffentlicht: (2026)
Structure Causal Models and LLMs Integration in Medical Visual Question Answering
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
von: Xu, Zibo, et al.
Veröffentlicht: (2025)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
von: Li, Xu, et al.
Veröffentlicht: (2025)
von: Li, Xu, et al.
Veröffentlicht: (2025)
Robust MLLM Unlearning via Visual Knowledge Distillation
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
von: Lee, Sua, et al.
Veröffentlicht: (2026)
von: Lee, Sua, et al.
Veröffentlicht: (2026)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
von: Wu, Yongjian, et al.
Veröffentlicht: (2024)
von: Wu, Yongjian, et al.
Veröffentlicht: (2024)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
SDPT: Synchronous Dual Prompt Tuning for Fusion-based Visual-Language Pre-trained Models
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
von: Munasinghe, Shehan, et al.
Veröffentlicht: (2024)
von: Munasinghe, Shehan, et al.
Veröffentlicht: (2024)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
von: Wu, Yixuan, et al.
Veröffentlicht: (2024)
von: Wu, Yixuan, et al.
Veröffentlicht: (2024)
MambaRefine-YOLO: A Dual-Modality Small Object Detector for UAV Imagery
von: Cao, Shuyu, et al.
Veröffentlicht: (2025)
von: Cao, Shuyu, et al.
Veröffentlicht: (2025)
MM-IFEngine: Towards Multimodal Instruction Following
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
von: Xin, Yi, et al.
Veröffentlicht: (2025)
von: Xin, Yi, et al.
Veröffentlicht: (2025)
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
von: Zhao, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhao, Pengfei, et al.
Veröffentlicht: (2025)
D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning
von: Zhang, Evelyn, et al.
Veröffentlicht: (2025)
von: Zhang, Evelyn, et al.
Veröffentlicht: (2025)
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
von: Yao, Lei, et al.
Veröffentlicht: (2026)
von: Yao, Lei, et al.
Veröffentlicht: (2026)
MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence
von: Yin, Xingyilang, et al.
Veröffentlicht: (2026)
von: Yin, Xingyilang, et al.
Veröffentlicht: (2026)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
von: Choi, Jae Young, et al.
Veröffentlicht: (2026)
von: Choi, Jae Young, et al.
Veröffentlicht: (2026)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
PIP: Prototypes-Injected Prompt for Federated Class Incremental Learning
von: Ma'sum, Muhammad Anwar, et al.
Veröffentlicht: (2024)
von: Ma'sum, Muhammad Anwar, et al.
Veröffentlicht: (2024)
VPTracker: Global Vision-Language Tracking via Visual Prompt
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
von: Song, Qi, et al.
Veröffentlicht: (2025)
von: Song, Qi, et al.
Veröffentlicht: (2025)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
von: Zhong, Liangyu, et al.
Veröffentlicht: (2025)
von: Zhong, Liangyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025) -
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024) -
Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings
von: Dai, Tianxiang, et al.
Veröffentlicht: (2026) -
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
von: Jiang, Songtao, et al.
Veröffentlicht: (2024) -
Elysium: Exploring Object-level Perception in Videos via MLLM
von: Wang, Han, et al.
Veröffentlicht: (2024)