Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiang, Li, Bolin, Li, Haoyuan, Lin, Tianwei, Zhang, Wenqiao, Zhong, Tao, Yu, Zhelun, Wei, Jinghao, Cheng, Hao, He, Wanggui, Shu, Fangxun, Jiang, Hao, Lv, Zheqi, Li, Juncheng, Tang, Siliang, Zhuang, Yueting |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
by: Zhang, Wenqiao, et al.
Published: (2024)
by: Zhang, Wenqiao, et al.
Published: (2024)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
by: Lin, Tianwei, et al.
Published: (2024)
by: Lin, Tianwei, et al.
Published: (2024)
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
by: Dai, Yang, et al.
Published: (2025)
by: Dai, Yang, et al.
Published: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
by: Xiao, Wenyi, et al.
Published: (2024)
by: Xiao, Wenyi, et al.
Published: (2024)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
Fast Thinking for Large Language Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
by: Huang, Hongzhe, et al.
Published: (2024)
by: Huang, Hongzhe, et al.
Published: (2024)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
by: Gao, Minghe, et al.
Published: (2023)
by: Gao, Minghe, et al.
Published: (2023)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
by: Zhang, Guanghao, et al.
Published: (2025)
by: Zhang, Guanghao, et al.
Published: (2025)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024)
by: Huang, Ziwei, et al.
Published: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025)
by: Lin, Tianwei, et al.
Published: (2025)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
by: Gao, Minghe, et al.
Published: (2024)
by: Gao, Minghe, et al.
Published: (2024)
Unified Personalized Understanding, Generating and Editing
by: Zhong, Yu, et al.
Published: (2026)
by: Zhong, Yu, et al.
Published: (2026)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026)
by: Lv, Zheqi, et al.
Published: (2026)
SAG: Style-Aligned Article Generation via Model Collaboration
by: Xu, Chenning, et al.
Published: (2024)
by: Xu, Chenning, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
by: He, Wanggui, et al.
Published: (2024)
by: He, Wanggui, et al.
Published: (2024)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
by: Li, Sijing, et al.
Published: (2025)
by: Li, Sijing, et al.
Published: (2025)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
by: Yin, Aoxiong, et al.
Published: (2024)
by: Yin, Aoxiong, et al.
Published: (2024)
Bridging Local Details and Global Context in Text-Attributed Graphs
by: Wang, Yaoke, et al.
Published: (2024)
by: Wang, Yaoke, et al.
Published: (2024)
DuetRAG: Collaborative Retrieval-Augmented Generation
by: Jiao, Dian, et al.
Published: (2024)
by: Jiao, Dian, et al.
Published: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
by: Miao, Bingchen, et al.
Published: (2024)
by: Miao, Bingchen, et al.
Published: (2024)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
by: Qiu, Haiyi, et al.
Published: (2026)
by: Qiu, Haiyi, et al.
Published: (2026)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025)
by: Yu, Binhe, et al.
Published: (2025)
IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization
by: Cao, Jie, et al.
Published: (2024)
by: Cao, Jie, et al.
Published: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
by: Wang, Keyu, et al.
Published: (2026)
by: Wang, Keyu, et al.
Published: (2026)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
by: Qian, Long, et al.
Published: (2024)
by: Qian, Long, et al.
Published: (2024)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
WorldGPT: Empowering LLM as Multimodal World Model
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models
by: Dai, Guangyu, et al.
Published: (2025)
by: Dai, Guangyu, et al.
Published: (2025)
Similar Items
-
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
by: Zhang, Wenqiao, et al.
Published: (2024) -
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
by: Lin, Tianwei, et al.
Published: (2024) -
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
by: Dai, Yang, et al.
Published: (2025) -
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025) -
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
by: Zheng, Haoyu, et al.
Published: (2024)