HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Wenqiao, Lin, Tianwei, Liu, Jiang, Shu, Fangxun, Li, Haoyuan, Zhang, Lei, Wanggui, He, Zhou, Hao, Lv, Zheqi, Jiang, Hao, Li, Juncheng, Tang, Siliang, Zhuang, Yueting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
di: Lin, Tianwei, et al.
Pubblicazione: (2024)
di: Lin, Tianwei, et al.
Pubblicazione: (2024)
Fast Thinking for Large Language Models
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
di: Liu, Jiang, et al.
Pubblicazione: (2024)
di: Liu, Jiang, et al.
Pubblicazione: (2024)
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
di: Huang, Hongzhe, et al.
Pubblicazione: (2024)
di: Huang, Hongzhe, et al.
Pubblicazione: (2024)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
di: Cao, Jie, et al.
Pubblicazione: (2025)
di: Cao, Jie, et al.
Pubblicazione: (2025)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
di: Dai, Yang, et al.
Pubblicazione: (2025)
di: Dai, Yang, et al.
Pubblicazione: (2025)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
di: Gao, Minghe, et al.
Pubblicazione: (2023)
di: Gao, Minghe, et al.
Pubblicazione: (2023)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
di: Lin, Tianwei, et al.
Pubblicazione: (2025)
di: Lin, Tianwei, et al.
Pubblicazione: (2025)
Enhancing Post-Training Quantization via Future Activation Awareness
di: Lv, Zheqi, et al.
Pubblicazione: (2026)
di: Lv, Zheqi, et al.
Pubblicazione: (2026)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
Unified Personalized Understanding, Generating and Editing
di: Zhong, Yu, et al.
Pubblicazione: (2026)
di: Zhong, Yu, et al.
Pubblicazione: (2026)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
di: Zheng, Haoyu, et al.
Pubblicazione: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
di: Di, Shangzhe, et al.
Pubblicazione: (2025)
di: Di, Shangzhe, et al.
Pubblicazione: (2025)
IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization
di: Cao, Jie, et al.
Pubblicazione: (2024)
di: Cao, Jie, et al.
Pubblicazione: (2024)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
di: Yin, Aoxiong, et al.
Pubblicazione: (2024)
di: Yin, Aoxiong, et al.
Pubblicazione: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
di: Wang, Wei, et al.
Pubblicazione: (2026)
di: Wang, Wei, et al.
Pubblicazione: (2026)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
di: Fan, Zhaoyu, et al.
Pubblicazione: (2025)
di: Fan, Zhaoyu, et al.
Pubblicazione: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
di: Qin, Bosheng, et al.
Pubblicazione: (2023)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
di: Li, Juncheng, et al.
Pubblicazione: (2023)
di: Li, Juncheng, et al.
Pubblicazione: (2023)
Bridging Local Details and Global Context in Text-Attributed Graphs
di: Wang, Yaoke, et al.
Pubblicazione: (2024)
di: Wang, Yaoke, et al.
Pubblicazione: (2024)
DuetRAG: Collaborative Retrieval-Augmented Generation
di: Jiao, Dian, et al.
Pubblicazione: (2024)
di: Jiao, Dian, et al.
Pubblicazione: (2024)
InstructSAM: Segment Any Instance with Any Instructions
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
di: Gao, Mingjian, et al.
Pubblicazione: (2026)
di: Gao, Mingjian, et al.
Pubblicazione: (2026)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
di: Qian, Long, et al.
Pubblicazione: (2024)
di: Qian, Long, et al.
Pubblicazione: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
di: Miao, Bingchen, et al.
Pubblicazione: (2024)
di: Miao, Bingchen, et al.
Pubblicazione: (2024)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
di: Ge, Zhiqi, et al.
Pubblicazione: (2024)
di: Ge, Zhiqi, et al.
Pubblicazione: (2024)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
di: Zhang, Guanghao, et al.
Pubblicazione: (2025)
di: Zhang, Guanghao, et al.
Pubblicazione: (2025)
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
di: Yu, Qifan, et al.
Pubblicazione: (2024)
di: Yu, Qifan, et al.
Pubblicazione: (2024)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
di: Liang, Han, et al.
Pubblicazione: (2024)
di: Liang, Han, et al.
Pubblicazione: (2024)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
di: Qiu, Haiyi, et al.
Pubblicazione: (2026)
di: Qiu, Haiyi, et al.
Pubblicazione: (2026)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
di: Wang, Keyu, et al.
Pubblicazione: (2026)
di: Wang, Keyu, et al.
Pubblicazione: (2026)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
di: Huang, Ziwei, et al.
Pubblicazione: (2024)
di: Huang, Ziwei, et al.
Pubblicazione: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
di: Li, Feng, et al.
Pubblicazione: (2024)
di: Li, Feng, et al.
Pubblicazione: (2024)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
di: Andersland, Michael
Pubblicazione: (2024)
di: Andersland, Michael
Pubblicazione: (2024)
Documenti analoghi
-
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
di: Lin, Tianwei, et al.
Pubblicazione: (2024) -
Fast Thinking for Large Language Models
di: Zheng, Haoyu, et al.
Pubblicazione: (2025) -
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
di: Liu, Jiang, et al.
Pubblicazione: (2024) -
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
di: Huang, Hongzhe, et al.
Pubblicazione: (2024) -
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
di: Zheng, Haoyu, et al.
Pubblicazione: (2024)