Efficient Multimodal Large Language Models: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Yizhang, Li, Jian, Liu, Yexin, Gu, Tianjun, Wu, Kai, Jiang, Zhengkai, He, Muyang, Zhao, Bo, Tan, Xin, Gan, Zhenye, Wang, Yabiao, Wang, Chengjie, Ma, Lizhuang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
SoftPatch+: Fully Unsupervised Anomaly Classification and Segmentation
von: Wang, Chengjie, et al.
Veröffentlicht: (2024)
von: Wang, Chengjie, et al.
Veröffentlicht: (2024)
Efficient Multimodal Learning from Data-centric Perspective
von: He, Muyang, et al.
Veröffentlicht: (2024)
von: He, Muyang, et al.
Veröffentlicht: (2024)
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
von: Wu, Kai, et al.
Veröffentlicht: (2024)
von: Wu, Kai, et al.
Veröffentlicht: (2024)
Improving Search Agent with One Line of Code
von: Li, Jian, et al.
Veröffentlicht: (2026)
von: Li, Jian, et al.
Veröffentlicht: (2026)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
von: Fan, Ke, et al.
Veröffentlicht: (2024)
von: Fan, Ke, et al.
Veröffentlicht: (2024)
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
von: Pan, Yanjie, et al.
Veröffentlicht: (2025)
von: Pan, Yanjie, et al.
Veröffentlicht: (2025)
Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection
von: Wang, Chengjie, et al.
Veröffentlicht: (2024)
von: Wang, Chengjie, et al.
Veröffentlicht: (2024)
A Comprehensive Library for Benchmarking Multi-class Visual Anomaly Detection
von: Zhang, Jiangning, et al.
Veröffentlicht: (2024)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2024)
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
von: He, Qingdong, et al.
Veröffentlicht: (2025)
von: He, Qingdong, et al.
Veröffentlicht: (2025)
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
von: He, Qingdong, et al.
Veröffentlicht: (2024)
von: He, Qingdong, et al.
Veröffentlicht: (2024)
AdapNet: Adaptive Noise-Based Network for Low-Quality Image Retrieval
von: Zhang, Sihe, et al.
Veröffentlicht: (2024)
von: Zhang, Sihe, et al.
Veröffentlicht: (2024)
MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
von: He, Haoyang, et al.
Veröffentlicht: (2024)
von: He, Haoyang, et al.
Veröffentlicht: (2024)
The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection
von: He, Qingdong, et al.
Veröffentlicht: (2025)
von: He, Qingdong, et al.
Veröffentlicht: (2025)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024)
von: Wang, Sen, et al.
Veröffentlicht: (2024)
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
von: Hu, Teng, et al.
Veröffentlicht: (2024)
von: Hu, Teng, et al.
Veröffentlicht: (2024)
Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
von: He, Liren, et al.
Veröffentlicht: (2024)
von: He, Liren, et al.
Veröffentlicht: (2024)
DMAD: Dual Memory Bank for Real-World Anomaly Detection
von: Hu, Jianlong, et al.
Veröffentlicht: (2024)
von: Hu, Jianlong, et al.
Veröffentlicht: (2024)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
von: Zhang, Jiangning, et al.
Veröffentlicht: (2025)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2025)
Vision-language models lag human performance on physical dynamics and intent reasoning
von: Gu, Tianjun, et al.
Veröffentlicht: (2026)
von: Gu, Tianjun, et al.
Veröffentlicht: (2026)
PSPU: Enhanced Positive and Unlabeled Learning by Leveraging Pseudo Supervision
von: Wang, Chengjie, et al.
Veröffentlicht: (2024)
von: Wang, Chengjie, et al.
Veröffentlicht: (2024)
RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
von: Ding, Hang, et al.
Veröffentlicht: (2025)
von: Ding, Hang, et al.
Veröffentlicht: (2025)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning
von: He, Qingdong, et al.
Veröffentlicht: (2024)
von: He, Qingdong, et al.
Veröffentlicht: (2024)
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
von: Wang, Qilin, et al.
Veröffentlicht: (2024)
von: Wang, Qilin, et al.
Veröffentlicht: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
von: Fu, Chencan, et al.
Veröffentlicht: (2024)
von: Fu, Chencan, et al.
Veröffentlicht: (2024)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
Semantic Frame Interpolation
von: Hong, Yijia, et al.
Veröffentlicht: (2025)
von: Hong, Yijia, et al.
Veröffentlicht: (2025)
SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
von: Li, Jian, et al.
Veröffentlicht: (2026)
von: Li, Jian, et al.
Veröffentlicht: (2026)
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation
von: Gu, Tianjun, et al.
Veröffentlicht: (2025)
von: Gu, Tianjun, et al.
Veröffentlicht: (2025)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
von: Fan, Ke, et al.
Veröffentlicht: (2024)
von: Fan, Ke, et al.
Veröffentlicht: (2024)
UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer
von: Wang, Haoxuan, et al.
Veröffentlicht: (2025)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2025)
Motif Counting in Complex Networks: A Comprehensive Survey
von: Yin, Haozhe, et al.
Veröffentlicht: (2025)
von: Yin, Haozhe, et al.
Veröffentlicht: (2025)
Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era
von: Zhu, Wenbing, et al.
Veröffentlicht: (2025)
von: Zhu, Wenbing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
von: Jin, Yizhang, et al.
Veröffentlicht: (2024) -
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024) -
SoftPatch+: Fully Unsupervised Anomaly Classification and Segmentation
von: Wang, Chengjie, et al.
Veröffentlicht: (2024) -
Efficient Multimodal Learning from Data-centric Perspective
von: He, Muyang, et al.
Veröffentlicht: (2024) -
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
von: Wu, Kai, et al.
Veröffentlicht: (2024)