A Survey of Multimodal Large Language Model from A Data-centric Perspective
Fuente:
arXiv
Guardado en:
| Autores principales: | Bai, Tianyi, Liang, Hao, Wan, Binwang, Xu, Yanran, Li, Xi, Li, Shiyu, Yang, Ling, Li, Bozhou, Wang, Yifan, Cui, Bin, Huang, Ping, Shan, Jiulong, He, Conghui, Yuan, Binhang, Zhang, Wentao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
por: Liang, Hao, et al.
Publicado: (2024)
por: Liang, Hao, et al.
Publicado: (2024)
PTA: Enhancing Multimodal Sentiment Analysis through Pipelined Prediction and Translation-based Alignment
por: Song, Shezheng, et al.
Publicado: (2024)
por: Song, Shezheng, et al.
Publicado: (2024)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
por: Hao, Shengyu, et al.
Publicado: (2024)
por: Hao, Shengyu, et al.
Publicado: (2024)
Deepfake Detection: A Comprehensive Survey from the Reliability Perspective
por: Wang, Tianyi, et al.
Publicado: (2022)
por: Wang, Tianyi, et al.
Publicado: (2022)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
por: Wan, Ninghao, et al.
Publicado: (2026)
por: Wan, Ninghao, et al.
Publicado: (2026)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
por: Song, Shezheng, et al.
Publicado: (2023)
por: Song, Shezheng, et al.
Publicado: (2023)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
por: Wang, Bingbing, et al.
Publicado: (2025)
por: Wang, Bingbing, et al.
Publicado: (2025)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
por: Cui, Shiyao, et al.
Publicado: (2023)
por: Cui, Shiyao, et al.
Publicado: (2023)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
por: Wang, Bing, et al.
Publicado: (2025)
por: Wang, Bing, et al.
Publicado: (2025)
Explainable Multimodal Emotion Recognition
por: Lian, Zheng, et al.
Publicado: (2023)
por: Lian, Zheng, et al.
Publicado: (2023)
A Survey of Information Disorder on Video-Sharing Platforms
por: Li, Meiyu, et al.
Publicado: (2025)
por: Li, Meiyu, et al.
Publicado: (2025)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
por: Wang, Bing, et al.
Publicado: (2025)
por: Wang, Bing, et al.
Publicado: (2025)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
por: Wang, Zihua, et al.
Publicado: (2025)
por: Wang, Zihua, et al.
Publicado: (2025)
Retrieval-Augmented Multimodal Model for Fake News Detection
por: Li, Yiheng, et al.
Publicado: (2026)
por: Li, Yiheng, et al.
Publicado: (2026)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
por: Li, Zhu, et al.
Publicado: (2025)
por: Li, Zhu, et al.
Publicado: (2025)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
por: Wu, Yi, et al.
Publicado: (2025)
por: Wu, Yi, et al.
Publicado: (2025)
THE WASTIVE: An Interactive Ebb and Flow of Digital Fabrication Waste
por: Shan, Yifan, et al.
Publicado: (2025)
por: Shan, Yifan, et al.
Publicado: (2025)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
por: Guan, Qinghao, et al.
Publicado: (2026)
por: Guan, Qinghao, et al.
Publicado: (2026)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
por: Yuan, Bo, et al.
Publicado: (2024)
por: Yuan, Bo, et al.
Publicado: (2024)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
por: Liu, Peipei, et al.
Publicado: (2023)
por: Liu, Peipei, et al.
Publicado: (2023)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
por: Zhang, Qintong, et al.
Publicado: (2024)
por: Zhang, Qintong, et al.
Publicado: (2024)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
por: Qian, Fan, et al.
Publicado: (2024)
por: Qian, Fan, et al.
Publicado: (2024)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
por: Li, Jia, et al.
Publicado: (2025)
por: Li, Jia, et al.
Publicado: (2025)
Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis
por: Liu, Hao, et al.
Publicado: (2025)
por: Liu, Hao, et al.
Publicado: (2025)
A Survey on Image-text Multimodal Models
por: Guo, Ruifeng, et al.
Publicado: (2023)
por: Guo, Ruifeng, et al.
Publicado: (2023)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2024)
por: Wang, Bing, et al.
Publicado: (2024)
ChemDFM-X: Towards Large Multimodal Model for Chemistry
por: Zhao, Zihan, et al.
Publicado: (2024)
por: Zhao, Zihan, et al.
Publicado: (2024)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
por: Cheng, Zhi-Qi, et al.
Publicado: (2024)
por: Cheng, Zhi-Qi, et al.
Publicado: (2024)
ViMo: Generating Motions from Casual Videos
por: Qiu, Liangdong, et al.
Publicado: (2024)
por: Qiu, Liangdong, et al.
Publicado: (2024)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
por: Zhao, Ruixiang, et al.
Publicado: (2024)
por: Zhao, Ruixiang, et al.
Publicado: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
por: Cui, Shiyao, et al.
Publicado: (2025)
por: Cui, Shiyao, et al.
Publicado: (2025)
Exploring the latent space of diffusion models directly through singular value decomposition
por: Wang, Li, et al.
Publicado: (2025)
por: Wang, Li, et al.
Publicado: (2025)
Scaling up Multimodal Pre-training for Sign Language Understanding
por: Zhou, Wengang, et al.
Publicado: (2024)
por: Zhou, Wengang, et al.
Publicado: (2024)
A Survey of Multimodal Composite Editing and Retrieval
por: Li, Suyan, et al.
Publicado: (2024)
por: Li, Suyan, et al.
Publicado: (2024)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
por: Li, Jia, et al.
Publicado: (2025)
por: Li, Jia, et al.
Publicado: (2025)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
por: Hu, Anwen, et al.
Publicado: (2023)
por: Hu, Anwen, et al.
Publicado: (2023)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
por: An, Wenbin, et al.
Publicado: (2025)
por: An, Wenbin, et al.
Publicado: (2025)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
por: Liu, Hui, et al.
Publicado: (2023)
por: Liu, Hui, et al.
Publicado: (2023)
Crafting Dynamic Virtual Activities with Advanced Multimodal Models
por: Li, Changyang, et al.
Publicado: (2024)
por: Li, Changyang, et al.
Publicado: (2024)
Ejemplares similares
-
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
por: Liang, Hao, et al.
Publicado: (2024) -
PTA: Enhancing Multimodal Sentiment Analysis through Pipelined Prediction and Translation-based Alignment
por: Song, Shezheng, et al.
Publicado: (2024) -
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
por: Hao, Shengyu, et al.
Publicado: (2024) -
Deepfake Detection: A Comprehensive Survey from the Reliability Perspective
por: Wang, Tianyi, et al.
Publicado: (2022) -
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
por: Wan, Ninghao, et al.
Publicado: (2026)