Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Soyeon Caren, Cao, Feiqi, Poon, Josiah, Navigli, Roberto |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-game Toxic Language Detection: Shared Task and Attention Residuals
by: Jia, Yuanzhe, et al.
Published: (2022)
by: Jia, Yuanzhe, et al.
Published: (2022)
Game-MUG: Multimodal Oriented Game Situation Understanding and Commentary Generation Dataset
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024)
by: Ng, Thye Shan, et al.
Published: (2024)
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
PEACH: Pretrained-embedding Explanation Across Contextual and Hierarchical Structure
by: Cao, Feiqi, et al.
Published: (2024)
by: Cao, Feiqi, et al.
Published: (2024)
Local Interpretations for Explainable Natural Language Processing: A Survey
by: Luo, Siwen, et al.
Published: (2021)
by: Luo, Siwen, et al.
Published: (2021)
A Survey of Large Language Models in Finance (FinLLMs)
by: Lee, Jean, et al.
Published: (2024)
by: Lee, Jean, et al.
Published: (2024)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
TriG-NER: Triplet-Grid Framework for Discontinuous Named Entity Recognition
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025)
by: Dai, Yue, et al.
Published: (2025)
MSG-Chart: Multimodal Scene Graph for ChartQA
by: Dai, Yue, et al.
Published: (2024)
by: Dai, Yue, et al.
Published: (2024)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)
by: Akasaka, Muku, et al.
Published: (2026)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Do Large Language Models Understand Word Senses?
by: Meconi, Domenico, et al.
Published: (2025)
by: Meconi, Domenico, et al.
Published: (2025)
MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models
by: Lee, Suchan, et al.
Published: (2025)
by: Lee, Suchan, et al.
Published: (2025)
Graph Neural Networks for Text Classification: A Survey
by: Wang, Kunze, et al.
Published: (2023)
by: Wang, Kunze, et al.
Published: (2023)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
by: Xiang, Biao, et al.
Published: (2026)
by: Xiang, Biao, et al.
Published: (2026)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
by: Pisano, Raffaele, et al.
Published: (2026)
by: Pisano, Raffaele, et al.
Published: (2026)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
by: Yuan, Angela Yifei, et al.
Published: (2025)
by: Yuan, Angela Yifei, et al.
Published: (2025)
Towards Robust Instruction Tuning on Multimodal Large Language Models
by: Han, Wei, et al.
Published: (2024)
by: Han, Wei, et al.
Published: (2024)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
by: Park, Jiwon, et al.
Published: (2025)
by: Park, Jiwon, et al.
Published: (2025)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
Beyond Accuracy Optimization: Computer Vision Losses for Large Language Model Fine-Tuning
by: Cambrin, Daniele Rege, et al.
Published: (2024)
by: Cambrin, Daniele Rege, et al.
Published: (2024)
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
by: Li, Haoyi, et al.
Published: (2025)
by: Li, Haoyi, et al.
Published: (2025)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
by: Cao, Qinglong, et al.
Published: (2026)
by: Cao, Qinglong, et al.
Published: (2026)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
by: Tedeschi, Simone, et al.
Published: (2024)
by: Tedeschi, Simone, et al.
Published: (2024)
DeFrame: Debiasing Large Language Models Against Framing Effects
by: Lim, Kahee, et al.
Published: (2026)
by: Lim, Kahee, et al.
Published: (2026)
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
Can a Unimodal Language Agent Provide Preferences to Tune a Multimodal Vision-Language Model?
by: Mim, Sazia Tabasum, et al.
Published: (2026)
by: Mim, Sazia Tabasum, et al.
Published: (2026)
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
by: Scirè, Alessandro, et al.
Published: (2024)
by: Scirè, Alessandro, et al.
Published: (2024)
Similar Items
-
In-game Toxic Language Detection: Shared Task and Attention Residuals
by: Jia, Yuanzhe, et al.
Published: (2022) -
Game-MUG: Multimodal Oriented Game Situation Understanding and Commentary Generation Dataset
by: Zhang, Zhihao, et al.
Published: (2024) -
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024) -
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024) -
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
by: Li, Yan, et al.
Published: (2024)