Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
Fuente:
arXiv
Salvato in:
| Autori principali: | Xing, Fuyu, Wang, Zimu, Wang, Wei, Zhang, Haiyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Event Extraction from Speech with Contextual Clues
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
di: Guan, Qinghao, et al.
Pubblicazione: (2026)
di: Guan, Qinghao, et al.
Pubblicazione: (2026)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
di: Zhang, Lei, et al.
Pubblicazione: (2025)
di: Zhang, Lei, et al.
Pubblicazione: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
Movie101v2: Improved Movie Narration Benchmark
di: Yue, Zihao, et al.
Pubblicazione: (2024)
di: Yue, Zihao, et al.
Pubblicazione: (2024)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
di: Chen, Qi, et al.
Pubblicazione: (2025)
di: Chen, Qi, et al.
Pubblicazione: (2025)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
di: Zhang, Qintong, et al.
Pubblicazione: (2024)
di: Zhang, Qintong, et al.
Pubblicazione: (2024)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
di: Lu, Jinghui, et al.
Pubblicazione: (2024)
di: Lu, Jinghui, et al.
Pubblicazione: (2024)
SpeechEE: A Novel Benchmark for Speech Event Extraction
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
di: Yuan, Xiang, et al.
Pubblicazione: (2025)
di: Yuan, Xiang, et al.
Pubblicazione: (2025)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
di: Zhang, Hanlei, et al.
Pubblicazione: (2024)
di: Zhang, Hanlei, et al.
Pubblicazione: (2024)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
di: Zhang, Chong, et al.
Pubblicazione: (2024)
di: Zhang, Chong, et al.
Pubblicazione: (2024)
MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
di: Peng, Yuezhang, et al.
Pubblicazione: (2025)
di: Peng, Yuezhang, et al.
Pubblicazione: (2025)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
di: Hu, Anwen, et al.
Pubblicazione: (2023)
di: Hu, Anwen, et al.
Pubblicazione: (2023)
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
di: Sun, Tao, et al.
Pubblicazione: (2025)
di: Sun, Tao, et al.
Pubblicazione: (2025)
Improving Gloss-free Sign Language Translation by Reducing Representation Density
di: Ye, Jinhui, et al.
Pubblicazione: (2024)
di: Ye, Jinhui, et al.
Pubblicazione: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
di: Maji, Arijit, et al.
Pubblicazione: (2025)
di: Maji, Arijit, et al.
Pubblicazione: (2025)
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection
di: Hachmeier, Simon, et al.
Pubblicazione: (2024)
di: Hachmeier, Simon, et al.
Pubblicazione: (2024)
RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture
di: Wang, Haofeng, et al.
Pubblicazione: (2025)
di: Wang, Haofeng, et al.
Pubblicazione: (2025)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
di: Zhang, Hanlei, et al.
Pubblicazione: (2025)
di: Zhang, Hanlei, et al.
Pubblicazione: (2025)
Double Mixture: Towards Continual Event Detection from Speech
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production
di: Feng, Liqian, et al.
Pubblicazione: (2025)
di: Feng, Liqian, et al.
Pubblicazione: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
di: Zhang, Bo, et al.
Pubblicazione: (2024)
di: Zhang, Bo, et al.
Pubblicazione: (2024)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
di: Wang, Bingbing, et al.
Pubblicazione: (2025)
di: Wang, Bingbing, et al.
Pubblicazione: (2025)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
di: Gan, Chengguang, et al.
Pubblicazione: (2025)
di: Gan, Chengguang, et al.
Pubblicazione: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
di: Luo, Wen, et al.
Pubblicazione: (2024)
di: Luo, Wen, et al.
Pubblicazione: (2024)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
di: Lu, Jinghui, et al.
Pubblicazione: (2025)
di: Lu, Jinghui, et al.
Pubblicazione: (2025)
ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
Towards Event-oriented Long Video Understanding
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
di: Sun, Huatuan, et al.
Pubblicazione: (2025)
di: Sun, Huatuan, et al.
Pubblicazione: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
EventCast: Hybrid Demand Forecasting in E-Commerce with LLM-Based Event Knowledge
di: Hu, Congcong, et al.
Pubblicazione: (2026)
di: Hu, Congcong, et al.
Pubblicazione: (2026)
Retrieval-Augmented Multimodal Model for Fake News Detection
di: Li, Yiheng, et al.
Pubblicazione: (2026)
di: Li, Yiheng, et al.
Pubblicazione: (2026)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Event Extraction from Speech with Contextual Clues
di: Kang, Jingqi, et al.
Pubblicazione: (2024) -
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
di: Cui, Shiyao, et al.
Pubblicazione: (2025) -
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
di: Guan, Qinghao, et al.
Pubblicazione: (2026) -
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
di: Zhang, Lei, et al.
Pubblicazione: (2025) -
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
di: Zhao, Fei, et al.
Pubblicazione: (2025)