MatViX: Multimodal Information Extraction from Visually Rich Articles
Fuente:
arXiv
Saved in:
| Main Authors: | Khalighinejad, Ghazal, Scott, Sharon, Liu, Ollie, Anderson, Kelly L., Stureborg, Rickard, Tyagi, Aman, Dhingra, Bhuwan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hierarchical Multi-Label Classification of Online Vaccine Concerns
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
A Platform for Investigating Public Health Content with Efficient Concern Classification
by: Li, Christopher, et al.
Published: (2025)
by: Li, Christopher, et al.
Published: (2025)
Extracting Polymer Nanocomposite Samples from Full-Length Documents
by: Khalighinejad, Ghazal, et al.
Published: (2024)
by: Khalighinejad, Ghazal, et al.
Published: (2024)
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
by: Fu, Deqing, et al.
Published: (2024)
by: Fu, Deqing, et al.
Published: (2024)
Document-as-Image Representations Fall Short for Scientific Retrieval
by: Khalighinejad, Ghazal, et al.
Published: (2026)
by: Khalighinejad, Ghazal, et al.
Published: (2026)
ChatShop: Interactive Information Seeking with Language Agents
by: Chen, Sanxing, et al.
Published: (2024)
by: Chen, Sanxing, et al.
Published: (2024)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
Tailoring Vaccine Messaging with Common-Ground Opinions
by: Stureborg, Rickard, et al.
Published: (2024)
by: Stureborg, Rickard, et al.
Published: (2024)
Atomic Self-Consistency for Better Long Form Generations
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)
by: Thirukovalluru, Raghuveer, et al.
Published: (2024)
Large Language Models are Inconsistent and Biased Evaluators
by: Stureborg, Rickard, et al.
Published: (2024)
by: Stureborg, Rickard, et al.
Published: (2024)
Your Large Language Models Are Leaving Fingerprints
by: McGovern, Hope, et al.
Published: (2024)
by: McGovern, Hope, et al.
Published: (2024)
Real-time Factuality Assessment from Adversarial Feedback
by: Chen, Sanxing, et al.
Published: (2024)
by: Chen, Sanxing, et al.
Published: (2024)
RVPO: Risk-Sensitive Alignment via Variance Regularization
by: Montero, Ivan, et al.
Published: (2026)
by: Montero, Ivan, et al.
Published: (2026)
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
Coding Agents are Effective Long-Context Processors
by: Cao, Weili, et al.
Published: (2026)
by: Cao, Weili, et al.
Published: (2026)
InData: Towards Secure Multi-Step, Tool-Based Data Analysis
by: K, Karthikeyan, et al.
Published: (2025)
by: K, Karthikeyan, et al.
Published: (2025)
Atomic Consistency Preference Optimization for Long-Form Question Answering
by: Chen, Jingfeng, et al.
Published: (2025)
by: Chen, Jingfeng, et al.
Published: (2025)
Dynamic Contexts for Generating Suggestion Questions in RAG Based Conversational Systems
by: Tayal, Anuja, et al.
Published: (2024)
by: Tayal, Anuja, et al.
Published: (2024)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
by: Chen, Sanxing, et al.
Published: (2025)
by: Chen, Sanxing, et al.
Published: (2025)
Calibrating Long-form Generations from Large Language Models
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals
by: Xie, Roy, et al.
Published: (2025)
by: Xie, Roy, et al.
Published: (2025)
Automated Benchmark Auditing for AI Agents and Large Language Models
by: Wang, Junlin, et al.
Published: (2026)
by: Wang, Junlin, et al.
Published: (2026)
Detecting Concrete Visual Tokens for Multimodal Machine Translation
by: Bowen, Braeden, et al.
Published: (2024)
by: Bowen, Braeden, et al.
Published: (2024)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
Interleaved Reasoning for Large Language Models via Reinforcement Learning
by: Xie, Roy, et al.
Published: (2025)
by: Xie, Roy, et al.
Published: (2025)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
by: Pala, Furkan, et al.
Published: (2024)
by: Pala, Furkan, et al.
Published: (2024)
RE$^2$: Region-Aware Relation Extraction from Visually Rich Documents
by: Ramu, Pritika, et al.
Published: (2023)
by: Ramu, Pritika, et al.
Published: (2023)
Enhancing Keyphrase Extraction from Academic Articles Using Section Structure Information
by: Zhang, Chengzhi, et al.
Published: (2025)
by: Zhang, Chengzhi, et al.
Published: (2025)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
by: Alqurnawi, Yahia, et al.
Published: (2026)
by: Alqurnawi, Yahia, et al.
Published: (2026)
ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once
by: Dhingra, Harnoor
Published: (2026)
by: Dhingra, Harnoor
Published: (2026)
Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff
by: Holsman, Maximilian, et al.
Published: (2025)
by: Holsman, Maximilian, et al.
Published: (2025)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
by: Ji, Yifan, et al.
Published: (2026)
by: Ji, Yifan, et al.
Published: (2026)
MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models
by: Belouadi, Jonas, et al.
Published: (2025)
by: Belouadi, Jonas, et al.
Published: (2025)
Similar Items
-
Hierarchical Multi-Label Classification of Online Vaccine Concerns
by: Zhu, Chloe Qinyu, et al.
Published: (2024) -
A Platform for Investigating Public Health Content with Efficient Concern Classification
by: Li, Christopher, et al.
Published: (2025) -
Extracting Polymer Nanocomposite Samples from Full-Length Documents
by: Khalighinejad, Ghazal, et al.
Published: (2024) -
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
by: Fu, Deqing, et al.
Published: (2024) -
Document-as-Image Representations Fall Short for Scientific Retrieval
by: Khalighinejad, Ghazal, et al.
Published: (2026)