Leveraging Vision Capabilities of Multimodal LLMs for Automated Data Extraction from Plots
Fuente:
arXiv
Saved in:
| Main Authors: | Polak, Maciej P., Morgan, Dane |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extracting Accurate Materials Data from Research Papers with Conversational Language Models and Prompt Engineering
by: Polak, Maciej P., et al.
Published: (2023)
by: Polak, Maciej P., et al.
Published: (2023)
Flexible, Model-Agnostic Method for Materials Data Extraction from Text Using General Purpose Language Models
by: Polak, Maciej P., et al.
Published: (2023)
by: Polak, Maciej P., et al.
Published: (2023)
Plots Unlock Time-Series Understanding in Multimodal Models
by: Daswani, Mayank, et al.
Published: (2024)
by: Daswani, Mayank, et al.
Published: (2024)
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
by: Sarmadi, Hamid, et al.
Published: (2025)
by: Sarmadi, Hamid, et al.
Published: (2025)
Accelerating Domain-Aware Electron Microscopy Analysis Using Deep Learning Models with Synthetic Data and Image-Wide Confidence Scoring
by: Lynch, Matthew J., et al.
Published: (2024)
by: Lynch, Matthew J., et al.
Published: (2024)
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
by: Chen, Zhuokun, et al.
Published: (2024)
by: Chen, Zhuokun, et al.
Published: (2024)
A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images
by: Chen, Yuxuan, et al.
Published: (2025)
by: Chen, Yuxuan, et al.
Published: (2025)
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
by: Roberts, Jonathan, et al.
Published: (2023)
by: Roberts, Jonathan, et al.
Published: (2023)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Predicting Performance of Object Detection Models in Electron Microscopy Using Random Forests
by: Li, Ni, et al.
Published: (2025)
by: Li, Ni, et al.
Published: (2025)
Automated Invoice Data Extraction: Using LLM and OCR
by: Khanchandani, Khushi, et al.
Published: (2025)
by: Khanchandani, Khushi, et al.
Published: (2025)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting Code
by: Lin, Haobo, et al.
Published: (2026)
by: Lin, Haobo, et al.
Published: (2026)
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction
by: Khan, Muhammad Tayyab, et al.
Published: (2024)
by: Khan, Muhammad Tayyab, et al.
Published: (2024)
Physics Informed Generative AI Enabling Labour Free Segmentation For Microscopy Analysis
by: Zahran, Salma, et al.
Published: (2026)
by: Zahran, Salma, et al.
Published: (2026)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2026)
by: Huang, Wenxuan, et al.
Published: (2026)
Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
by: Yao, Siyuan, et al.
Published: (2026)
by: Yao, Siyuan, et al.
Published: (2026)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
Image and Data Mining in Reticular Chemistry Using GPT-4V
by: Zheng, Zhiling, et al.
Published: (2023)
by: Zheng, Zhiling, et al.
Published: (2023)
KM-GPT: An Automated Pipeline for Reconstructing Individual Patient Data from Kaplan-Meier Plots
by: Zhao, Yao, et al.
Published: (2025)
by: Zhao, Yao, et al.
Published: (2025)
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
by: Onsu, Murat Arda, et al.
Published: (2025)
by: Onsu, Murat Arda, et al.
Published: (2025)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
by: Li, Yunxin, et al.
Published: (2023)
by: Li, Yunxin, et al.
Published: (2023)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction
by: Hu, Miaobo, et al.
Published: (2026)
by: Hu, Miaobo, et al.
Published: (2026)
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025)
by: Guo, Yifu, et al.
Published: (2025)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs
by: Lei, Zhi, et al.
Published: (2026)
by: Lei, Zhi, et al.
Published: (2026)
UlcerGPT: A Multimodal Approach Leveraging Large Language and Vision Models for Diabetic Foot Ulcer Image Transcription
by: Basiri, Reza, et al.
Published: (2024)
by: Basiri, Reza, et al.
Published: (2024)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Leveraging Speech for Gesture Detection in Multimodal Communication
by: Ghaleb, Esam, et al.
Published: (2024)
by: Ghaleb, Esam, et al.
Published: (2024)
Leveraging Spatial and Semantic Feature Extraction for Skin Cancer Diagnosis with Capsule Networks and Graph Neural Networks
by: Santoso, K. P., et al.
Published: (2024)
by: Santoso, K. P., et al.
Published: (2024)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
by: Shiri, Fatemeh, et al.
Published: (2024)
by: Shiri, Fatemeh, et al.
Published: (2024)
Target Prompting for Information Extraction with Vision Language Model
by: Medhi, Dipankar
Published: (2024)
by: Medhi, Dipankar
Published: (2024)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
by: Li, Hengzhuang, et al.
Published: (2025)
by: Li, Hengzhuang, et al.
Published: (2025)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
by: Lui, Nicholas, et al.
Published: (2023)
by: Lui, Nicholas, et al.
Published: (2023)
Similar Items
-
Extracting Accurate Materials Data from Research Papers with Conversational Language Models and Prompt Engineering
by: Polak, Maciej P., et al.
Published: (2023) -
Flexible, Model-Agnostic Method for Materials Data Extraction from Text Using General Purpose Language Models
by: Polak, Maciej P., et al.
Published: (2023) -
Plots Unlock Time-Series Understanding in Multimodal Models
by: Daswani, Mayank, et al.
Published: (2024) -
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
by: Sarmadi, Hamid, et al.
Published: (2025) -
Accelerating Domain-Aware Electron Microscopy Analysis Using Deep Learning Models with Synthetic Data and Image-Wide Confidence Scoring
by: Lynch, Matthew J., et al.
Published: (2024)