Towards Making Flowchart Images Machine Interpretable
Fuente:
arXiv
Guardado en:
| Autores principales: | Shukla, Shreya, Gatti, Prajwal, Kumar, Yogesh, Yadav, Vikash, Mishra, Anand |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
por: Samsonau, Sergey V
Publicado: (2026)
por: Samsonau, Sergey V
Publicado: (2026)
Identity resolution of software metadata using Large Language Models
por: del Pico, Eva Martín, et al.
Publicado: (2025)
por: del Pico, Eva Martín, et al.
Publicado: (2025)
An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand
por: Isom, Joshua D.
Publicado: (2025)
por: Isom, Joshua D.
Publicado: (2025)
EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering
por: Dou, Zhifei, et al.
Publicado: (2026)
por: Dou, Zhifei, et al.
Publicado: (2026)
Docling Technical Report
por: Auer, Christoph, et al.
Publicado: (2024)
por: Auer, Christoph, et al.
Publicado: (2024)
Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation
por: Vaidya, Shreyas, et al.
Publicado: (2023)
por: Vaidya, Shreyas, et al.
Publicado: (2023)
WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
por: Liang, Shanchao, et al.
Publicado: (2024)
por: Liang, Shanchao, et al.
Publicado: (2024)
Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
por: Livathinos, Nikolaos, et al.
Publicado: (2025)
por: Livathinos, Nikolaos, et al.
Publicado: (2025)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
por: Yang, Cheng, et al.
Publicado: (2024)
por: Yang, Cheng, et al.
Publicado: (2024)
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
por: Xu, Zhenran, et al.
Publicado: (2025)
por: Xu, Zhenran, et al.
Publicado: (2025)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
por: Lu, Zimu, et al.
Publicado: (2026)
por: Lu, Zimu, et al.
Publicado: (2026)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
por: Li, Kaixin, et al.
Publicado: (2024)
por: Li, Kaixin, et al.
Publicado: (2024)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
por: Duan, Changxu
Publicado: (2025)
por: Duan, Changxu
Publicado: (2025)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
por: Aich, Abhishek, et al.
Publicado: (2026)
por: Aich, Abhishek, et al.
Publicado: (2026)
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
por: Sun, Qiushi, et al.
Publicado: (2025)
por: Sun, Qiushi, et al.
Publicado: (2025)
Automating Categorization of Scientific Texts with In-Context Learning and Prompt-Chaining in Large Language Models
por: Shahi, Gautam Kishore, et al.
Publicado: (2026)
por: Shahi, Gautam Kishore, et al.
Publicado: (2026)
Automatic Modeling of Social Concepts Evoked by Art Images as Multimodal Frames
por: Pandiani, Delfina Sol Martinez, et al.
Publicado: (2021)
por: Pandiani, Delfina Sol Martinez, et al.
Publicado: (2021)
MVOS_HSI: A Python Library for Preprocessing Agricultural Crop Hyperspectral Data
por: Aggarwal, Rishik, et al.
Publicado: (2026)
por: Aggarwal, Rishik, et al.
Publicado: (2026)
Automatic Recognition of Learning Resource Category in a Digital Library
por: Banerjee, Soumya, et al.
Publicado: (2023)
por: Banerjee, Soumya, et al.
Publicado: (2023)
Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
por: Wu, Yu, et al.
Publicado: (2025)
por: Wu, Yu, et al.
Publicado: (2025)
PubMed-OCR: PMC Open Access OCR Annotations
por: Heidenreich, Hunter, et al.
Publicado: (2026)
por: Heidenreich, Hunter, et al.
Publicado: (2026)
TRIDIS: A Comprehensive Medieval and Early Modern Corpus for HTR and NER
por: Aguilar, Sergio Torres
Publicado: (2025)
por: Aguilar, Sergio Torres
Publicado: (2025)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
por: Karamolegkou, Antonia, et al.
Publicado: (2026)
por: Karamolegkou, Antonia, et al.
Publicado: (2026)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
por: Zhu, Minjun, et al.
Publicado: (2026)
por: Zhu, Minjun, et al.
Publicado: (2026)
Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents
por: Humphries, Mark, et al.
Publicado: (2024)
por: Humphries, Mark, et al.
Publicado: (2024)
Callico: a Versatile Open-Source Document Image Annotation Platform
por: Kermorvant, Christopher, et al.
Publicado: (2024)
por: Kermorvant, Christopher, et al.
Publicado: (2024)
Towards a Metadata Schema for Energy Research Software
por: Ferenz, Stephan, et al.
Publicado: (2026)
por: Ferenz, Stephan, et al.
Publicado: (2026)
SWARM-SLR -- Streamlined Workflow Automation for Machine-actionable Systematic Literature Reviews
por: Wittenborg, Tim, et al.
Publicado: (2024)
por: Wittenborg, Tim, et al.
Publicado: (2024)
Interpretable Gallbladder Ultrasound Diagnosis: A Lightweight Web-Mobile Software Platform with Real-Time XAI
por: Bhoyan, Fuyad Hasan, et al.
Publicado: (2025)
por: Bhoyan, Fuyad Hasan, et al.
Publicado: (2025)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
por: Wang, Wenxuan, et al.
Publicado: (2024)
por: Wang, Wenxuan, et al.
Publicado: (2024)
Running a Data Integration Lab in the Context of the EHRI Project: Challenges, Lessons Learnt and Future Directions
por: García-González, Herminio, et al.
Publicado: (2025)
por: García-González, Herminio, et al.
Publicado: (2025)
VideoGameBunny: Towards vision assistants for video games
por: Taesiri, Mohammad Reza, et al.
Publicado: (2024)
por: Taesiri, Mohammad Reza, et al.
Publicado: (2024)
Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
por: Shi, Xiang, et al.
Publicado: (2024)
por: Shi, Xiang, et al.
Publicado: (2024)
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
por: Wu, Xueqing, et al.
Publicado: (2025)
por: Wu, Xueqing, et al.
Publicado: (2025)
Benchmarking Image Perturbations for Testing Automated Driving Assistance Systems
por: Lambertenghi, Stefano Carlo, et al.
Publicado: (2025)
por: Lambertenghi, Stefano Carlo, et al.
Publicado: (2025)
ITKIT: Feasible CT Image Analysis based on SimpleITK and MMEngine
por: Zhang, Yiqin, et al.
Publicado: (2026)
por: Zhang, Yiqin, et al.
Publicado: (2026)
Cross-Breed Pig Identification Using Auricular Vein Pattern Recognition: A Machine Learning Approach for Small-Scale Farming Applications
por: Nsengiyumvaa, Emmanuel, et al.
Publicado: (2025)
por: Nsengiyumvaa, Emmanuel, et al.
Publicado: (2025)
SWAN -- Enabling Fast and Mobile Histopathology Image Annotation through Swipeable Interfaces
por: Banerjee, Sweta, et al.
Publicado: (2025)
por: Banerjee, Sweta, et al.
Publicado: (2025)
Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?
por: Lyu, Yunbo, et al.
Publicado: (2025)
por: Lyu, Yunbo, et al.
Publicado: (2025)
AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding
por: Khan, Safwat Ali, et al.
Publicado: (2024)
por: Khan, Safwat Ali, et al.
Publicado: (2024)
Ejemplares similares
-
sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
por: Samsonau, Sergey V
Publicado: (2026) -
Identity resolution of software metadata using Large Language Models
por: del Pico, Eva Martín, et al.
Publicado: (2025) -
An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand
por: Isom, Joshua D.
Publicado: (2025) -
EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering
por: Dou, Zhifei, et al.
Publicado: (2026) -
Docling Technical Report
por: Auer, Christoph, et al.
Publicado: (2024)