An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Giang Son, Lim, Zi Pong, Modi, Sarthak Ketanbhai, Teo, Yon Shin, Wang, Wenya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding
por: Bao, Jianzhu, et al.
Publicado: (2026)
por: Bao, Jianzhu, et al.
Publicado: (2026)
Towards Calibration Enhanced Network by Inverse Adversarial Attack
por: Cheng, Yupeng, et al.
Publicado: (2025)
por: Cheng, Yupeng, et al.
Publicado: (2025)
Towards Universal & Efficient Model Compression via Exponential Torque Pruning
por: Modi, Sarthak Ketanbhai, et al.
Publicado: (2025)
por: Modi, Sarthak Ketanbhai, et al.
Publicado: (2025)
SOSAE: Self-Organizing Sparse AutoEncoder
por: Modi, Sarthak Ketanbhai, et al.
Publicado: (2025)
por: Modi, Sarthak Ketanbhai, et al.
Publicado: (2025)
Towards Making Flowchart Images Machine Interpretable
por: Shukla, Shreya, et al.
Publicado: (2025)
por: Shukla, Shreya, et al.
Publicado: (2025)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
por: Jiang, Yifan, et al.
Publicado: (2026)
por: Jiang, Yifan, et al.
Publicado: (2026)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
por: Ye, Junyi, et al.
Publicado: (2024)
por: Ye, Junyi, et al.
Publicado: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
por: Wu, Yin, et al.
Publicado: (2025)
por: Wu, Yin, et al.
Publicado: (2025)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
por: Nguyen, Ngoc Son, et al.
Publicado: (2024)
por: Nguyen, Ngoc Son, et al.
Publicado: (2024)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
por: Omasa, Takamitsu, et al.
Publicado: (2025)
por: Omasa, Takamitsu, et al.
Publicado: (2025)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
por: Ahmadi, Saba, et al.
Publicado: (2023)
por: Ahmadi, Saba, et al.
Publicado: (2023)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
por: Singh, Shubhankar, et al.
Publicado: (2024)
por: Singh, Shubhankar, et al.
Publicado: (2024)
A Unified Agentic Framework for Evaluating Conditional Image Generation
por: Wang, Jifang, et al.
Publicado: (2025)
por: Wang, Jifang, et al.
Publicado: (2025)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
por: Lee, Yebin, et al.
Publicado: (2024)
por: Lee, Yebin, et al.
Publicado: (2024)
Orchestration-Free Customer Service Automation: A Privacy-Preserving and Flowchart-Guided Framework
por: Hong, Mengze, et al.
Publicado: (2026)
por: Hong, Mengze, et al.
Publicado: (2026)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
Phrase-Instance Alignment for Generalized Referring Segmentation
por: Nguyen, E-Ro, et al.
Publicado: (2024)
por: Nguyen, E-Ro, et al.
Publicado: (2024)
CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
por: Zheng, Mingzhe, et al.
Publicado: (2025)
por: Zheng, Mingzhe, et al.
Publicado: (2025)
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
por: Nguyen, Ngoc-Son, et al.
Publicado: (2025)
por: Nguyen, Ngoc-Son, et al.
Publicado: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
por: Teo, Rachel S. Y., et al.
Publicado: (2025)
por: Teo, Rachel S. Y., et al.
Publicado: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
por: Teo, Rachel S. Y., et al.
Publicado: (2024)
por: Teo, Rachel S. Y., et al.
Publicado: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
por: Teo, Rachel S. Y., et al.
Publicado: (2024)
por: Teo, Rachel S. Y., et al.
Publicado: (2024)
Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
por: Thai, Triet Minh, et al.
Publicado: (2023)
por: Thai, Triet Minh, et al.
Publicado: (2023)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
por: Vasilev, Viacheslav, et al.
Publicado: (2025)
por: Vasilev, Viacheslav, et al.
Publicado: (2025)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
Spatial Semantic Recurrent Mining for Referring Image Segmentation
por: Yang, Jiaxing, et al.
Publicado: (2024)
por: Yang, Jiaxing, et al.
Publicado: (2024)
TechING: Towards Real World Technical Image Understanding via VLMs
por: Nadeem, Tafazzul, et al.
Publicado: (2026)
por: Nadeem, Tafazzul, et al.
Publicado: (2026)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
por: Wu, Chengyue, et al.
Publicado: (2024)
por: Wu, Chengyue, et al.
Publicado: (2024)
Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
por: Yun, Sukmin, et al.
Publicado: (2024)
por: Yun, Sukmin, et al.
Publicado: (2024)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
por: Hu, Zhe, et al.
Publicado: (2025)
por: Hu, Zhe, et al.
Publicado: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
por: Yang, Yue, et al.
Publicado: (2025)
por: Yang, Yue, et al.
Publicado: (2025)
CAP: Evaluation of Persuasive and Creative Image Generation
por: Aghazadeh, Aysan, et al.
Publicado: (2024)
por: Aghazadeh, Aysan, et al.
Publicado: (2024)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
por: Shin, Philip Wootaek, et al.
Publicado: (2024)
por: Shin, Philip Wootaek, et al.
Publicado: (2024)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
por: Lim, Hyeonseok, et al.
Publicado: (2024)
por: Lim, Hyeonseok, et al.
Publicado: (2024)
Multimedia Generative Script Learning for Task Planning
por: Wang, Qingyun, et al.
Publicado: (2022)
por: Wang, Qingyun, et al.
Publicado: (2022)
Barriers in Integrating Medical Visual Question Answering into Radiology Workflows: A Scoping Review and Clinicians' Insights
por: Mishra, Deepali, et al.
Publicado: (2025)
por: Mishra, Deepali, et al.
Publicado: (2025)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
por: Bao, Zhijie, et al.
Publicado: (2026)
por: Bao, Zhijie, et al.
Publicado: (2026)
Ejemplares similares
-
Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding
por: Bao, Jianzhu, et al.
Publicado: (2026) -
Towards Calibration Enhanced Network by Inverse Adversarial Attack
por: Cheng, Yupeng, et al.
Publicado: (2025) -
Towards Universal & Efficient Model Compression via Exponential Torque Pruning
por: Modi, Sarthak Ketanbhai, et al.
Publicado: (2025) -
SOSAE: Self-Organizing Sparse AutoEncoder
por: Modi, Sarthak Ketanbhai, et al.
Publicado: (2025) -
Towards Making Flowchart Images Machine Interpretable
por: Shukla, Shreya, et al.
Publicado: (2025)