AltChart: Enhancing VLM-based Chart Summarization Through Multi-Pretext Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Moured, Omar, Zhang, Jiaming, Sarfraz, M. Saquib, Stiefelhagen, Rainer |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alt4Blind: A User Interface to Simplify Charts Alt-Text Creation
by: Moured, Omar, et al.
Published: (2024)
by: Moured, Omar, et al.
Published: (2024)
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
by: Vogel, Alexander, et al.
Published: (2025)
by: Vogel, Alexander, et al.
Published: (2025)
Chart4Blind: An Intelligent Interface for Chart Accessibility Conversion
by: Moured, Omar, et al.
Published: (2024)
by: Moured, Omar, et al.
Published: (2024)
CHAOS: Chart Analysis with Outlier Samples
by: Moured, Omar, et al.
Published: (2025)
by: Moured, Omar, et al.
Published: (2025)
ChartFormer: A Large Vision Language Model for Converting Chart Images into Tactile Accessible SVGs
by: Moured, Omar, et al.
Published: (2024)
by: Moured, Omar, et al.
Published: (2024)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
by: Kondic, Jovana, et al.
Published: (2025)
by: Kondic, Jovana, et al.
Published: (2025)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
by: Yang, Kaichun, et al.
Published: (2025)
by: Yang, Kaichun, et al.
Published: (2025)
Generalization of CNNs on Relational Reasoning with Bar Charts
by: Cui, Zhenxing, et al.
Published: (2025)
by: Cui, Zhenxing, et al.
Published: (2025)
Strike the Balance: On-the-Fly Uncertainty based User Interactions for Long-Term Video Object Segmentation
by: Vujasinović, Stéphane, et al.
Published: (2024)
by: Vujasinović, Stéphane, et al.
Published: (2024)
Rethinking Annotator Simulation: Realistic Evaluation of Whole-Body PET Lesion Interactive Segmentation Methods
by: Marinov, Zdravko, et al.
Published: (2024)
by: Marinov, Zdravko, et al.
Published: (2024)
SFDLA: Source-Free Document Layout Analysis
by: Tewes, Sebastian, et al.
Published: (2025)
by: Tewes, Sebastian, et al.
Published: (2025)
Spacewalker: Traversing Representation Spaces for Fast Interactive Exploration and Annotation of Unstructured Data
by: Heine, Lukas, et al.
Published: (2024)
by: Heine, Lukas, et al.
Published: (2024)
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends
by: Al-Shetairy, Mirna, et al.
Published: (2024)
by: Al-Shetairy, Mirna, et al.
Published: (2024)
Enhancing Saliency Prediction in Monitoring Tasks: The Role of Visual Highlights
by: Wu, Zekun, et al.
Published: (2024)
by: Wu, Zekun, et al.
Published: (2024)
VLM-driven Behavior Tree for Context-aware Task Planning
by: Wake, Naoki, et al.
Published: (2025)
by: Wake, Naoki, et al.
Published: (2025)
AltCanvas: A Tile-Based Image Editor with Generative AI for Blind or Visually Impaired People
by: Lee, Seonghee, et al.
Published: (2024)
by: Lee, Seonghee, et al.
Published: (2024)
ChartOptimiser: Task-driven Optimisation of Chart Designs
by: Wang, Yao, et al.
Published: (2025)
by: Wang, Yao, et al.
Published: (2025)
Extend Your Horizon: A Device-Agnostic Surgical Tool Tracking Framework with Multi-View Optimization for Augmented Reality
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness
by: Mukhopadhyay, Srija, et al.
Published: (2024)
by: Mukhopadhyay, Srija, et al.
Published: (2024)
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
by: Kawamura, Kazuki, et al.
Published: (2024)
by: Kawamura, Kazuki, et al.
Published: (2024)
WebAccessVL: Violation-Aware VLM for Web Accessibility
by: Zheng, Amber Yijia, et al.
Published: (2025)
by: Zheng, Amber Yijia, et al.
Published: (2025)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2025)
by: Garg, Mallika, et al.
Published: (2025)
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
by: Li, Xun, et al.
Published: (2025)
by: Li, Xun, et al.
Published: (2025)
Vid2Coach: Transforming How-To Videos into Task Assistants
by: Huh, Mina, et al.
Published: (2025)
by: Huh, Mina, et al.
Published: (2025)
VideoMix: Aggregating How-To Videos for Task-Oriented Learning
by: Yang, Saelyne, et al.
Published: (2025)
by: Yang, Saelyne, et al.
Published: (2025)
PrivatEyes: Appearance-based Gaze Estimation Using Federated Secure Multi-Party Computation
by: Elfares, Mayar, et al.
Published: (2024)
by: Elfares, Mayar, et al.
Published: (2024)
A Monocular SLAM-based Multi-User Positioning System with Image Occlusion in Augmented Reality
by: Lien, Wei-Hsiang, et al.
Published: (2024)
by: Lien, Wei-Hsiang, et al.
Published: (2024)
Reviving Static Charts into Live Charts
by: Ying, Lu, et al.
Published: (2023)
by: Ying, Lu, et al.
Published: (2023)
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
Few-Shot VLM-Based G-Code and HMI Verification in CNC Machining
by: Pour, Yasaman Hashem, et al.
Published: (2025)
by: Pour, Yasaman Hashem, et al.
Published: (2025)
Unsupervised Domain Adaptation for RF-based Gesture Recognition
by: Zhang, Bin-Bin, et al.
Published: (2021)
by: Zhang, Bin-Bin, et al.
Published: (2021)
Category-aware EEG image generation based on wavelet transform and contrast semantic loss
by: Zhang, Enshang, et al.
Published: (2025)
by: Zhang, Enshang, et al.
Published: (2025)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
by: Liu, Can, et al.
Published: (2025)
by: Liu, Can, et al.
Published: (2025)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
by: Xenos, Alexandros, et al.
Published: (2024)
by: Xenos, Alexandros, et al.
Published: (2024)
Enhanced Automated Quality Assessment Network for Interactive Building Segmentation in High-Resolution Remote Sensing Imagery
by: Zhang, Zhili, et al.
Published: (2024)
by: Zhang, Zhili, et al.
Published: (2024)
egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks
by: Jammot, Matthias, et al.
Published: (2025)
by: Jammot, Matthias, et al.
Published: (2025)
ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale
by: Huang, Jinbin, et al.
Published: (2024)
by: Huang, Jinbin, et al.
Published: (2024)
VisionCAD: An Integration-Free Radiology Copilot Framework
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Dreamcrafter: Immersive Editing of 3D Radiance Fields Through Flexible, Generative Inputs and Outputs
by: Vachha, Cyrus, et al.
Published: (2025)
by: Vachha, Cyrus, et al.
Published: (2025)
Similar Items
-
Alt4Blind: A User Interface to Simplify Charts Alt-Text Creation
by: Moured, Omar, et al.
Published: (2024) -
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
by: Vogel, Alexander, et al.
Published: (2025) -
Chart4Blind: An Intelligent Interface for Chart Accessibility Conversion
by: Moured, Omar, et al.
Published: (2024) -
CHAOS: Chart Analysis with Outlier Samples
by: Moured, Omar, et al.
Published: (2025) -
ChartFormer: A Large Vision Language Model for Converting Chart Images into Tactile Accessible SVGs
by: Moured, Omar, et al.
Published: (2024)