METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Bingxuan, Wang, Yiwei, Gu, Jiuxiang, Chang, Kai-Wei, Peng, Nanyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
Contrastive Visual Data Augmentation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
von: Kondic, Jovana, et al.
Veröffentlicht: (2026)
von: Kondic, Jovana, et al.
Veröffentlicht: (2026)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
von: Zhao, Haozhe, et al.
Veröffentlicht: (2025)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2025)
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
von: Wang, Xingqi, et al.
Veröffentlicht: (2025)
von: Wang, Xingqi, et al.
Veröffentlicht: (2025)
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
von: Hong, Yining, et al.
Veröffentlicht: (2025)
von: Hong, Yining, et al.
Veröffentlicht: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2024)
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2024)
ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
von: Du, Yong, et al.
Veröffentlicht: (2025)
von: Du, Yong, et al.
Veröffentlicht: (2025)
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
On Pre-training of Multimodal Language Models Customized for Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
von: Kaur, Rachneet, et al.
Veröffentlicht: (2025)
von: Kaur, Rachneet, et al.
Veröffentlicht: (2025)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
von: Zhang, Junlei, et al.
Veröffentlicht: (2025)
von: Zhang, Junlei, et al.
Veröffentlicht: (2025)
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
von: Tao, Xingjian, et al.
Veröffentlicht: (2025)
von: Tao, Xingjian, et al.
Veröffentlicht: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
von: Zhao, Haozhe, et al.
Veröffentlicht: (2026)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2026)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024) -
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026) -
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024) -
Contrastive Visual Data Augmentation
von: Zhou, Yu, et al.
Veröffentlicht: (2025) -
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
von: Kondic, Jovana, et al.
Veröffentlicht: (2026)