START: Spatial and Textual Learning for Chart Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Zhuoming, Gao, Xiaofeng, Niu, Feiyang, Gao, Qiaozi, Liu, Liu, Piramuthu, Robinson |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AskChart: Universal Chart Understanding through Textual Enhancement
di: Yang, Xudong, et al.
Pubblicazione: (2024)
di: Yang, Xudong, et al.
Pubblicazione: (2024)
ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
di: Fang, Zhenghan, et al.
Pubblicazione: (2025)
di: Fang, Zhenghan, et al.
Pubblicazione: (2025)
T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design
di: Li, Jiachen, et al.
Pubblicazione: (2024)
di: Li, Jiachen, et al.
Pubblicazione: (2024)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
di: Xie, Tianchi, et al.
Pubblicazione: (2025)
di: Xie, Tianchi, et al.
Pubblicazione: (2025)
On Pre-training of Multimodal Language Models Customized for Chart Understanding
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2024)
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2024)
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
di: Zhang, Zhiwei, et al.
Pubblicazione: (2023)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2023)
Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
di: Li, Zhuoming, et al.
Pubblicazione: (2025)
di: Li, Zhuoming, et al.
Pubblicazione: (2025)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
di: Liu, Yuhong, et al.
Pubblicazione: (2025)
di: Liu, Yuhong, et al.
Pubblicazione: (2025)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
di: Zhong, Yiwu, et al.
Pubblicazione: (2024)
di: Zhong, Yiwu, et al.
Pubblicazione: (2024)
Enhancing Spatial Reasoning through Visual and Textual Thinking
di: Liang, Xun, et al.
Pubblicazione: (2025)
di: Liang, Xun, et al.
Pubblicazione: (2025)
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
di: Kondic, Jovana, et al.
Pubblicazione: (2026)
di: Kondic, Jovana, et al.
Pubblicazione: (2026)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
di: Liu, Zichuan, et al.
Pubblicazione: (2025)
di: Liu, Zichuan, et al.
Pubblicazione: (2025)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
di: Wang, Haoming, et al.
Pubblicazione: (2026)
di: Wang, Haoming, et al.
Pubblicazione: (2026)
ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2024)
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2024)
Benchmarking Multimodal RAG through a Chart-based Document Question-Answering Generation Framework
di: Yang, Yuming, et al.
Pubblicazione: (2025)
di: Yang, Yuming, et al.
Pubblicazione: (2025)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2025)
di: Fan, Wan-Cyuan, et al.
Pubblicazione: (2025)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
di: Kondic, Jovana, et al.
Pubblicazione: (2025)
di: Kondic, Jovana, et al.
Pubblicazione: (2025)
ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
di: Xiao, Shishi, et al.
Pubblicazione: (2026)
di: Xiao, Shishi, et al.
Pubblicazione: (2026)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
di: Li, Yuan, et al.
Pubblicazione: (2026)
di: Li, Yuan, et al.
Pubblicazione: (2026)
Saliency Guided Longitudinal Medical Visual Question Answering
di: Wu, Jialin, et al.
Pubblicazione: (2025)
di: Wu, Jialin, et al.
Pubblicazione: (2025)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
di: Lee, Yujian, et al.
Pubblicazione: (2026)
di: Lee, Yujian, et al.
Pubblicazione: (2026)
Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing
di: Zhang, Weiyu, et al.
Pubblicazione: (2026)
di: Zhang, Weiyu, et al.
Pubblicazione: (2026)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
di: Kong, Fei, et al.
Pubblicazione: (2025)
di: Kong, Fei, et al.
Pubblicazione: (2025)
TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations
di: Gao, Mingze, et al.
Pubblicazione: (2024)
di: Gao, Mingze, et al.
Pubblicazione: (2024)
HumanCM: One Step Human Motion Prediction
di: Haojie, Liu, et al.
Pubblicazione: (2025)
di: Haojie, Liu, et al.
Pubblicazione: (2025)
FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image Translation
di: Gao, Xiang, et al.
Pubblicazione: (2024)
di: Gao, Xiang, et al.
Pubblicazione: (2024)
VaPR -- Vision-language Preference alignment for Reasoning
di: Wadhawan, Rohan, et al.
Pubblicazione: (2025)
di: Wadhawan, Rohan, et al.
Pubblicazione: (2025)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
di: Zhang, Haoyu, et al.
Pubblicazione: (2025)
di: Zhang, Haoyu, et al.
Pubblicazione: (2025)
Figuring out Figures: Using Textual References to Caption Scientific Figures
di: Cao, Stanley, et al.
Pubblicazione: (2024)
di: Cao, Stanley, et al.
Pubblicazione: (2024)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
di: Ropero, Fernando, et al.
Pubblicazione: (2026)
di: Ropero, Fernando, et al.
Pubblicazione: (2026)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
di: Zhang, Chenkai, et al.
Pubblicazione: (2025)
di: Zhang, Chenkai, et al.
Pubblicazione: (2025)
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
di: Wang, Xingqi, et al.
Pubblicazione: (2025)
di: Wang, Xingqi, et al.
Pubblicazione: (2025)
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
di: Xue, Qiyao, et al.
Pubblicazione: (2025)
di: Xue, Qiyao, et al.
Pubblicazione: (2025)
ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single Image
di: Gao, Yunqi, et al.
Pubblicazione: (2025)
di: Gao, Yunqi, et al.
Pubblicazione: (2025)
Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions
di: Yang, Donglu, et al.
Pubblicazione: (2025)
di: Yang, Donglu, et al.
Pubblicazione: (2025)
GAFR-Net: A Graph Attention and Fuzzy-Rule Network for Interpretable Breast Cancer Image Classification
di: Gao, Lin-Guo, et al.
Pubblicazione: (2026)
di: Gao, Lin-Guo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AskChart: Universal Chart Understanding through Textual Enhancement
di: Yang, Xudong, et al.
Pubblicazione: (2024) -
ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
di: Fang, Zhenghan, et al.
Pubblicazione: (2025) -
T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design
di: Li, Jiachen, et al.
Pubblicazione: (2024) -
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
di: Zhang, Yichi, et al.
Pubblicazione: (2024) -
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
di: Guo, Zhongbin, et al.
Pubblicazione: (2026)