Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Deng, Naihao, Sun, Zhenjie, He, Ruiqi, Sikka, Aman, Chen, Yulong, Ma, Lin, Zhang, Yue, Mihalcea, Rada |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Table Instruction Tuning
di: Deng, Naihao, et al.
Pubblicazione: (2025)
di: Deng, Naihao, et al.
Pubblicazione: (2025)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
di: Deng, Naihao, et al.
Pubblicazione: (2025)
di: Deng, Naihao, et al.
Pubblicazione: (2025)
Table as Thought: Exploring Structured Thoughts in LLM Reasoning
di: Sun, Zhenjie, et al.
Pubblicazione: (2025)
di: Sun, Zhenjie, et al.
Pubblicazione: (2025)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
di: Arif, Samee, et al.
Pubblicazione: (2026)
di: Arif, Samee, et al.
Pubblicazione: (2026)
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
di: He, Ruiqi, et al.
Pubblicazione: (2024)
di: He, Ruiqi, et al.
Pubblicazione: (2024)
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
di: Deng, Naihao, et al.
Pubblicazione: (2025)
di: Deng, Naihao, et al.
Pubblicazione: (2025)
$R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks
di: Xia, Hanchen, et al.
Pubblicazione: (2024)
di: Xia, Hanchen, et al.
Pubblicazione: (2024)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
di: Bai, Longju, et al.
Pubblicazione: (2024)
di: Bai, Longju, et al.
Pubblicazione: (2024)
Benchmarking and Improving LLM Robustness for Personalized Generation
di: Okite, Chimaobi, et al.
Pubblicazione: (2025)
di: Okite, Chimaobi, et al.
Pubblicazione: (2025)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
di: He, Wei, et al.
Pubblicazione: (2024)
di: He, Wei, et al.
Pubblicazione: (2024)
The Curious Case of Curiosity across Human Cultures and LLMs
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
di: Ignat, Oana, et al.
Pubblicazione: (2024)
di: Ignat, Oana, et al.
Pubblicazione: (2024)
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
di: Ignat, Oana, et al.
Pubblicazione: (2023)
di: Ignat, Oana, et al.
Pubblicazione: (2023)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
di: Castro, Santiago, et al.
Pubblicazione: (2024)
di: Castro, Santiago, et al.
Pubblicazione: (2024)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
di: Borah, Angana, et al.
Pubblicazione: (2024)
di: Borah, Angana, et al.
Pubblicazione: (2024)
Towards Region-aware Bias Evaluation Metrics
di: Borah, Angana, et al.
Pubblicazione: (2024)
di: Borah, Angana, et al.
Pubblicazione: (2024)
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
di: Torres-Fonseca, Josue, et al.
Pubblicazione: (2026)
di: Torres-Fonseca, Josue, et al.
Pubblicazione: (2026)
Mind the (Belief) Gap: Group Identity in the World of LLMs
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
di: Shankarampeta, Abhilash, et al.
Pubblicazione: (2025)
di: Shankarampeta, Abhilash, et al.
Pubblicazione: (2025)
Effective Distillation of Table-based Reasoning Ability from LLMs
di: Yang, Bohao, et al.
Pubblicazione: (2023)
di: Yang, Bohao, et al.
Pubblicazione: (2023)
Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba
di: He, Ruiqi, et al.
Pubblicazione: (2024)
di: He, Ruiqi, et al.
Pubblicazione: (2024)
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025)
di: Fan, Yue, et al.
Pubblicazione: (2025)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
di: Guo, Zichun, et al.
Pubblicazione: (2026)
di: Guo, Zichun, et al.
Pubblicazione: (2026)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
di: Hong, Pengfei, et al.
Pubblicazione: (2024)
di: Hong, Pengfei, et al.
Pubblicazione: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
di: Patel, Maitreya, et al.
Pubblicazione: (2023)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
di: Chen, Renmiao, et al.
Pubblicazione: (2026)
di: Chen, Renmiao, et al.
Pubblicazione: (2026)
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
di: Xu, Zitong, et al.
Pubblicazione: (2026)
di: Xu, Zitong, et al.
Pubblicazione: (2026)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
di: Nwatu, Joan, et al.
Pubblicazione: (2025)
di: Nwatu, Joan, et al.
Pubblicazione: (2025)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2025)
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2025)
Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
di: Lompo, Boammani Aser, et al.
Pubblicazione: (2025)
di: Lompo, Boammani Aser, et al.
Pubblicazione: (2025)
Evaluating Numerical Reasoning in Text-to-Image Models
di: Kajić, Ivana, et al.
Pubblicazione: (2024)
di: Kajić, Ivana, et al.
Pubblicazione: (2024)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
di: Arif, Samee, et al.
Pubblicazione: (2026)
di: Arif, Samee, et al.
Pubblicazione: (2026)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
di: Jiang, Houcheng, et al.
Pubblicazione: (2026)
di: Jiang, Houcheng, et al.
Pubblicazione: (2026)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
di: Dong, Qihua, et al.
Pubblicazione: (2026)
di: Dong, Qihua, et al.
Pubblicazione: (2026)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
di: Chen, Yangyi, et al.
Pubblicazione: (2023)
di: Chen, Yangyi, et al.
Pubblicazione: (2023)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
di: Ji, Yikun, et al.
Pubblicazione: (2025)
di: Ji, Yikun, et al.
Pubblicazione: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
di: Fang, I-Sheng, et al.
Pubblicazione: (2025)
di: Fang, I-Sheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rethinking Table Instruction Tuning
di: Deng, Naihao, et al.
Pubblicazione: (2025) -
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
di: Deng, Naihao, et al.
Pubblicazione: (2025) -
Table as Thought: Exploring Structured Thoughts in LLM Reasoning
di: Sun, Zhenjie, et al.
Pubblicazione: (2025) -
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
di: Arif, Samee, et al.
Pubblicazione: (2026) -
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
di: He, Ruiqi, et al.
Pubblicazione: (2024)