Hierarchical structure understanding in complex tables with VLLMs: a benchmark and experiments
Fuente:
arXiv
Saved in:
| Main Authors: | Bindini, Luca, Giovannini, Simone, Marinai, Simone, Nardoni, Valeria, Ali, Kimiya Noor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
by: Giovannini, Simone, et al.
Published: (2025)
by: Giovannini, Simone, et al.
Published: (2025)
Towards Reliable and Interpretable Document Question Answering via VLMs
by: Chen, Alessio, et al.
Published: (2025)
by: Chen, Alessio, et al.
Published: (2025)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
by: Min, Junghyun, et al.
Published: (2025)
by: Min, Junghyun, et al.
Published: (2025)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
by: Jassim, Serwan, et al.
Published: (2023)
by: Jassim, Serwan, et al.
Published: (2023)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
by: Corbo, Simone, et al.
Published: (2025)
by: Corbo, Simone, et al.
Published: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
LLMs in Interpreting Legal Documents
by: Corbo, Simone
Published: (2025)
by: Corbo, Simone
Published: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
by: Conti, Alessandro, et al.
Published: (2024)
by: Conti, Alessandro, et al.
Published: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Code-Switching and Syntax: A Large-Scale Experiment
by: Sterner, Igor, et al.
Published: (2025)
by: Sterner, Igor, et al.
Published: (2025)
Minimal Pair-Based Evaluation of Code-Switching
by: Sterner, Igor, et al.
Published: (2025)
by: Sterner, Igor, et al.
Published: (2025)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection
by: Tedeschini, Luca, et al.
Published: (2026)
by: Tedeschini, Luca, et al.
Published: (2026)
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
by: Ali, Wazir, et al.
Published: (2026)
by: Ali, Wazir, et al.
Published: (2026)
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
by: Ali, Muhammad, et al.
Published: (2025)
by: Ali, Muhammad, et al.
Published: (2025)
Towards Non-Latin Text and Layout Personalization for Enhanced Readability
by: Buoy, Rina, et al.
Published: (2026)
by: Buoy, Rina, et al.
Published: (2026)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
by: Bombari, Simone, et al.
Published: (2024)
by: Bombari, Simone, et al.
Published: (2024)
Inverse Language Modeling towards Robust and Grounded LLMs
by: Gabrielli, Davide, et al.
Published: (2025)
by: Gabrielli, Davide, et al.
Published: (2025)
Can Language Models Rival Mathematics Students? Evaluating Mathematical Reasoning through Textual Manipulation and Human Experiments
by: Nikolaiev, Andrii, et al.
Published: (2024)
by: Nikolaiev, Andrii, et al.
Published: (2024)
Line of Sight: On Linear Representations in VLLMs
by: Rajaram, Achyuta, et al.
Published: (2025)
by: Rajaram, Achyuta, et al.
Published: (2025)
An Expert-grounded benchmark of General Purpose LLMs in LCA
by: Donaldson, Artur, et al.
Published: (2025)
by: Donaldson, Artur, et al.
Published: (2025)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
by: Ailem, Melissa, et al.
Published: (2024)
by: Ailem, Melissa, et al.
Published: (2024)
Linguini: A benchmark for language-agnostic linguistic reasoning
by: Sánchez, Eduardo, et al.
Published: (2024)
by: Sánchez, Eduardo, et al.
Published: (2024)
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
by: Ali, Wazir, et al.
Published: (2026)
by: Ali, Wazir, et al.
Published: (2026)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
by: Perlitz, Yotam, et al.
Published: (2026)
by: Perlitz, Yotam, et al.
Published: (2026)
A geometric framework for interstellar discourse on fundamental physical structures
by: Esposito, Giampiero, et al.
Published: (2024)
by: Esposito, Giampiero, et al.
Published: (2024)
NaviQAte: Functionality-Guided Web Application Navigation
by: Shahbandeh, Mobina, et al.
Published: (2024)
by: Shahbandeh, Mobina, et al.
Published: (2024)
Using Shapley interactions to understand how models use structure
by: Singhvi, Divyansh, et al.
Published: (2024)
by: Singhvi, Divyansh, et al.
Published: (2024)
A multilingual hallucination benchmark: MultiWikiQHalluA
by: Thoresen, Freja, et al.
Published: (2026)
by: Thoresen, Freja, et al.
Published: (2026)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
by: Mousavi, Seyed Mahed, et al.
Published: (2025)
by: Mousavi, Seyed Mahed, et al.
Published: (2025)
Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs
by: Bombieri, Marco, et al.
Published: (2026)
by: Bombieri, Marco, et al.
Published: (2026)
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
by: Mousavi, Seyed Mahed, et al.
Published: (2026)
by: Mousavi, Seyed Mahed, et al.
Published: (2026)
ROUGE-K: Do Your Summaries Have Keywords?
by: Takeshita, Sotaro, et al.
Published: (2024)
by: Takeshita, Sotaro, et al.
Published: (2024)
Suvach -- Generated Hindi QA benchmark
by: Narayanan, Vaishak, et al.
Published: (2024)
by: Narayanan, Vaishak, et al.
Published: (2024)
What does it mean to understand language?
by: Casto, Colton, et al.
Published: (2025)
by: Casto, Colton, et al.
Published: (2025)
Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain
by: Royer, Corentin, et al.
Published: (2026)
by: Royer, Corentin, et al.
Published: (2026)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
by: Kankowski, Florian, et al.
Published: (2025)
by: Kankowski, Florian, et al.
Published: (2025)
Hierarchical Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
by: Morbiato, Filippo, et al.
Published: (2026)
by: Morbiato, Filippo, et al.
Published: (2026)
Attention Sinks in Diffusion Language Models
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
by: Rulli, Maximo Eduardo, et al.
Published: (2025)
CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
by: Zhang, Feng, et al.
Published: (2025)
by: Zhang, Feng, et al.
Published: (2025)
Similar Items
-
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
by: Giovannini, Simone, et al.
Published: (2025) -
Towards Reliable and Interpretable Document Question Answering via VLMs
by: Chen, Alessio, et al.
Published: (2025) -
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
by: Min, Junghyun, et al.
Published: (2025) -
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
by: Jassim, Serwan, et al.
Published: (2023) -
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
by: Corbo, Simone, et al.
Published: (2025)