Hierarchical structure understanding in complex tables with VLLMs: a benchmark and experiments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bindini, Luca, Giovannini, Simone, Marinai, Simone, Nardoni, Valeria, Ali, Kimiya Noor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
von: Giovannini, Simone, et al.
Veröffentlicht: (2025)
von: Giovannini, Simone, et al.
Veröffentlicht: (2025)
Towards Reliable and Interpretable Document Question Answering via VLMs
von: Chen, Alessio, et al.
Veröffentlicht: (2025)
von: Chen, Alessio, et al.
Veröffentlicht: (2025)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
von: Min, Junghyun, et al.
Veröffentlicht: (2025)
von: Min, Junghyun, et al.
Veröffentlicht: (2025)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
von: Corbo, Simone, et al.
Veröffentlicht: (2025)
von: Corbo, Simone, et al.
Veröffentlicht: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
LLMs in Interpreting Legal Documents
von: Corbo, Simone
Veröffentlicht: (2025)
von: Corbo, Simone
Veröffentlicht: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
Code-Switching and Syntax: A Large-Scale Experiment
von: Sterner, Igor, et al.
Veröffentlicht: (2025)
von: Sterner, Igor, et al.
Veröffentlicht: (2025)
Minimal Pair-Based Evaluation of Code-Switching
von: Sterner, Igor, et al.
Veröffentlicht: (2025)
von: Sterner, Igor, et al.
Veröffentlicht: (2025)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection
von: Tedeschini, Luca, et al.
Veröffentlicht: (2026)
von: Tedeschini, Luca, et al.
Veröffentlicht: (2026)
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
von: Ali, Wazir, et al.
Veröffentlicht: (2026)
von: Ali, Wazir, et al.
Veröffentlicht: (2026)
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
von: Ali, Muhammad, et al.
Veröffentlicht: (2025)
von: Ali, Muhammad, et al.
Veröffentlicht: (2025)
Towards Non-Latin Text and Layout Personalization for Enhanced Readability
von: Buoy, Rina, et al.
Veröffentlicht: (2026)
von: Buoy, Rina, et al.
Veröffentlicht: (2026)
Towards Understanding the Word Sensitivity of Attention Layers: A Study via Random Features
von: Bombari, Simone, et al.
Veröffentlicht: (2024)
von: Bombari, Simone, et al.
Veröffentlicht: (2024)
Inverse Language Modeling towards Robust and Grounded LLMs
von: Gabrielli, Davide, et al.
Veröffentlicht: (2025)
von: Gabrielli, Davide, et al.
Veröffentlicht: (2025)
Can Language Models Rival Mathematics Students? Evaluating Mathematical Reasoning through Textual Manipulation and Human Experiments
von: Nikolaiev, Andrii, et al.
Veröffentlicht: (2024)
von: Nikolaiev, Andrii, et al.
Veröffentlicht: (2024)
Line of Sight: On Linear Representations in VLLMs
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2025)
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2025)
An Expert-grounded benchmark of General Purpose LLMs in LCA
von: Donaldson, Artur, et al.
Veröffentlicht: (2025)
von: Donaldson, Artur, et al.
Veröffentlicht: (2025)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
von: Ailem, Melissa, et al.
Veröffentlicht: (2024)
von: Ailem, Melissa, et al.
Veröffentlicht: (2024)
Linguini: A benchmark for language-agnostic linguistic reasoning
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
von: Ali, Wazir, et al.
Veröffentlicht: (2026)
von: Ali, Wazir, et al.
Veröffentlicht: (2026)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
von: Perlitz, Yotam, et al.
Veröffentlicht: (2026)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2026)
A geometric framework for interstellar discourse on fundamental physical structures
von: Esposito, Giampiero, et al.
Veröffentlicht: (2024)
von: Esposito, Giampiero, et al.
Veröffentlicht: (2024)
NaviQAte: Functionality-Guided Web Application Navigation
von: Shahbandeh, Mobina, et al.
Veröffentlicht: (2024)
von: Shahbandeh, Mobina, et al.
Veröffentlicht: (2024)
Using Shapley interactions to understand how models use structure
von: Singhvi, Divyansh, et al.
Veröffentlicht: (2024)
von: Singhvi, Divyansh, et al.
Veröffentlicht: (2024)
A multilingual hallucination benchmark: MultiWikiQHalluA
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
von: Thoresen, Freja, et al.
Veröffentlicht: (2026)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2025)
Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs
von: Bombieri, Marco, et al.
Veröffentlicht: (2026)
von: Bombieri, Marco, et al.
Veröffentlicht: (2026)
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2026)
ROUGE-K: Do Your Summaries Have Keywords?
von: Takeshita, Sotaro, et al.
Veröffentlicht: (2024)
von: Takeshita, Sotaro, et al.
Veröffentlicht: (2024)
Suvach -- Generated Hindi QA benchmark
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
von: Narayanan, Vaishak, et al.
Veröffentlicht: (2024)
What does it mean to understand language?
von: Casto, Colton, et al.
Veröffentlicht: (2025)
von: Casto, Colton, et al.
Veröffentlicht: (2025)
Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain
von: Royer, Corentin, et al.
Veröffentlicht: (2026)
von: Royer, Corentin, et al.
Veröffentlicht: (2026)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
von: Kankowski, Florian, et al.
Veröffentlicht: (2025)
von: Kankowski, Florian, et al.
Veröffentlicht: (2025)
Hierarchical Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
von: Morbiato, Filippo, et al.
Veröffentlicht: (2026)
von: Morbiato, Filippo, et al.
Veröffentlicht: (2026)
Attention Sinks in Diffusion Language Models
von: Rulli, Maximo Eduardo, et al.
Veröffentlicht: (2025)
von: Rulli, Maximo Eduardo, et al.
Veröffentlicht: (2025)
CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
von: Zhang, Feng, et al.
Veröffentlicht: (2025)
von: Zhang, Feng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
von: Giovannini, Simone, et al.
Veröffentlicht: (2025) -
Towards Reliable and Interpretable Document Question Answering via VLMs
von: Chen, Alessio, et al.
Veröffentlicht: (2025) -
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
von: Min, Junghyun, et al.
Veröffentlicht: (2025) -
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
von: Jassim, Serwan, et al.
Veröffentlicht: (2023) -
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
von: Corbo, Simone, et al.
Veröffentlicht: (2025)