TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Alonso, Iñigo, Miranda, Imanol, Agirre, Eneko, Lapata, Mirella |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PixT3: Pixel-based Table-To-Text Generation
por: Alonso, Iñigo, et al.
Publicado: (2023)
por: Alonso, Iñigo, et al.
Publicado: (2023)
Adding simple structure at inference improves Vision-Language Compositionality
por: Miranda, Imanol, et al.
Publicado: (2025)
por: Miranda, Imanol, et al.
Publicado: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
por: Miranda, Imanol, et al.
Publicado: (2024)
por: Miranda, Imanol, et al.
Publicado: (2024)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
por: Miranda, Imanol, et al.
Publicado: (2026)
por: Miranda, Imanol, et al.
Publicado: (2026)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
por: Gupta, Akash, et al.
Publicado: (2025)
por: Gupta, Akash, et al.
Publicado: (2025)
Parameter-free Video Segmentation for Vision and Language Understanding
por: Mahon, Louis, et al.
Publicado: (2025)
por: Mahon, Louis, et al.
Publicado: (2025)
Automatic Logical Forms improve fidelity in Table-to-Text generation
por: Alonso, Iñigo, et al.
Publicado: (2023)
por: Alonso, Iñigo, et al.
Publicado: (2023)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
por: Liu, Dongqi, et al.
Publicado: (2025)
por: Liu, Dongqi, et al.
Publicado: (2025)
Finding the Right Moment: Human-Assisted Trailer Creation via Task Composition
por: Papalampidi, Pinelopi, et al.
Publicado: (2021)
por: Papalampidi, Pinelopi, et al.
Publicado: (2021)
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
por: Mahon, Louis, et al.
Publicado: (2024)
por: Mahon, Louis, et al.
Publicado: (2024)
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
por: Kondic, Jovana, et al.
Publicado: (2026)
por: Kondic, Jovana, et al.
Publicado: (2026)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
por: Tanaka, Ryota, et al.
Publicado: (2024)
por: Tanaka, Ryota, et al.
Publicado: (2024)
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
por: Li, Lei, et al.
Publicado: (2024)
por: Li, Lei, et al.
Publicado: (2024)
RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events
por: Chen, Zhenyuan, et al.
Publicado: (2025)
por: Chen, Zhenyuan, et al.
Publicado: (2025)
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
por: Butsanets, Léo, et al.
Publicado: (2025)
por: Butsanets, Léo, et al.
Publicado: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
por: Zhu, Yingjie, et al.
Publicado: (2024)
por: Zhu, Yingjie, et al.
Publicado: (2024)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
por: Dong, Xuanzhao, et al.
Publicado: (2026)
por: Dong, Xuanzhao, et al.
Publicado: (2026)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
por: Sugiura, Issa, et al.
Publicado: (2025)
por: Sugiura, Issa, et al.
Publicado: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
por: Chung, Jiwan, et al.
Publicado: (2024)
por: Chung, Jiwan, et al.
Publicado: (2024)
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
por: Li, Zhen, et al.
Publicado: (2025)
por: Li, Zhen, et al.
Publicado: (2025)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
por: Al-Homoud, Haneen, et al.
Publicado: (2025)
por: Al-Homoud, Haneen, et al.
Publicado: (2025)
Understanding Bias in Large-Scale Visual Datasets
por: Zeng, Boya, et al.
Publicado: (2024)
por: Zeng, Boya, et al.
Publicado: (2024)
ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
por: Abdallah, Abdelrahman, et al.
Publicado: (2024)
por: Abdallah, Abdelrahman, et al.
Publicado: (2024)
Hierarchical Windowed Graph Attention Network and a Large Scale Dataset for Isolated Indian Sign Language Recognition
por: Patra, Suvajit, et al.
Publicado: (2024)
por: Patra, Suvajit, et al.
Publicado: (2024)
VISaGE: Understanding Visual Generics and Exceptions
por: Frank, Stella, et al.
Publicado: (2025)
por: Frank, Stella, et al.
Publicado: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
por: Bai, Tianyi, et al.
Publicado: (2025)
por: Bai, Tianyi, et al.
Publicado: (2025)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
por: Wang, Yuxuan, et al.
Publicado: (2024)
por: Wang, Yuxuan, et al.
Publicado: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
por: Song, Wei, et al.
Publicado: (2025)
por: Song, Wei, et al.
Publicado: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
por: Cheng, Zihui, et al.
Publicado: (2025)
por: Cheng, Zihui, et al.
Publicado: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
por: Li, Jiaang, et al.
Publicado: (2025)
por: Li, Jiaang, et al.
Publicado: (2025)
Do Vision-Language Models Understand Visual Persuasiveness?
por: Park, Gyuwon
Publicado: (2025)
por: Park, Gyuwon
Publicado: (2025)
Harnessing Webpage UIs for Text-Rich Visual Understanding
por: Liu, Junpeng, et al.
Publicado: (2024)
por: Liu, Junpeng, et al.
Publicado: (2024)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
por: Ma, David, et al.
Publicado: (2025)
por: Ma, David, et al.
Publicado: (2025)
ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
por: Barua, Deeparghya Dutta, et al.
Publicado: (2024)
por: Barua, Deeparghya Dutta, et al.
Publicado: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
por: Li, Wenyan, et al.
Publicado: (2024)
por: Li, Wenyan, et al.
Publicado: (2024)
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
por: Atuhurra, Jesse, et al.
Publicado: (2024)
por: Atuhurra, Jesse, et al.
Publicado: (2024)
Video Understanding with Large Language Models: A Survey
por: Tang, Yolo Y., et al.
Publicado: (2023)
por: Tang, Yolo Y., et al.
Publicado: (2023)
Deep Learning based Visually Rich Document Content Understanding: A Survey
por: Ding, Yihao, et al.
Publicado: (2024)
por: Ding, Yihao, et al.
Publicado: (2024)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
por: Kuang, Jiayi, et al.
Publicado: (2024)
por: Kuang, Jiayi, et al.
Publicado: (2024)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
por: Yue, Xiang, et al.
Publicado: (2024)
por: Yue, Xiang, et al.
Publicado: (2024)
Ejemplares similares
-
PixT3: Pixel-based Table-To-Text Generation
por: Alonso, Iñigo, et al.
Publicado: (2023) -
Adding simple structure at inference improves Vision-Language Compositionality
por: Miranda, Imanol, et al.
Publicado: (2025) -
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
por: Miranda, Imanol, et al.
Publicado: (2024) -
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
por: Miranda, Imanol, et al.
Publicado: (2026) -
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
por: Gupta, Akash, et al.
Publicado: (2025)