KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
Fuente:
arXiv
Guardado en:
| Autores principales: | Hwang, Taebaek, Kim, Minseo, Lee, Gisang, Kim, Seonuk, Eun, Hyunjun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
por: Chung, Jiwan, et al.
Publicado: (2024)
por: Chung, Jiwan, et al.
Publicado: (2024)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
por: Kim, Yoonshik, et al.
Publicado: (2025)
por: Kim, Yoonshik, et al.
Publicado: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
por: Han, Donghoon, et al.
Publicado: (2024)
por: Han, Donghoon, et al.
Publicado: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
por: Rostamkhani, Mohammadmostafa, et al.
Publicado: (2024)
por: Rostamkhani, Mohammadmostafa, et al.
Publicado: (2024)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
por: Kim, Geewook, et al.
Publicado: (2024)
por: Kim, Geewook, et al.
Publicado: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
por: Shrestha, Robik, et al.
Publicado: (2020)
por: Shrestha, Robik, et al.
Publicado: (2020)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
por: Wang, Haibo, et al.
Publicado: (2024)
por: Wang, Haibo, et al.
Publicado: (2024)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
por: Yu, Suhao, et al.
Publicado: (2025)
por: Yu, Suhao, et al.
Publicado: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
por: Chung, Jiwan, et al.
Publicado: (2025)
por: Chung, Jiwan, et al.
Publicado: (2025)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
por: Rahman, Mizanur, et al.
Publicado: (2025)
por: Rahman, Mizanur, et al.
Publicado: (2025)
GRAM: Global Reasoning for Multi-Page VQA
por: Blau, Tsachi, et al.
Publicado: (2024)
por: Blau, Tsachi, et al.
Publicado: (2024)
Harnessing Webpage UIs for Text-Rich Visual Understanding
por: Liu, Junpeng, et al.
Publicado: (2024)
por: Liu, Junpeng, et al.
Publicado: (2024)
See the Text: From Tokenization to Visual Reading
por: Xing, Ling, et al.
Publicado: (2025)
por: Xing, Ling, et al.
Publicado: (2025)
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
por: Ju, Jeongho, et al.
Publicado: (2024)
por: Ju, Jeongho, et al.
Publicado: (2024)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
por: Li, Xuchen, et al.
Publicado: (2024)
por: Li, Xuchen, et al.
Publicado: (2024)
Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
por: Cho, Yeongjae, et al.
Publicado: (2025)
por: Cho, Yeongjae, et al.
Publicado: (2025)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
por: Dong, Qihua, et al.
Publicado: (2026)
por: Dong, Qihua, et al.
Publicado: (2026)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
por: Ma, Dongsheng, et al.
Publicado: (2026)
por: Ma, Dongsheng, et al.
Publicado: (2026)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
por: Ge, Jinchao, et al.
Publicado: (2025)
por: Ge, Jinchao, et al.
Publicado: (2025)
Attribute Diversity Determines the Systematicity Gap in VQA
por: Berlot-Attwell, Ian, et al.
Publicado: (2023)
por: Berlot-Attwell, Ian, et al.
Publicado: (2023)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
por: Liang, Jianxin, et al.
Publicado: (2025)
por: Liang, Jianxin, et al.
Publicado: (2025)
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective
por: Yu, Xinmiao, et al.
Publicado: (2024)
por: Yu, Xinmiao, et al.
Publicado: (2024)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
por: Zhang, Yanzhe, et al.
Publicado: (2023)
por: Zhang, Yanzhe, et al.
Publicado: (2023)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
por: Hwang, Yerin, et al.
Publicado: (2025)
por: Hwang, Yerin, et al.
Publicado: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
por: Fallah, Forouzan, et al.
Publicado: (2025)
por: Fallah, Forouzan, et al.
Publicado: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
por: Guo, Pinxue, et al.
Publicado: (2025)
por: Guo, Pinxue, et al.
Publicado: (2025)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
por: Su, Xin, et al.
Publicado: (2024)
por: Su, Xin, et al.
Publicado: (2024)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
por: Kim, Taewhan, et al.
Publicado: (2024)
por: Kim, Taewhan, et al.
Publicado: (2024)
Evaluating Multimodal Generative AI with Korean Educational Standards
por: Park, Sanghee, et al.
Publicado: (2025)
por: Park, Sanghee, et al.
Publicado: (2025)
Entropic Context Shaping: Information-Theoretic Filtering for Context-Aware LLM Agents
por: Kim, Hyunjun
Publicado: (2026)
por: Kim, Hyunjun
Publicado: (2026)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
por: Ye, Maoyuan, et al.
Publicado: (2025)
por: Ye, Maoyuan, et al.
Publicado: (2025)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
por: Li, Zhuowan, et al.
Publicado: (2024)
por: Li, Zhuowan, et al.
Publicado: (2024)
Hallucination Benchmark in Medical Visual Question Answering
por: Wu, Jinge, et al.
Publicado: (2024)
por: Wu, Jinge, et al.
Publicado: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
por: Fang, I-Sheng, et al.
Publicado: (2025)
por: Fang, I-Sheng, et al.
Publicado: (2025)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
por: Sood, Ekta, et al.
Publicado: (2021)
por: Sood, Ekta, et al.
Publicado: (2021)
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
por: Kim, Keonwoo, et al.
Publicado: (2025)
por: Kim, Keonwoo, et al.
Publicado: (2025)
LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
por: Xiao, Yijia, et al.
Publicado: (2024)
por: Xiao, Yijia, et al.
Publicado: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
por: Ding, Yihao, et al.
Publicado: (2024)
por: Ding, Yihao, et al.
Publicado: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
por: Wu, Yin, et al.
Publicado: (2025)
por: Wu, Yin, et al.
Publicado: (2025)
Ejemplares similares
-
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
por: Chung, Jiwan, et al.
Publicado: (2024) -
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
por: Kim, Yoonshik, et al.
Publicado: (2025) -
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
por: Han, Donghoon, et al.
Publicado: (2024) -
CommVQA: Situating Visual Question Answering in Communicative Contexts
por: Naik, Nandita Shankar, et al.
Publicado: (2024) -
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
por: Rostamkhani, Mohammadmostafa, et al.
Publicado: (2024)