Lightweight and Production-Ready PDF Visual Element Parsing
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Meizhu, Abbasi, Yassi, Rowe, Matthew, Avendi, Michael, Li, Paul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do Image-Text Metrics Respect Semantic Invariances?
por: Agarwal, Amit, et al.
Publicado: (2026)
por: Agarwal, Amit, et al.
Publicado: (2026)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
por: He, Jiashu, et al.
Publicado: (2026)
por: He, Jiashu, et al.
Publicado: (2026)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
por: Xie, Xudong, et al.
Publicado: (2024)
por: Xie, Xudong, et al.
Publicado: (2024)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
por: Yang, Minglai, et al.
Publicado: (2026)
por: Yang, Minglai, et al.
Publicado: (2026)
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
por: Xu, Bangrui, et al.
Publicado: (2026)
por: Xu, Bangrui, et al.
Publicado: (2026)
Lightweight Operations for Visual Speech Recognition
por: Panagos, Iason Ioannis, et al.
Publicado: (2025)
por: Panagos, Iason Ioannis, et al.
Publicado: (2025)
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
por: Zhu, Fanwei, et al.
Publicado: (2025)
por: Zhu, Fanwei, et al.
Publicado: (2025)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
por: Baek, Jeonghun, et al.
Publicado: (2025)
por: Baek, Jeonghun, et al.
Publicado: (2025)
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
por: Tang, Yiwen, et al.
Publicado: (2025)
por: Tang, Yiwen, et al.
Publicado: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
Red Teaming Visual Language Models
por: Li, Mukai, et al.
Publicado: (2024)
por: Li, Mukai, et al.
Publicado: (2024)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
por: Zhang, Qintong, et al.
Publicado: (2024)
por: Zhang, Qintong, et al.
Publicado: (2024)
InfoDet: A Dataset for Infographic Element Detection
por: Zhu, Jiangning, et al.
Publicado: (2025)
por: Zhu, Jiangning, et al.
Publicado: (2025)
Unbiased Visual Reasoning with Controlled Visual Inputs
por: Li, Zhaonan, et al.
Publicado: (2025)
por: Li, Zhaonan, et al.
Publicado: (2025)
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
por: Li, Mukai, et al.
Publicado: (2024)
por: Li, Mukai, et al.
Publicado: (2024)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
por: Wang, Baode, et al.
Publicado: (2025)
por: Wang, Baode, et al.
Publicado: (2025)
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
por: Wang, Zifu, et al.
Publicado: (2025)
por: Wang, Zifu, et al.
Publicado: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
por: Kim, Taewhan, et al.
Publicado: (2024)
por: Kim, Taewhan, et al.
Publicado: (2024)
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
por: Lai, Xin, et al.
Publicado: (2025)
por: Lai, Xin, et al.
Publicado: (2025)
VisualRWKV-HD and UHD: Advancing High-Resolution Processing for Visual Language Models
por: Li, Zihang, et al.
Publicado: (2024)
por: Li, Zihang, et al.
Publicado: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
por: Li, Bo, et al.
Publicado: (2024)
por: Li, Bo, et al.
Publicado: (2024)
Semantically-Prompted Language Models Improve Visual Descriptions
por: Ogezi, Michael, et al.
Publicado: (2023)
por: Ogezi, Michael, et al.
Publicado: (2023)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
por: Ouyang, Linke, et al.
Publicado: (2024)
por: Ouyang, Linke, et al.
Publicado: (2024)
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
por: Son, Jaemin, et al.
Publicado: (2025)
por: Son, Jaemin, et al.
Publicado: (2025)
AC-Lite : A Lightweight Image Captioning Model for Low-Resource Assamese Language
por: Choudhury, Pankaj, et al.
Publicado: (2025)
por: Choudhury, Pankaj, et al.
Publicado: (2025)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
por: Zhu, Minjun, et al.
Publicado: (2026)
por: Zhu, Minjun, et al.
Publicado: (2026)
On Data Synthesis and Post-training for Visual Abstract Reasoning
por: Zhu, Ke, et al.
Publicado: (2025)
por: Zhu, Ke, et al.
Publicado: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
por: Jing, Hongyi, et al.
Publicado: (2025)
por: Jing, Hongyi, et al.
Publicado: (2025)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
por: Li, Kailing, et al.
Publicado: (2025)
por: Li, Kailing, et al.
Publicado: (2025)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Chatting with Images for Introspective Visual Thinking
por: Wu, Junfei, et al.
Publicado: (2026)
por: Wu, Junfei, et al.
Publicado: (2026)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
por: Li, Bin, et al.
Publicado: (2022)
por: Li, Bin, et al.
Publicado: (2022)
VGR: Visual Grounded Reasoning
por: Wang, Jiacong, et al.
Publicado: (2025)
por: Wang, Jiacong, et al.
Publicado: (2025)
VIALM: A Survey and Benchmark of Visually Impaired Assistance with Large Models
por: Zhao, Yi, et al.
Publicado: (2024)
por: Zhao, Yi, et al.
Publicado: (2024)
CognArtive: Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements
por: Khadangi, Afshin, et al.
Publicado: (2025)
por: Khadangi, Afshin, et al.
Publicado: (2025)
The Role of Entropy in Visual Grounding: Analysis and Optimization
por: Li, Shuo, et al.
Publicado: (2025)
por: Li, Shuo, et al.
Publicado: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
por: Sarch, Gabriel, et al.
Publicado: (2026)
por: Sarch, Gabriel, et al.
Publicado: (2026)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
por: Sharma, Aditya, et al.
Publicado: (2024)
por: Sharma, Aditya, et al.
Publicado: (2024)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
por: Wu, Chengyue, et al.
Publicado: (2024)
por: Wu, Chengyue, et al.
Publicado: (2024)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
por: Liu, Yuhan, et al.
Publicado: (2025)
por: Liu, Yuhan, et al.
Publicado: (2025)
Ejemplares similares
-
Do Image-Text Metrics Respect Semantic Invariances?
por: Agarwal, Amit, et al.
Publicado: (2026) -
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
por: He, Jiashu, et al.
Publicado: (2026) -
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
por: Xie, Xudong, et al.
Publicado: (2024) -
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
por: Yang, Minglai, et al.
Publicado: (2026) -
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
por: Xu, Bangrui, et al.
Publicado: (2026)