Hierarchical Document Parsing via Large Margin Feature Matching and Heuristics
Fuente:
arXiv
Guardado en:
| Autor principal: | Kiet, Duong Anh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Elderly Activity Recognition in the Wild: Results from the EAR Challenge
por: Duong, Anh-Kiet
Publicado: (2025)
por: Duong, Anh-Kiet
Publicado: (2025)
Efficient Document Parsing via Parallel Token Prediction
por: Li, Lei, et al.
Publicado: (2026)
por: Li, Lei, et al.
Publicado: (2026)
Addressing Out-of-Label Hazard Detection in Dashcam Videos: Insights from the COOOL Challenge
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
Action Recognition Using Temporal Shift Module and Ensemble Learning
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
Scalable Framework for Classifying AI-Generated Content Across Modalities
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
por: Duong, Anh-Kiet, et al.
Publicado: (2025)
Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models
por: Saim, Mohammad, et al.
Publicado: (2025)
por: Saim, Mohammad, et al.
Publicado: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
por: Tang, Zihan, et al.
Publicado: (2026)
por: Tang, Zihan, et al.
Publicado: (2026)
Leveraging feature communication in federated learning for remote sensing image classification
por: Duong, Anh-Kiet, et al.
Publicado: (2024)
por: Duong, Anh-Kiet, et al.
Publicado: (2024)
Enhanced Face Authentication With Separate Loss Functions
por: Duong, Anh-Kiet, et al.
Publicado: (2023)
por: Duong, Anh-Kiet, et al.
Publicado: (2023)
Collision-Resistant Single-Pass Method for Unsupervised Fine-Grained Image Hashing
por: Duong, Anh-Kiet, et al.
Publicado: (2026)
por: Duong, Anh-Kiet, et al.
Publicado: (2026)
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
por: Wang, Bin, et al.
Publicado: (2026)
por: Wang, Bin, et al.
Publicado: (2026)
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
por: Niu, Junbo, et al.
Publicado: (2025)
por: Niu, Junbo, et al.
Publicado: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
por: Xu, Bangrui, et al.
Publicado: (2026)
por: Xu, Bangrui, et al.
Publicado: (2026)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
por: Nguyen, Thong, et al.
Publicado: (2024)
por: Nguyen, Thong, et al.
Publicado: (2024)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
por: Yang, Minglai, et al.
Publicado: (2026)
por: Yang, Minglai, et al.
Publicado: (2026)
From Image Hashing to Scene Change Detection
por: Duong, Anh-Kiet, et al.
Publicado: (2026)
por: Duong, Anh-Kiet, et al.
Publicado: (2026)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
por: Yu, Wenwen, et al.
Publicado: (2025)
por: Yu, Wenwen, et al.
Publicado: (2025)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
por: Zhang, Qintong, et al.
Publicado: (2024)
por: Zhang, Qintong, et al.
Publicado: (2024)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
por: Wei, Xinyu, et al.
Publicado: (2025)
por: Wei, Xinyu, et al.
Publicado: (2025)
Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions
por: Zhang, Xiao, et al.
Publicado: (2025)
por: Zhang, Xiao, et al.
Publicado: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
por: Wang, Zhengren, et al.
Publicado: (2026)
por: Wang, Zhengren, et al.
Publicado: (2026)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
por: Wang, Baode, et al.
Publicado: (2025)
por: Wang, Baode, et al.
Publicado: (2025)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
por: Zhang, Kaichen, et al.
Publicado: (2024)
por: Zhang, Kaichen, et al.
Publicado: (2024)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
por: Guo, Yansong, et al.
Publicado: (2026)
por: Guo, Yansong, et al.
Publicado: (2026)
ParseBench: A Document Parsing Benchmark for AI Agents
por: Zhang, Boyang, et al.
Publicado: (2026)
por: Zhang, Boyang, et al.
Publicado: (2026)
Lightweight and Production-Ready PDF Visual Element Parsing
por: Liu, Meizhu, et al.
Publicado: (2026)
por: Liu, Meizhu, et al.
Publicado: (2026)
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
por: Li, Mingcheng, et al.
Publicado: (2024)
por: Li, Mingcheng, et al.
Publicado: (2024)
HSD: Training-Free Acceleration for Document Parsing Vision-Language Model with Hierarchical Speculative Decoding
por: Liao, Wenhui, et al.
Publicado: (2026)
por: Liao, Wenhui, et al.
Publicado: (2026)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
por: Yu, Zhou, et al.
Publicado: (2023)
por: Yu, Zhou, et al.
Publicado: (2023)
Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
por: Shinozaki, Taiga, et al.
Publicado: (2025)
por: Shinozaki, Taiga, et al.
Publicado: (2025)
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
por: Feng, Hao, et al.
Publicado: (2025)
por: Feng, Hao, et al.
Publicado: (2025)
Hierarchical Windowed Graph Attention Network and a Large Scale Dataset for Isolated Indian Sign Language Recognition
por: Patra, Suvajit, et al.
Publicado: (2024)
por: Patra, Suvajit, et al.
Publicado: (2024)
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
por: Futeral, Matthieu, et al.
Publicado: (2024)
por: Futeral, Matthieu, et al.
Publicado: (2024)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
por: Luo, Chuwei, et al.
Publicado: (2024)
por: Luo, Chuwei, et al.
Publicado: (2024)
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
por: Fuller, Harrison, et al.
Publicado: (2025)
por: Fuller, Harrison, et al.
Publicado: (2025)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
por: Ji, Yifan, et al.
Publicado: (2026)
por: Ji, Yifan, et al.
Publicado: (2026)
GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
por: Gao, Nan, et al.
Publicado: (2023)
por: Gao, Nan, et al.
Publicado: (2023)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
por: Nguyen, Cong-Duy, et al.
Publicado: (2025)
por: Nguyen, Cong-Duy, et al.
Publicado: (2025)
Ejemplares similares
-
Elderly Activity Recognition in the Wild: Results from the EAR Challenge
por: Duong, Anh-Kiet
Publicado: (2025) -
Efficient Document Parsing via Parallel Token Prediction
por: Li, Lei, et al.
Publicado: (2026) -
Addressing Out-of-Label Hazard Detection in Dashcam Videos: Insights from the COOOL Challenge
por: Duong, Anh-Kiet, et al.
Publicado: (2025) -
Action Recognition Using Temporal Shift Module and Ensemble Learning
por: Duong, Anh-Kiet, et al.
Publicado: (2025) -
Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
por: Duong, Anh-Kiet, et al.
Publicado: (2025)