Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jaeyoo, Choi, Jin Young, Park, Jeonghyung, Han, Bohyung |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Class Feature Augmentation for Class Incremental Learning
by: Kim, Taehoon, et al.
Published: (2023)
by: Kim, Taehoon, et al.
Published: (2023)
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025)
by: Park, Jaeyoo, et al.
Published: (2025)
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
by: Baek, Kanghyun, et al.
Published: (2025)
by: Baek, Kanghyun, et al.
Published: (2025)
Free-Grained Hierarchical Visual Recognition
by: Park, Seulki, et al.
Published: (2025)
by: Park, Seulki, et al.
Published: (2025)
A Training-Free Defense Framework for Robust Learned Image Compression
by: Song, Myungseo, et al.
Published: (2024)
by: Song, Myungseo, et al.
Published: (2024)
VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
by: Abramovich, Ofir, et al.
Published: (2024)
by: Abramovich, Ofir, et al.
Published: (2024)
Video Deblurring with Deconvolution and Aggregation Networks
by: Choi, Giyong, et al.
Published: (2025)
by: Choi, Giyong, et al.
Published: (2025)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
Exploring OCR-augmented Generation for Bilingual VQA
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
by: Kang, Hyolim, et al.
Published: (2025)
by: Kang, Hyolim, et al.
Published: (2025)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
by: Lee, Junsung, et al.
Published: (2025)
by: Lee, Junsung, et al.
Published: (2025)
Do Vision-Language Models Understand Visual Persuasiveness?
by: Park, Gyuwon
Published: (2025)
by: Park, Gyuwon
Published: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
by: Tien, Dong Nguyen, et al.
Published: (2025)
by: Tien, Dong Nguyen, et al.
Published: (2025)
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
by: Le, Binh M., et al.
Published: (2025)
by: Le, Binh M., et al.
Published: (2025)
Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance
by: Lee, Hyunsoo, et al.
Published: (2024)
by: Lee, Hyunsoo, et al.
Published: (2024)
Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation
by: Lee, Junsung, et al.
Published: (2024)
by: Lee, Junsung, et al.
Published: (2024)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
by: Kim, Mijeong, et al.
Published: (2026)
by: Kim, Mijeong, et al.
Published: (2026)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
by: Kim, Minji, et al.
Published: (2025)
by: Kim, Minji, et al.
Published: (2025)
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
by: Hu, Anwen, et al.
Published: (2024)
by: Hu, Anwen, et al.
Published: (2024)
FIFO-Diffusion: Generating Infinite Videos from Text without Training
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
by: Hu, Anwen, et al.
Published: (2024)
by: Hu, Anwen, et al.
Published: (2024)
Visually Consistent Hierarchical Image Classification
by: Park, Seulki, et al.
Published: (2024)
by: Park, Seulki, et al.
Published: (2024)
A Training-Free Style-Personalization via SVD-Based Feature Decomposition
by: Lee, Kyoungmin, et al.
Published: (2025)
by: Lee, Kyoungmin, et al.
Published: (2025)
ODGS: 3D Scene Reconstruction from Omnidirectional Images with 3D Gaussian Splattings
by: Lee, Suyoung, et al.
Published: (2024)
by: Lee, Suyoung, et al.
Published: (2024)
CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding
by: Jeon, Yongjun, et al.
Published: (2026)
by: Jeon, Yongjun, et al.
Published: (2026)
HCF: Hierarchical Cascade Framework for Distributed Multi-Stage Image Compression
by: Cai, Junhao, et al.
Published: (2025)
by: Cai, Junhao, et al.
Published: (2025)
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Understanding Open-Set Recognition by Jacobian Norm and Inter-Class Separation
by: Park, Jaewoo, et al.
Published: (2022)
by: Park, Jaewoo, et al.
Published: (2022)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
by: Byun, Ji Young, et al.
Published: (2025)
by: Byun, Ji Young, et al.
Published: (2025)
CGTrack: Cascade Gating Network with Hierarchical Feature Aggregation for UAV Tracking
by: Li, Weihong, et al.
Published: (2025)
by: Li, Weihong, et al.
Published: (2025)
Multimodal OCR: Parse Anything from Documents
by: Zheng, Handong, et al.
Published: (2026)
by: Zheng, Handong, et al.
Published: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
Beyond the Ground Truth: Enhanced Supervision for Image Restoration
by: Ryou, Donghun, et al.
Published: (2025)
by: Ryou, Donghun, et al.
Published: (2025)
Task-Agnostic Noisy Label Detection via Standardized Loss Aggregation
by: Park, Inhyuk, et al.
Published: (2026)
by: Park, Inhyuk, et al.
Published: (2026)
UL-VIO: Ultra-lightweight Visual-Inertial Odometry with Noise Robust Test-time Adaptation
by: Park, Jinho, et al.
Published: (2024)
by: Park, Jinho, et al.
Published: (2024)
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR
by: Yuan, Ziye, et al.
Published: (2026)
by: Yuan, Ziye, et al.
Published: (2026)
Similar Items
-
Cross-Class Feature Augmentation for Class Incremental Learning
by: Kim, Taehoon, et al.
Published: (2023) -
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025) -
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
by: Baek, Kanghyun, et al.
Published: (2025) -
Free-Grained Hierarchical Visual Recognition
by: Park, Seulki, et al.
Published: (2025) -
A Training-Free Defense Framework for Robust Learned Image Compression
by: Song, Myungseo, et al.
Published: (2024)