HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Rujiao, Wang, Pengfei, Yang, Zhibo, Yao, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training
by: Long, Rujiao, et al.
Published: (2024)
by: Long, Rujiao, et al.
Published: (2024)
Generative Compositor for Few-Shot Visual Information Extraction
by: Yang, Zhibo, et al.
Published: (2025)
by: Yang, Zhibo, et al.
Published: (2025)
Robust Fine-tuning for Pre-trained 3D Point Cloud Models
by: Zhang, Zhibo, et al.
Published: (2024)
by: Zhang, Zhibo, et al.
Published: (2024)
VILA: On Pre-training for Visual Language Models
by: Lin, Ji, et al.
Published: (2023)
by: Lin, Ji, et al.
Published: (2023)
Towards Scalable Pre-training of Visual Tokenizers for Generation
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
In Pursuit of Pixel Supervision for Visual Pre-training
by: Yang, Lihe, et al.
Published: (2025)
by: Yang, Lihe, et al.
Published: (2025)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
by: Wan, Jianqiang, et al.
Published: (2024)
by: Wan, Jianqiang, et al.
Published: (2024)
Micro-Expression Recognition by Motion Feature Extraction based on Pre-training
by: Li, Ruolin, et al.
Published: (2024)
by: Li, Ruolin, et al.
Published: (2024)
Pre-training Point Cloud Compact Model with Partial-aware Reconstruction
by: Zha, Yaohua, et al.
Published: (2024)
by: Zha, Yaohua, et al.
Published: (2024)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
by: Ji, Yifan, et al.
Published: (2026)
by: Ji, Yifan, et al.
Published: (2026)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
by: Zhang, Yuyi, et al.
Published: (2024)
by: Zhang, Yuyi, et al.
Published: (2024)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
by: Tian, Qingyao, et al.
Published: (2025)
by: Tian, Qingyao, et al.
Published: (2025)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
Formula-Supervised Visual-Geometric Pre-training
by: Yamada, Ryosuke, et al.
Published: (2024)
by: Yamada, Ryosuke, et al.
Published: (2024)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark
by: Cui, Ziteng, et al.
Published: (2025)
by: Cui, Ziteng, et al.
Published: (2025)
VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained Models
by: Yin, Ziyi, et al.
Published: (2024)
by: Yin, Ziyi, et al.
Published: (2024)
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
by: Liu, Liwen, et al.
Published: (2025)
by: Liu, Liwen, et al.
Published: (2025)
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition
by: Liu, Ran, et al.
Published: (2025)
by: Liu, Ran, et al.
Published: (2025)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
SDPT: Synchronous Dual Prompt Tuning for Fusion-based Visual-Language Pre-trained Models
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
Platypus: A Generalized Specialist Model for Reading Text in Various Forms
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
4D Visual Pre-training for Robot Learning
by: Hou, Chengkai, et al.
Published: (2025)
by: Hou, Chengkai, et al.
Published: (2025)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images
by: Cui, Ziteng, et al.
Published: (2024)
by: Cui, Ziteng, et al.
Published: (2024)
Hierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language Model
by: Sun, Weilin, et al.
Published: (2025)
by: Sun, Weilin, et al.
Published: (2025)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
by: Wang, Hengkang, et al.
Published: (2025)
by: Wang, Hengkang, et al.
Published: (2025)
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
by: Zhang, Duzhen, et al.
Published: (2025)
by: Zhang, Duzhen, et al.
Published: (2025)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
by: Lin, Zhiwei, et al.
Published: (2022)
by: Lin, Zhiwei, et al.
Published: (2022)
Continual Forgetting for Pre-trained Vision Models
by: Zhao, Hongbo, et al.
Published: (2024)
by: Zhao, Hongbo, et al.
Published: (2024)
Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation
by: Yao, Yuanhang, et al.
Published: (2026)
by: Yao, Yuanhang, et al.
Published: (2026)
LapFM: A Laparoscopic Segmentation Foundation Model via Hierarchical Concept Evolving Pre-training
by: Xu, Qing, et al.
Published: (2025)
by: Xu, Qing, et al.
Published: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
by: Albastaki, Shahad, et al.
Published: (2025)
by: Albastaki, Shahad, et al.
Published: (2025)
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
Efficient and Effective Universal Adversarial Attack against Vision-Language Pre-training Models
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Similar Items
-
LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training
by: Long, Rujiao, et al.
Published: (2024) -
Generative Compositor for Few-Shot Visual Information Extraction
by: Yang, Zhibo, et al.
Published: (2025) -
Robust Fine-tuning for Pre-trained 3D Point Cloud Models
by: Zhang, Zhibo, et al.
Published: (2024) -
VILA: On Pre-training for Visual Language Models
by: Lin, Ji, et al.
Published: (2023) -
Towards Scalable Pre-training of Visual Tokenizers for Generation
by: Yao, Jingfeng, et al.
Published: (2025)