Gespeichert in:
| Hauptverfasser: | Long, Rujiao, Xing, Hangdi, Yang, Zhibo, Zheng, Qi, Yu, Zhi, Yao, Cong, Huang, Fei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2401.01522 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
von: Long, Rujiao, et al.
Veröffentlicht: (2024)
von: Long, Rujiao, et al.
Veröffentlicht: (2024)
WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
von: Shao, Zirui, et al.
Veröffentlicht: (2024)
von: Shao, Zirui, et al.
Veröffentlicht: (2024)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024)
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024)
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
von: Zhu, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Zhu, Zhaoqing, et al.
Veröffentlicht: (2025)
Robust Fine-tuning for Pre-trained 3D Point Cloud Models
von: Zhang, Zhibo, et al.
Veröffentlicht: (2024)
von: Zhang, Zhibo, et al.
Veröffentlicht: (2024)
LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders
von: Khodabandeh, Borna, et al.
Veröffentlicht: (2025)
von: Khodabandeh, Borna, et al.
Veröffentlicht: (2025)
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
von: Zhang, Duzhen, et al.
Veröffentlicht: (2025)
von: Zhang, Duzhen, et al.
Veröffentlicht: (2025)
Platypus: A Generalized Specialist Model for Reading Text in Various Forms
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
SepFormer: Coarse-to-fine Separator Regression Network for Table Structure Recognition
von: Nguyen, Nam Quan, et al.
Veröffentlicht: (2025)
von: Nguyen, Nam Quan, et al.
Veröffentlicht: (2025)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
von: Zhang, Yuyi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuyi, et al.
Veröffentlicht: (2024)
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Fake It Right: Injecting Anatomical Logic into Synthetic Supervised Pre-training for Medical Segmentation
von: Tang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Tang, Jiaqi, et al.
Veröffentlicht: (2026)
Let ViT Speak: Generative Language-Image Pre-training
von: Fang, Yan, et al.
Veröffentlicht: (2026)
von: Fang, Yan, et al.
Veröffentlicht: (2026)
Long-Tailed Recognition on Binary Networks by Calibrating A Pre-trained Model
von: Kim, Jihun, et al.
Veröffentlicht: (2024)
von: Kim, Jihun, et al.
Veröffentlicht: (2024)
Micro-Expression Recognition by Motion Feature Extraction based on Pre-training
von: Li, Ruolin, et al.
Veröffentlicht: (2024)
von: Li, Ruolin, et al.
Veröffentlicht: (2024)
MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2025)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2025)
Deep Radar Inverse Sensor Models for Dynamic Occupancy Grid Maps
von: Wei, Zihang, et al.
Veröffentlicht: (2023)
von: Wei, Zihang, et al.
Veröffentlicht: (2023)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
Visual Text Generation in the Wild
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
Stylized Structural Patterns for Improved Neural Network Pre-training
von: Salehi, Farnood, et al.
Veröffentlicht: (2025)
von: Salehi, Farnood, et al.
Veröffentlicht: (2025)
Pre-training for Action Recognition with Automatically Generated Fractal Datasets
von: Svyezhentsev, Davyd, et al.
Veröffentlicht: (2024)
von: Svyezhentsev, Davyd, et al.
Veröffentlicht: (2024)
IPAD: Iterative, Parallel, and Diffusion-based Network for Scene Text Recognition
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2023)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2023)
A Foundation Model for DAS Signal Recognition and Visual Prompt Tuning of the Pre-trained Model for Downstream Tasks
von: Gui, Kun, et al.
Veröffentlicht: (2025)
von: Gui, Kun, et al.
Veröffentlicht: (2025)
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
von: Liu, Liwen, et al.
Veröffentlicht: (2025)
von: Liu, Liwen, et al.
Veröffentlicht: (2025)
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
von: Shen, Shufan, et al.
Veröffentlicht: (2025)
von: Shen, Shufan, et al.
Veröffentlicht: (2025)
Sparse Reasoning is Enough: Biological-Inspired Framework for Video Anomaly Detection with Large Pre-trained Models
von: Huang, He, et al.
Veröffentlicht: (2025)
von: Huang, He, et al.
Veröffentlicht: (2025)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
von: Huang, Ziyuan, et al.
Veröffentlicht: (2024)
von: Huang, Ziyuan, et al.
Veröffentlicht: (2024)
Logos as a Well-Tempered Pre-train for Sign Language Recognition
von: Ovodov, Ilya, et al.
Veröffentlicht: (2025)
von: Ovodov, Ilya, et al.
Veröffentlicht: (2025)
AdFair-CLIP: Adversarial Fair Contrastive Language-Image Pre-training for Chest X-rays
von: Yi, Chenlang, et al.
Veröffentlicht: (2025)
von: Yi, Chenlang, et al.
Veröffentlicht: (2025)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
von: Zheng, Weijie, et al.
Veröffentlicht: (2024)
ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data
von: Shen, Yufan, et al.
Veröffentlicht: (2024)
von: Shen, Yufan, et al.
Veröffentlicht: (2024)
SR-Stereo & DAPE: Stepwise Regression and Pre-trained Edges for Practical Stereo Matching
von: Xiao, Weiqing, et al.
Veröffentlicht: (2024)
von: Xiao, Weiqing, et al.
Veröffentlicht: (2024)
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
von: Lu, Feng, et al.
Veröffentlicht: (2024)
von: Lu, Feng, et al.
Veröffentlicht: (2024)
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
von: Haliassos, Alexandros, et al.
Veröffentlicht: (2024)
von: Haliassos, Alexandros, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction
von: Long, Rujiao, et al.
Veröffentlicht: (2024) -
WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
von: Shao, Zirui, et al.
Veröffentlicht: (2024) -
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024) -
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
von: Zhu, Zhaoqing, et al.
Veröffentlicht: (2025) -
Robust Fine-tuning for Pre-trained 3D Point Cloud Models
von: Zhang, Zhibo, et al.
Veröffentlicht: (2024)