LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Huawen, Li, Gengluo, Zhong, Jinwen, Zhou, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
von: Li, Gengluo, et al.
Veröffentlicht: (2025)
von: Li, Gengluo, et al.
Veröffentlicht: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
von: Luo, Yaxin, et al.
Veröffentlicht: (2026)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025)
Parrot: Multilingual Visual Instruction Tuning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
von: You, Zebin, et al.
Veröffentlicht: (2025)
von: You, Zebin, et al.
Veröffentlicht: (2025)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
Context-Aware Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
von: Ging, Simon, et al.
Veröffentlicht: (2026)
von: Ging, Simon, et al.
Veröffentlicht: (2026)
Unifying Specialized Visual Encoders for Video Language Models
von: Chung, Jihoon, et al.
Veröffentlicht: (2025)
von: Chung, Jihoon, et al.
Veröffentlicht: (2025)
Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
von: Li, Sijie, et al.
Veröffentlicht: (2026)
von: Li, Sijie, et al.
Veröffentlicht: (2026)
RealKIE: Five Novel Datasets for Enterprise Key Information Extraction
von: Townsend, Benjamin, et al.
Veröffentlicht: (2024)
von: Townsend, Benjamin, et al.
Veröffentlicht: (2024)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
von: Ma, Yan, et al.
Veröffentlicht: (2025)
von: Ma, Yan, et al.
Veröffentlicht: (2025)
A Practitioner's Guide to Continual Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
Translation-Enhanced Multilingual Text-to-Image Generation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2023)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2023)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
von: Fields, Clayton, et al.
Veröffentlicht: (2024)
von: Fields, Clayton, et al.
Veröffentlicht: (2024)
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
von: Li, Kaican, et al.
Veröffentlicht: (2025)
von: Li, Kaican, et al.
Veröffentlicht: (2025)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
Latent Action Pretraining from Videos
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2024)
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2024)
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
Aya Vision: Advancing the Frontier of Multilingual Multimodality
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
von: Favero, Alessandro, et al.
Veröffentlicht: (2024)
von: Favero, Alessandro, et al.
Veröffentlicht: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
von: Bini, Massimo, et al.
Veröffentlicht: (2025)
von: Bini, Massimo, et al.
Veröffentlicht: (2025)
Mordal: Automated Pretrained Model Selection for Vision Language Models
von: He, Shiqi, et al.
Veröffentlicht: (2025)
von: He, Shiqi, et al.
Veröffentlicht: (2025)
A U-Net and Transformer Pipeline for Multilingual Image Translation
von: Sahay, Siddharth, et al.
Veröffentlicht: (2025)
von: Sahay, Siddharth, et al.
Veröffentlicht: (2025)
SpikeCLIP: A Contrastive Language-Image Pretrained Spiking Neural Network
von: Lv, Changze, et al.
Veröffentlicht: (2023)
von: Lv, Changze, et al.
Veröffentlicht: (2023)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
Composition-Grounded Data Synthesis for Visual Reasoning
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
von: Li, Gengluo, et al.
Veröffentlicht: (2025) -
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
von: Luo, Yaxin, et al.
Veröffentlicht: (2026) -
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
von: Shi, Weijia, et al.
Veröffentlicht: (2024) -
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
von: Zhang, Wenqi, et al.
Veröffentlicht: (2025) -
Parrot: Multilingual Visual Instruction Tuning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)