Hierarchical Pre-Training of Vision Encoders with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Eugene, Chang, Ting-Yu, Tsai, Jui-Huang, Diao, Jiajie, Lee, Chen-Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Select Visual In-Context Demonstrations
by: Lee, Eugene, et al.
Published: (2026)
by: Lee, Eugene, et al.
Published: (2026)
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
by: Whittaker, Edward, et al.
Published: (2024)
by: Whittaker, Edward, et al.
Published: (2024)
A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
by: Schäfer, Henning, et al.
Published: (2025)
by: Schäfer, Henning, et al.
Published: (2025)
Detection of Personal Data in Structured Datasets Using a Large Language Model
by: Ntwali, Albert Agisha, et al.
Published: (2025)
by: Ntwali, Albert Agisha, et al.
Published: (2025)
Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
by: Neptune, Nathalie, et al.
Published: (2025)
by: Neptune, Nathalie, et al.
Published: (2025)
Real-World En Call Center Transcripts Dataset with PII Redaction
by: Dao, Ha, et al.
Published: (2025)
by: Dao, Ha, et al.
Published: (2025)
The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
by: Fenech-Borg, Emanuel Z., et al.
Published: (2025)
by: Fenech-Borg, Emanuel Z., et al.
Published: (2025)
Studying Maps at Scale: A Digital Investigation of Cartography and the Evolution of Figuration
by: Petitpierre, Remi
Published: (2025)
by: Petitpierre, Remi
Published: (2025)
The Table of Media Bias Elements: A sentence-level taxonomy of media bias types and propaganda techniques
by: Menzner, Tim, et al.
Published: (2026)
by: Menzner, Tim, et al.
Published: (2026)
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
by: Wang, Junxin, et al.
Published: (2026)
by: Wang, Junxin, et al.
Published: (2026)
Leveraging Large Language Models for Semantic Query Processing in a Scholarly Knowledge Graph
by: Jia, Runsong, et al.
Published: (2024)
by: Jia, Runsong, et al.
Published: (2024)
SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding
by: Gomi, Keisuke, et al.
Published: (2026)
by: Gomi, Keisuke, et al.
Published: (2026)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
by: Hacheme, Gilles Quentin, et al.
Published: (2025)
by: Hacheme, Gilles Quentin, et al.
Published: (2025)
VLM4Rec: Multimodal Semantic Representation for Recommendation with Large Vision-Language Models
by: Valencia, Ty, et al.
Published: (2026)
by: Valencia, Ty, et al.
Published: (2026)
An Approach for Detection of Entities in Dynamic Media Contents
by: Mbongo, Nzakiese, et al.
Published: (2025)
by: Mbongo, Nzakiese, et al.
Published: (2025)
ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
by: Ahmed, Md Shamim, et al.
Published: (2026)
by: Ahmed, Md Shamim, et al.
Published: (2026)
From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
by: Wang, Mo, et al.
Published: (2026)
by: Wang, Mo, et al.
Published: (2026)
DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
by: Wright, Devin R., et al.
Published: (2025)
by: Wright, Devin R., et al.
Published: (2025)
A Prompt-Aware Structuring Framework for Reliable Reuse of AI-Generated Content in the Agentic Web
by: Egami, Shusaku, et al.
Published: (2026)
by: Egami, Shusaku, et al.
Published: (2026)
Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling
by: Beskorovainyi, Vladimir
Published: (2026)
by: Beskorovainyi, Vladimir
Published: (2026)
Adapting PromptORE for Modern History: Information Extraction from Hispanic Monarchy Documents of the XVIth Century
by: Hidalgo, Hèctor Loopez, et al.
Published: (2024)
by: Hidalgo, Hèctor Loopez, et al.
Published: (2024)
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
by: Quy, Nguyen Lam Phu, et al.
Published: (2025)
by: Quy, Nguyen Lam Phu, et al.
Published: (2025)
Experimenting active and sequential learning in a medieval music manuscript
by: Sharma, Sachin, et al.
Published: (2025)
by: Sharma, Sachin, et al.
Published: (2025)
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
by: Feuer, Benjamin, et al.
Published: (2023)
by: Feuer, Benjamin, et al.
Published: (2023)
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)
by: Slyman, Eric, et al.
Published: (2023)
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
Ontology-based Semantic Similarity Measures for Clustering Medical Concepts in Drug Safety
by: Painter, Jeffery L, et al.
Published: (2025)
by: Painter, Jeffery L, et al.
Published: (2025)
Semantic Similarity-Informed Bayesian Borrowing for Quantitative Signal Detection of Adverse Events
by: Haguinet, François, et al.
Published: (2025)
by: Haguinet, François, et al.
Published: (2025)
Bottleneck-based Encoder-decoder ARchitecture (BEAR) for Learning Unbiased Consumer-to-Consumer Image Representations
by: Rivas, Pablo, et al.
Published: (2024)
by: Rivas, Pablo, et al.
Published: (2024)
IDDR-NGP: Incorporating Detectors for Distractor Removal with Instant Neural Radiance Field
by: Huang, Xianliang, et al.
Published: (2026)
by: Huang, Xianliang, et al.
Published: (2026)
LLMStructBench: Benchmarking Large Language Model Structured Data Extraction
by: Tenckhoff, Sönke, et al.
Published: (2026)
by: Tenckhoff, Sönke, et al.
Published: (2026)
Frequency-Decomposed INR for NIR-Assisted Low-Light RGB Image Denoising
by: Shi, Ligen, et al.
Published: (2026)
by: Shi, Ligen, et al.
Published: (2026)
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
by: Ou, Yiwei, et al.
Published: (2026)
by: Ou, Yiwei, et al.
Published: (2026)
AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety
by: Levi, Adi, et al.
Published: (2025)
by: Levi, Adi, et al.
Published: (2025)
Advancing Annotat3D with Harpia: A CUDA-Accelerated Library For Large-Scale Volumetric Data Segmentation
by: de Araujo, Camila Machado, et al.
Published: (2025)
by: de Araujo, Camila Machado, et al.
Published: (2025)
When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
by: Mbongo, Nzakiese, et al.
Published: (2025)
by: Mbongo, Nzakiese, et al.
Published: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
by: Shopnil, Mir Nafis Sharear, et al.
Published: (2025)
by: Shopnil, Mir Nafis Sharear, et al.
Published: (2025)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
Similar Items
-
Learning to Select Visual In-Context Demonstrations
by: Lee, Eugene, et al.
Published: (2026) -
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
by: Whittaker, Edward, et al.
Published: (2024) -
A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports
by: Schäfer, Henning, et al.
Published: (2025) -
Detection of Personal Data in Structured Datasets Using a Large Language Model
by: Ntwali, Albert Agisha, et al.
Published: (2025) -
Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
by: Neptune, Nathalie, et al.
Published: (2025)