BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xiao, Wu, Chenfei, Rosenman, Shachar, Lal, Vasudev, Che, Wanxiang, Duan, Nan |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
by: Xu, Xiao, et al.
Published: (2025)
by: Xu, Xiao, et al.
Published: (2025)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
by: Luo, Man, et al.
Published: (2025)
by: Luo, Man, et al.
Published: (2025)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
by: Aflalo, Estelle, et al.
Published: (2024)
by: Aflalo, Estelle, et al.
Published: (2024)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
by: Nath, Sujoy, et al.
Published: (2025)
by: Nath, Sujoy, et al.
Published: (2025)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
by: Rosenman, Shachar, et al.
Published: (2023)
by: Rosenman, Shachar, et al.
Published: (2023)
Semi-Instruct: Bridging Natural-Instruct and Self-Instruct for Code Large Language Models
by: Luo, Xianzhen, et al.
Published: (2024)
by: Luo, Xianzhen, et al.
Published: (2024)
Why do LLaVA Vision-Language Models Reply to Images in English?
by: Hinck, Musashi, et al.
Published: (2024)
by: Hinck, Musashi, et al.
Published: (2024)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
by: Xu, Xiao, et al.
Published: (2024)
by: Xu, Xiao, et al.
Published: (2024)
Chitrarth: Bridging Vision and Language for a Billion People
by: Khan, Shaharukh, et al.
Published: (2025)
by: Khan, Shaharukh, et al.
Published: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
by: Bachu, Saketh, et al.
Published: (2024)
by: Bachu, Saketh, et al.
Published: (2024)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
by: Kundu, Souvik, et al.
Published: (2025)
by: Kundu, Souvik, et al.
Published: (2025)
Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization
by: Islam, Md Moinul, et al.
Published: (2025)
by: Islam, Md Moinul, et al.
Published: (2025)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
by: Niu, Tianhao, et al.
Published: (2026)
by: Niu, Tianhao, et al.
Published: (2026)
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
by: Liu, Zixuan, et al.
Published: (2025)
by: Liu, Zixuan, et al.
Published: (2025)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
by: Chen, Qiguang, et al.
Published: (2026)
by: Chen, Qiguang, et al.
Published: (2026)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
by: Wang, Zihu, et al.
Published: (2025)
by: Wang, Zihu, et al.
Published: (2025)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
by: Bandraupalli, Srihari, et al.
Published: (2025)
by: Bandraupalli, Srihari, et al.
Published: (2025)
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization
by: Luo, Richard, et al.
Published: (2024)
by: Luo, Richard, et al.
Published: (2024)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
by: Zhu, Yingjie, et al.
Published: (2025)
by: Zhu, Yingjie, et al.
Published: (2025)
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
by: Du, Yiyang, et al.
Published: (2026)
by: Du, Yiyang, et al.
Published: (2026)
PromptSync: Bridging Domain Gaps in Vision-Language Models through Class-Aware Prototype Alignment and Discrimination
by: Khandelwal, Anant
Published: (2024)
by: Khandelwal, Anant
Published: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale
by: Gong, ZeMing, et al.
Published: (2024)
by: Gong, ZeMing, et al.
Published: (2024)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation
by: Nilukshi, Shashini, et al.
Published: (2026)
by: Nilukshi, Shashini, et al.
Published: (2026)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
by: Kim, Yunsoo, et al.
Published: (2025)
by: Kim, Yunsoo, et al.
Published: (2025)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
by: Fields, Clayton, et al.
Published: (2024)
by: Fields, Clayton, et al.
Published: (2024)
Similar Items
-
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
by: Xu, Xiao, et al.
Published: (2025) -
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
by: Madasu, Avinash, et al.
Published: (2025) -
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024) -
DPO Learning with LLMs-Judge Signal for Computer Use Agents
by: Luo, Man, et al.
Published: (2025) -
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)