CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Hao, Zhao, Zhuokai, Yan, Shen, Korycki, Lukasz, Wang, Jianyu, He, Baosheng, Liu, Jiayi, Zhang, Lizhu, Fan, Xiangjun, Yu, Hanchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909552166830080
author Yu, Hao
Zhao, Zhuokai
Yan, Shen
Korycki, Lukasz
Wang, Jianyu
He, Baosheng
Liu, Jiayi
Zhang, Lizhu
Fan, Xiangjun
Yu, Hanchao
author_facet Yu, Hao
Zhao, Zhuokai
Yan, Shen
Korycki, Lukasz
Wang, Jianyu
He, Baosheng
Liu, Jiayi
Zhang, Lizhu
Fan, Xiangjun
Yu, Hanchao
contents The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across both visual and textual domains. While excelling in generative tasks, existing LVLMs often face limitations in tasks requiring high-fidelity representation learning, such as generating image or text embeddings for retrieval. Recent work has proposed finetuning LVLMs for representational learning, but the fine-tuned model often loses its generative capabilities due to the representational learning training paradigm. To address this trade-off, we introduce CAFe, a contrastive-autoregressive fine-tuning framework that enhances LVLMs for both representation and generative tasks. By integrating a contrastive objective with autoregressive language modeling, our approach unifies these traditionally separate tasks, achieving state-of-the-art results in both multimodal retrieval and multimodal generative benchmarks, including object hallucination (OH) mitigation. CAFe establishes a novel framework that synergizes embedding and generative functionalities in a single model, setting a foundation for future multimodal models that excel in both retrieval precision and coherent output generation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
Yu, Hao
Zhao, Zhuokai
Yan, Shen
Korycki, Lukasz
Wang, Jianyu
He, Baosheng
Liu, Jiayi
Zhang, Lizhu
Fan, Xiangjun
Yu, Hanchao
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The rapid advancement of large vision-language models (LVLMs) has driven significant progress in multimodal tasks, enabling models to interpret, reason, and generate outputs across both visual and textual domains. While excelling in generative tasks, existing LVLMs often face limitations in tasks requiring high-fidelity representation learning, such as generating image or text embeddings for retrieval. Recent work has proposed finetuning LVLMs for representational learning, but the fine-tuned model often loses its generative capabilities due to the representational learning training paradigm. To address this trade-off, we introduce CAFe, a contrastive-autoregressive fine-tuning framework that enhances LVLMs for both representation and generative tasks. By integrating a contrastive objective with autoregressive language modeling, our approach unifies these traditionally separate tasks, achieving state-of-the-art results in both multimodal retrieval and multimodal generative benchmarks, including object hallucination (OH) mitigation. CAFe establishes a novel framework that synergizes embedding and generative functionalities in a single model, setting a foundation for future multimodal models that excel in both retrieval precision and coherent output generation.
title CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.19900