Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Jiaming, Zhou, Jiayi, Lou, Hantao, Chen, Boyuan, Hong, Donghai, Wang, Xuyao, Chen, Wenqi, Wang, Kaile, Pan, Rui, Li, Jiahao, Wang, Mohan, Dai, Josef, Qiu, Tianyi, Xu, Hua, Li, Dong, Chen, Weipeng, Song, Jun, Zheng, Bo, Yang, Yaodong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910766060273664
author Ji, Jiaming
Zhou, Jiayi
Lou, Hantao
Chen, Boyuan
Hong, Donghai
Wang, Xuyao
Chen, Wenqi
Wang, Kaile
Pan, Rui
Li, Jiahao
Wang, Mohan
Dai, Josef
Qiu, Tianyi
Xu, Hua
Li, Dong
Chen, Weipeng
Song, Jun
Zheng, Bo
Yang, Yaodong
author_facet Ji, Jiaming
Zhou, Jiayi
Lou, Hantao
Chen, Boyuan
Hong, Donghai
Wang, Xuyao
Chen, Wenqi
Wang, Kaile
Pan, Rui
Li, Jiahao
Wang, Mohan
Dai, Josef
Qiu, Tianyi
Xu, Hua
Li, Dong
Chen, Weipeng
Song, Jun
Zheng, Bo
Yang, Yaodong
contents Reinforcement learning from human feedback (RLHF) has proven effective in enhancing the instruction-following capabilities of large language models; however, it remains underexplored in the cross-modality domain. As the number of modalities increases, aligning all-modality models with human intentions -- such as instruction following -- becomes a pressing challenge. In this work, we make the first attempt to fine-tune all-modality models (i.e. input and output with any modality, also named any-to-any models) using human preference data across all modalities (including text, image, audio, and video), ensuring its behavior aligns with human intentions. This endeavor presents several challenges. First, there is no large-scale all-modality human preference data in existing open-source resources, as most datasets are limited to specific modalities, predominantly text and image. Secondly, the effectiveness of binary preferences in RLHF for post-training alignment in complex all-modality scenarios remains an unexplored area. Finally, there is a lack of a systematic framework to evaluate the capabilities of all-modality models, particularly regarding modality selection and synergy. To address these challenges, we propose the align-anything framework, which includes meticulously annotated 200k all-modality human preference data. Then, we introduce an alignment method that learns from unified language feedback, effectively capturing complex modality-specific human preferences and enhancing the model's instruction-following capabilities. Furthermore, to assess performance improvements in all-modality models after post-training alignment, we construct a challenging all-modality capability evaluation framework -- eval-anything. All data, models, and code frameworks have been open-sourced for the community. For more details, please refer to https://github.com/PKU-Alignment/align-anything.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Ji, Jiaming
Zhou, Jiayi
Lou, Hantao
Chen, Boyuan
Hong, Donghai
Wang, Xuyao
Chen, Wenqi
Wang, Kaile
Pan, Rui
Li, Jiahao
Wang, Mohan
Dai, Josef
Qiu, Tianyi
Xu, Hua
Li, Dong
Chen, Weipeng
Song, Jun
Zheng, Bo
Yang, Yaodong
Artificial Intelligence
Computation and Language
Reinforcement learning from human feedback (RLHF) has proven effective in enhancing the instruction-following capabilities of large language models; however, it remains underexplored in the cross-modality domain. As the number of modalities increases, aligning all-modality models with human intentions -- such as instruction following -- becomes a pressing challenge. In this work, we make the first attempt to fine-tune all-modality models (i.e. input and output with any modality, also named any-to-any models) using human preference data across all modalities (including text, image, audio, and video), ensuring its behavior aligns with human intentions. This endeavor presents several challenges. First, there is no large-scale all-modality human preference data in existing open-source resources, as most datasets are limited to specific modalities, predominantly text and image. Secondly, the effectiveness of binary preferences in RLHF for post-training alignment in complex all-modality scenarios remains an unexplored area. Finally, there is a lack of a systematic framework to evaluate the capabilities of all-modality models, particularly regarding modality selection and synergy. To address these challenges, we propose the align-anything framework, which includes meticulously annotated 200k all-modality human preference data. Then, we introduce an alignment method that learns from unified language feedback, effectively capturing complex modality-specific human preferences and enhancing the model's instruction-following capabilities. Furthermore, to assess performance improvements in all-modality models after post-training alignment, we construct a challenging all-modality capability evaluation framework -- eval-anything. All data, models, and code frameworks have been open-sourced for the community. For more details, please refer to https://github.com/PKU-Alignment/align-anything.
title Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.15838