RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Tianyu, Zhang, Haoye, Li, Qiming, Xu, Qixin, Yao, Yuan, Chen, Da, Lu, Xiaoman, Cui, Ganqu, Dang, Yunkai, He, Taiwen, Feng, Xiaocheng, Song, Jun, Zheng, Bo, Liu, Zhiyuan, Chua, Tat-Seng, Sun, Maosong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
RLPR: Extrapolating RLVR to General Domains without Verifiers
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
di: Jin, Zhe, et al.
Pubblicazione: (2025)
di: Jin, Zhe, et al.
Pubblicazione: (2025)
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
di: Yao, Yuan, et al.
Pubblicazione: (2024)
di: Yao, Yuan, et al.
Pubblicazione: (2024)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
di: Chen, Yiyang, et al.
Pubblicazione: (2022)
di: Chen, Yiyang, et al.
Pubblicazione: (2022)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
di: Beck, Jacob
Pubblicazione: (2025)
di: Beck, Jacob
Pubblicazione: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
di: Lee, Harrison, et al.
Pubblicazione: (2023)
di: Lee, Harrison, et al.
Pubblicazione: (2023)
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
di: Lin, Jiaye, et al.
Pubblicazione: (2025)
di: Lin, Jiaye, et al.
Pubblicazione: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
di: Chu, Meng, et al.
Pubblicazione: (2025)
di: Chu, Meng, et al.
Pubblicazione: (2025)
Universal Scene Graph Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
Learning to Ask Critical Questions for Assisting Product Search
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
di: Cui, Ganqu, et al.
Pubblicazione: (2023)
di: Cui, Ganqu, et al.
Pubblicazione: (2023)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
di: Yang, Qing, et al.
Pubblicazione: (2025)
di: Yang, Qing, et al.
Pubblicazione: (2025)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
di: Deng, Yang, et al.
Pubblicazione: (2023)
di: Deng, Yang, et al.
Pubblicazione: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
di: Zhang, An, et al.
Pubblicazione: (2024)
di: Zhang, An, et al.
Pubblicazione: (2024)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
di: Ma, Weijian, et al.
Pubblicazione: (2026)
di: Ma, Weijian, et al.
Pubblicazione: (2026)
Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V
di: Zhi, Peiyuan, et al.
Pubblicazione: (2024)
di: Zhi, Peiyuan, et al.
Pubblicazione: (2024)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
di: Xu, Ruyi, et al.
Pubblicazione: (2024)
di: Xu, Ruyi, et al.
Pubblicazione: (2024)
Length Controlled Generation for Black-box LLMs
di: Gu, Yuxuan, et al.
Pubblicazione: (2024)
di: Gu, Yuxuan, et al.
Pubblicazione: (2024)
Why Does RLAIF Work At All?
di: Young, Robin
Pubblicazione: (2026)
di: Young, Robin
Pubblicazione: (2026)
Can I Trust Your Answer? Visually Grounded Video Question Answering
di: Xiao, Junbin, et al.
Pubblicazione: (2023)
di: Xiao, Junbin, et al.
Pubblicazione: (2023)
Contrastive Pre-training for Deep Session Data Understanding
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
di: Fang, Xiang, et al.
Pubblicazione: (2026)
di: Fang, Xiang, et al.
Pubblicazione: (2026)
Extending Visual Dynamics for Video-to-Music Generation
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Enhancing Spectral Graph Neural Networks with LLM-Predicted Homophily
di: Lu, Kangkang, et al.
Pubblicazione: (2025)
di: Lu, Kangkang, et al.
Pubblicazione: (2025)
3D-TAFS: A Training-free Framework for 3D Affordance Segmentation
di: Chu, Meng, et al.
Pubblicazione: (2024)
di: Chu, Meng, et al.
Pubblicazione: (2024)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
di: Zhou, Zhenglin, et al.
Pubblicazione: (2025)
XNLP: An Interactive Demonstration System for Universal Structured NLP
di: Fei, Hao, et al.
Pubblicazione: (2023)
di: Fei, Hao, et al.
Pubblicazione: (2023)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
di: Guo, Shasha, et al.
Pubblicazione: (2024)
di: Guo, Shasha, et al.
Pubblicazione: (2024)
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
di: Liu, Zhiyuan, et al.
Pubblicazione: (2024)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2024)
Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules
di: Liu, Zhiyuan, et al.
Pubblicazione: (2023)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2023)
NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
di: Dai, Sunhao, et al.
Pubblicazione: (2025)
di: Dai, Sunhao, et al.
Pubblicazione: (2025)
Inverting the wedge map and Gauss composition
di: Chua, Kok Seng
Pubblicazione: (2024)
di: Chua, Kok Seng
Pubblicazione: (2024)
Chebyshev polynomials and a refinement of the local residue/non-residue structure at a prime
di: Chua, Kok Seng
Pubblicazione: (2026)
di: Chua, Kok Seng
Pubblicazione: (2026)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
di: Shen, Fei, et al.
Pubblicazione: (2025)
di: Shen, Fei, et al.
Pubblicazione: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
di: Qi, Ji, et al.
Pubblicazione: (2025)
di: Qi, Ji, et al.
Pubblicazione: (2025)
Principled Multimodal Representation Learning
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
On Generative Agents in Recommendation
di: Zhang, An, et al.
Pubblicazione: (2023)
di: Zhang, An, et al.
Pubblicazione: (2023)
LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation
di: He, Yingzhi, et al.
Pubblicazione: (2025)
di: He, Yingzhi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
di: Yu, Tianyu, et al.
Pubblicazione: (2023) -
RLPR: Extrapolating RLVR to General Domains without Verifiers
di: Yu, Tianyu, et al.
Pubblicazione: (2025) -
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
di: Jin, Zhe, et al.
Pubblicazione: (2025) -
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023) -
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
di: Yao, Yuan, et al.
Pubblicazione: (2024)