RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Tianyu, Zhang, Haoye, Li, Qiming, Xu, Qixin, Yao, Yuan, Chen, Da, Lu, Xiaoman, Cui, Ganqu, Dang, Yunkai, He, Taiwen, Feng, Xiaocheng, Song, Jun, Zheng, Bo, Liu, Zhiyuan, Chua, Tat-Seng, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
von: Yu, Tianyu, et al.
Veröffentlicht: (2023)
von: Yu, Tianyu, et al.
Veröffentlicht: (2023)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
NExT-GPT: Any-to-Any Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023)
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
von: Yao, Yuan, et al.
Veröffentlicht: (2024)
von: Yao, Yuan, et al.
Veröffentlicht: (2024)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
von: Beck, Jacob
Veröffentlicht: (2025)
von: Beck, Jacob
Veröffentlicht: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
von: Lin, Jiaye, et al.
Veröffentlicht: (2025)
von: Lin, Jiaye, et al.
Veröffentlicht: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Learning to Ask Critical Questions for Assisting Product Search
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
von: Cui, Ganqu, et al.
Veröffentlicht: (2023)
von: Cui, Ganqu, et al.
Veröffentlicht: (2023)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
von: Yang, Qing, et al.
Veröffentlicht: (2025)
von: Yang, Qing, et al.
Veröffentlicht: (2025)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
von: Deng, Yang, et al.
Veröffentlicht: (2023)
von: Deng, Yang, et al.
Veröffentlicht: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
von: Zhang, An, et al.
Veröffentlicht: (2024)
von: Zhang, An, et al.
Veröffentlicht: (2024)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
von: Ma, Weijian, et al.
Veröffentlicht: (2026)
von: Ma, Weijian, et al.
Veröffentlicht: (2026)
Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V
von: Zhi, Peiyuan, et al.
Veröffentlicht: (2024)
von: Zhi, Peiyuan, et al.
Veröffentlicht: (2024)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
von: Xu, Ruyi, et al.
Veröffentlicht: (2024)
von: Xu, Ruyi, et al.
Veröffentlicht: (2024)
Length Controlled Generation for Black-box LLMs
von: Gu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Gu, Yuxuan, et al.
Veröffentlicht: (2024)
Why Does RLAIF Work At All?
von: Young, Robin
Veröffentlicht: (2026)
von: Young, Robin
Veröffentlicht: (2026)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
Contrastive Pre-training for Deep Session Data Understanding
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Extending Visual Dynamics for Video-to-Music Generation
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Enhancing Spectral Graph Neural Networks with LLM-Predicted Homophily
von: Lu, Kangkang, et al.
Veröffentlicht: (2025)
von: Lu, Kangkang, et al.
Veröffentlicht: (2025)
3D-TAFS: A Training-free Framework for 3D Affordance Segmentation
von: Chu, Meng, et al.
Veröffentlicht: (2024)
von: Chu, Meng, et al.
Veröffentlicht: (2024)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
XNLP: An Interactive Demonstration System for Universal Structured NLP
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
von: Guo, Shasha, et al.
Veröffentlicht: (2024)
von: Guo, Shasha, et al.
Veröffentlicht: (2024)
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2024)
Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2023)
NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
von: Dai, Sunhao, et al.
Veröffentlicht: (2025)
von: Dai, Sunhao, et al.
Veröffentlicht: (2025)
Inverting the wedge map and Gauss composition
von: Chua, Kok Seng
Veröffentlicht: (2024)
von: Chua, Kok Seng
Veröffentlicht: (2024)
Chebyshev polynomials and a refinement of the local residue/non-residue structure at a prime
von: Chua, Kok Seng
Veröffentlicht: (2026)
von: Chua, Kok Seng
Veröffentlicht: (2026)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
von: Shen, Fei, et al.
Veröffentlicht: (2025)
von: Shen, Fei, et al.
Veröffentlicht: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
von: Qi, Ji, et al.
Veröffentlicht: (2025)
von: Qi, Ji, et al.
Veröffentlicht: (2025)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
On Generative Agents in Recommendation
von: Zhang, An, et al.
Veröffentlicht: (2023)
von: Zhang, An, et al.
Veröffentlicht: (2023)
LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation
von: He, Yingzhi, et al.
Veröffentlicht: (2025)
von: He, Yingzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
von: Yu, Tianyu, et al.
Veröffentlicht: (2023) -
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025) -
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025) -
NExT-GPT: Any-to-Any Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023) -
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
von: Yao, Yuan, et al.
Veröffentlicht: (2024)