Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models through Reinforcement Learning from Ranking Feedback
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Derek, Glatt, Ruben, Klymko, Christine, Mohole, Shubham, Choi, Hongjun, Kushwaha, Shashank, Sakla, Sam, da Silva, Felipe Leno |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VERIRAG: A Post-Retrieval Auditing of Scientific Study Summaries
por: Mohole, Shubham, et al.
Publicado: (2025)
por: Mohole, Shubham, et al.
Publicado: (2025)
SIFOTL: A Principled, Statistically-Informed Fidelity-Optimization Method for Tabular Learning
por: Mohole, Shubham, et al.
Publicado: (2025)
por: Mohole, Shubham, et al.
Publicado: (2025)
VeriMinder: Mitigating Analytical Vulnerabilities in NL2SQL
por: Mohole, Shubham, et al.
Publicado: (2025)
por: Mohole, Shubham, et al.
Publicado: (2025)
Enhancing Accuracy and Parameter-Efficiency of Neural Representations for Network Parameterization
por: Choi, Hongjun, et al.
Publicado: (2024)
por: Choi, Hongjun, et al.
Publicado: (2024)
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
por: Lin, Jiaye, et al.
Publicado: (2025)
por: Lin, Jiaye, et al.
Publicado: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
por: Lee, Harrison, et al.
Publicado: (2023)
por: Lee, Harrison, et al.
Publicado: (2023)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
por: Beck, Jacob
Publicado: (2025)
por: Beck, Jacob
Publicado: (2025)
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
por: Ahn, Daechul, et al.
Publicado: (2024)
por: Ahn, Daechul, et al.
Publicado: (2024)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
por: Yang, Qing, et al.
Publicado: (2025)
por: Yang, Qing, et al.
Publicado: (2025)
Oracle modalities
por: Swan, Andrew W
Publicado: (2024)
por: Swan, Andrew W
Publicado: (2024)
Safe, Efficient, and Robust Reinforcement Learning for Ranking and Diffusion Models
por: Gupta, Shashank
Publicado: (2025)
por: Gupta, Shashank
Publicado: (2025)
RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
por: Yu, Tianyu, et al.
Publicado: (2024)
por: Yu, Tianyu, et al.
Publicado: (2024)
Why Does RLAIF Work At All?
por: Young, Robin
Publicado: (2026)
por: Young, Robin
Publicado: (2026)
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
por: Hengle, Amey, et al.
Publicado: (2024)
por: Hengle, Amey, et al.
Publicado: (2024)
Sparse Autoencoders as a Steering Basis for Phase Synchronization in Graph-Based CFD Surrogates
por: Hu, Yeping, et al.
Publicado: (2026)
por: Hu, Yeping, et al.
Publicado: (2026)
Improving Robustness In Sparse Autoencoders via Masked Regularization
por: Narayanaswamy, Vivek, et al.
Publicado: (2026)
por: Narayanaswamy, Vivek, et al.
Publicado: (2026)
Learning nuclear cross sections across the chart of nuclides with graph neural networks
por: Choi, Hongjun, et al.
Publicado: (2024)
por: Choi, Hongjun, et al.
Publicado: (2024)
Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles
por: Tang, Zhiwei, et al.
Publicado: (2023)
por: Tang, Zhiwei, et al.
Publicado: (2023)
Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
por: Chen, Jingyi, et al.
Publicado: (2025)
por: Chen, Jingyi, et al.
Publicado: (2025)
Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning
por: Ye, Kai, et al.
Publicado: (2025)
por: Ye, Kai, et al.
Publicado: (2025)
Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
por: Liu, Boyin, et al.
Publicado: (2025)
por: Liu, Boyin, et al.
Publicado: (2025)
Applying RLAIF for Code Generation with API-usage in Lightweight LLMs
por: Dutta, Sujan, et al.
Publicado: (2024)
por: Dutta, Sujan, et al.
Publicado: (2024)
Enhancing the Aesthetic Appeal of AI-Generated Physical Product Designs through LoRA Fine-Tuning with Human Feedback
por: Liao, Dinuo, et al.
Publicado: (2025)
por: Liao, Dinuo, et al.
Publicado: (2025)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
por: Wang, Qi, et al.
Publicado: (2025)
por: Wang, Qi, et al.
Publicado: (2025)
Oracle Bone Inscriptions Multi-modal Dataset
por: Li, Bang, et al.
Publicado: (2024)
por: Li, Bang, et al.
Publicado: (2024)
TeleOracle: Fine-Tuned Retrieval-Augmented Generation with Long-Context Support for Network
por: Alabbasi, Nouf, et al.
Publicado: (2024)
por: Alabbasi, Nouf, et al.
Publicado: (2024)
TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
por: Jana, Prithwish, et al.
Publicado: (2026)
por: Jana, Prithwish, et al.
Publicado: (2026)
GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning
por: Bao, Xiaoyi, et al.
Publicado: (2025)
por: Bao, Xiaoyi, et al.
Publicado: (2025)
Reinforcement Learning without Human Feedback for Last Mile Fine-Tuning of Large Language Models
por: Solway, Alec
Publicado: (2024)
por: Solway, Alec
Publicado: (2024)
Attractor Patch Networks: Reducing Catastrophic Forgetting with Routed Low-Rank Patch Experts
por: Shashank
Publicado: (2026)
por: Shashank
Publicado: (2026)
The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation
por: Pundhir, Shubham, et al.
Publicado: (2025)
por: Pundhir, Shubham, et al.
Publicado: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
por: CH-Wang, Sky, et al.
Publicado: (2025)
por: CH-Wang, Sky, et al.
Publicado: (2025)
Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation
por: Yang, Guangjing, et al.
Publicado: (2026)
por: Yang, Guangjing, et al.
Publicado: (2026)
IMPORTANCIA DE UN DIAGNÓSTICO PRECOZ Y CUIDADOS DE ENFERMERÍA EN DIABETES GESTACIONAL.
por: D. Leno González
Publicado: (2005)
por: D. Leno González
Publicado: (2005)
América latina, o discurso filosófico-sociológico da modernidade, a ce-gueira histórico-sociológica das teorias da modernidade: notas programá-ticas para uma práxis decolonial latino-americana
por: Leno Francisco Danner
Publicado: (2018)
por: Leno Francisco Danner
Publicado: (2018)
Um mundo sem mediações: descolonização africana como teoria política da modernização periférica
por: Leno Francisco Danner
Publicado: (2022)
por: Leno Francisco Danner
Publicado: (2022)
Pacificando o branco: uma história da modernidade contada pelos indígenas
por: Leno Francisco Danner
Publicado: (2022)
por: Leno Francisco Danner
Publicado: (2022)
Educação, resistência e politização: sobre o sentido da educação na literatura indígena brasileira contemporânea
por: Leno Francisco Danner
Publicado: (2020)
por: Leno Francisco Danner
Publicado: (2020)
Um xamã yanomami frente ao discurso filosófico-sociológico da modernidade
por: Leno Francisco Danner
Publicado: (2018)
por: Leno Francisco Danner
Publicado: (2018)
Estado, política e evolução social: uma tendência para este século XXI
por: Leno Francisco Danner
Publicado: (2017)
por: Leno Francisco Danner
Publicado: (2017)
Ejemplares similares
-
VERIRAG: A Post-Retrieval Auditing of Scientific Study Summaries
por: Mohole, Shubham, et al.
Publicado: (2025) -
SIFOTL: A Principled, Statistically-Informed Fidelity-Optimization Method for Tabular Learning
por: Mohole, Shubham, et al.
Publicado: (2025) -
VeriMinder: Mitigating Analytical Vulnerabilities in NL2SQL
por: Mohole, Shubham, et al.
Publicado: (2025) -
Enhancing Accuracy and Parameter-Efficiency of Neural Representations for Network Parameterization
por: Choi, Hongjun, et al.
Publicado: (2024) -
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
por: Lin, Jiaye, et al.
Publicado: (2025)