$i$REPO: $i$mplicit Reward Pairwise Difference based Empirical Preference Optimization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Le, Long Tan, Shu, Han, Nguyen, Tung-Anh, Hong, Choong Seon, Tran, Nguyen H. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Federated PCA on Grassmann Manifold for IoT Anomaly Detection
par: Nguyen, Tung-Anh, et autres
Publié: (2024)
par: Nguyen, Tung-Anh, et autres
Publié: (2024)
Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
par: Le, Long Tan, et autres
Publié: (2025)
par: Le, Long Tan, et autres
Publié: (2025)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
par: Wang, Shiqi, et autres
Publié: (2024)
par: Wang, Shiqi, et autres
Publié: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
par: Tran, Phuong, et autres
Publié: (2025)
par: Tran, Phuong, et autres
Publié: (2025)
CDKT-FL: Cross-Device Knowledge Transfer using Proxy Dataset in Federated Learning
par: Le, Huy Q., et autres
Publié: (2022)
par: Le, Huy Q., et autres
Publié: (2022)
iMoT: Inertial Motion Transformer for Inertial Navigation
par: Nguyen, Son Minh, et autres
Publié: (2024)
par: Nguyen, Son Minh, et autres
Publié: (2024)
Federated Deep Equilibrium Learning: Harnessing Compact Global Representations to Enhance Personalization
par: Le, Long Tan, et autres
Publié: (2023)
par: Le, Long Tan, et autres
Publié: (2023)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
par: Le, Huy, et autres
Publié: (2023)
par: Le, Huy, et autres
Publié: (2023)
Robust Federated Learning on Edge Devices with Domain Heterogeneity
par: Le, Huy Q., et autres
Publié: (2025)
par: Le, Huy Q., et autres
Publié: (2025)
InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting
par: Vu, Duc, et autres
Publié: (2026)
par: Vu, Duc, et autres
Publié: (2026)
Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation
par: Cao, Guining, et autres
Publié: (2026)
par: Cao, Guining, et autres
Publié: (2026)
C2L-Net: A Data-Driven Model for State-of-Charge Estimation of Lithium-Ion Batteries During Discharge
par: Tran, Khoa, et autres
Publié: (2026)
par: Tran, Khoa, et autres
Publié: (2026)
A Complete Survey on LLM-based AI Chatbots
par: Dam, Sumit Kumar, et autres
Publié: (2024)
par: Dam, Sumit Kumar, et autres
Publié: (2024)
MOTIF: Multi-strategy Optimization via Turn-based Interactive Framework
par: Kiet, Nguyen Viet Tuan, et autres
Publié: (2025)
par: Kiet, Nguyen Viet Tuan, et autres
Publié: (2025)
Resource-efficient Layer-wise Federated Self-supervised Learning
par: Tun, Ye Lin, et autres
Publié: (2024)
par: Tun, Ye Lin, et autres
Publié: (2024)
Anti-I2V: Safeguarding your photos from malicious image-to-video generation
par: Vu, Duc, et autres
Publié: (2026)
par: Vu, Duc, et autres
Publié: (2026)
Learning to Stop Overthinking at Test Time
par: Bao, Hieu Tran, et autres
Publié: (2025)
par: Bao, Hieu Tran, et autres
Publié: (2025)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
par: Nguyen, Truong, et autres
Publié: (2026)
par: Nguyen, Truong, et autres
Publié: (2026)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
par: Minh, Nguyen Huu Nhat, et autres
Publié: (2025)
par: Minh, Nguyen Huu Nhat, et autres
Publié: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
par: Tuong, Nguyen Anh, et autres
Publié: (2026)
par: Tuong, Nguyen Anh, et autres
Publié: (2026)
FedDAP: Domain-Aware Prototype Learning for Federated Learning under Domain Shift
par: Le, Huy Q., et autres
Publié: (2026)
par: Le, Huy Q., et autres
Publié: (2026)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
par: Le, Huy, et autres
Publié: (2025)
par: Le, Huy, et autres
Publié: (2025)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
par: Nguyen, Hieu Trung, et autres
Publié: (2026)
par: Nguyen, Hieu Trung, et autres
Publié: (2026)
DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey
par: Qiao, Yu, et autres
Publié: (2025)
par: Qiao, Yu, et autres
Publié: (2025)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
par: Nguyen, Toan, et autres
Publié: (2025)
par: Nguyen, Toan, et autres
Publié: (2025)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
par: Nguyen, Anh Thi-Hoang, et autres
Publié: (2026)
par: Nguyen, Anh Thi-Hoang, et autres
Publié: (2026)
Mitigating Domain Shift in Federated Learning via Intra- and Inter-Domain Prototypes
par: Le, Huy Q., et autres
Publié: (2025)
par: Le, Huy Q., et autres
Publié: (2025)
Leveraging AI for Enhanced Software Effort Estimation: A Comprehensive Study and Framework Proposal
par: Tran, Nhi, et autres
Publié: (2024)
par: Tran, Nhi, et autres
Publié: (2024)
An Evaluation of LLMs Inference on Popular Single-board Computers
par: Tung, et autres
Publié: (2025)
par: Tung, et autres
Publié: (2025)
Pairwise Calibrated Rewards for Pluralistic Alignment
par: Halpern, Daniel, et autres
Publié: (2025)
par: Halpern, Daniel, et autres
Publié: (2025)
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
par: Tran, Quyen, et autres
Publié: (2024)
par: Tran, Quyen, et autres
Publié: (2024)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
par: Nguyen, Quang-Binh, et autres
Publié: (2025)
par: Nguyen, Quang-Binh, et autres
Publié: (2025)
Fast Stochastic Greedy Algorithm for $k$-Submodular Cover Problem
par: Nguyen, Hue T., et autres
Publié: (2025)
par: Nguyen, Hue T., et autres
Publié: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
par: Jiang, Zaifan, et autres
Publié: (2023)
par: Jiang, Zaifan, et autres
Publié: (2023)
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
par: Sajadi, Amirali, et autres
Publié: (2025)
par: Sajadi, Amirali, et autres
Publié: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
par: Nguyen, Trong-Hieu, et autres
Publié: (2024)
par: Nguyen, Trong-Hieu, et autres
Publié: (2024)
LICO: Large Language Models for In-Context Molecular Optimization
par: Nguyen, Tung, et autres
Publié: (2024)
par: Nguyen, Tung, et autres
Publié: (2024)
Robust SDE Parameter Estimation Under Missing Time Information Setting
par: Van Tran, Long, et autres
Publié: (2026)
par: Van Tran, Long, et autres
Publié: (2026)
ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
par: Nguyen, Hong, et autres
Publié: (2024)
par: Nguyen, Hong, et autres
Publié: (2024)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
par: Jian, Ai, et autres
Publié: (2025)
par: Jian, Ai, et autres
Publié: (2025)
Documents similaires
-
Federated PCA on Grassmann Manifold for IoT Anomaly Detection
par: Nguyen, Tung-Anh, et autres
Publié: (2024) -
Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
par: Le, Long Tan, et autres
Publié: (2025) -
Reward Difference Optimization For Sample Reweighting In Offline RLHF
par: Wang, Shiqi, et autres
Publié: (2024) -
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
par: Tran, Phuong, et autres
Publié: (2025) -
CDKT-FL: Cross-Device Knowledge Transfer using Proxy Dataset in Federated Learning
par: Le, Huy Q., et autres
Publié: (2022)