$i$REPO: $i$mplicit Reward Pairwise Difference based Empirical Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Le, Long Tan, Shu, Han, Nguyen, Tung-Anh, Hong, Choong Seon, Tran, Nguyen H. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Federated PCA on Grassmann Manifold for IoT Anomaly Detection
von: Nguyen, Tung-Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Tung-Anh, et al.
Veröffentlicht: (2024)
Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
von: Le, Long Tan, et al.
Veröffentlicht: (2025)
von: Le, Long Tan, et al.
Veröffentlicht: (2025)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
CDKT-FL: Cross-Device Knowledge Transfer using Proxy Dataset in Federated Learning
von: Le, Huy Q., et al.
Veröffentlicht: (2022)
von: Le, Huy Q., et al.
Veröffentlicht: (2022)
iMoT: Inertial Motion Transformer for Inertial Navigation
von: Nguyen, Son Minh, et al.
Veröffentlicht: (2024)
von: Nguyen, Son Minh, et al.
Veröffentlicht: (2024)
Federated Deep Equilibrium Learning: Harnessing Compact Global Representations to Enhance Personalization
von: Le, Long Tan, et al.
Veröffentlicht: (2023)
von: Le, Long Tan, et al.
Veröffentlicht: (2023)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
von: Le, Huy, et al.
Veröffentlicht: (2023)
von: Le, Huy, et al.
Veröffentlicht: (2023)
Robust Federated Learning on Edge Devices with Domain Heterogeneity
von: Le, Huy Q., et al.
Veröffentlicht: (2025)
von: Le, Huy Q., et al.
Veröffentlicht: (2025)
InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting
von: Vu, Duc, et al.
Veröffentlicht: (2026)
von: Vu, Duc, et al.
Veröffentlicht: (2026)
Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation
von: Cao, Guining, et al.
Veröffentlicht: (2026)
von: Cao, Guining, et al.
Veröffentlicht: (2026)
C2L-Net: A Data-Driven Model for State-of-Charge Estimation of Lithium-Ion Batteries During Discharge
von: Tran, Khoa, et al.
Veröffentlicht: (2026)
von: Tran, Khoa, et al.
Veröffentlicht: (2026)
A Complete Survey on LLM-based AI Chatbots
von: Dam, Sumit Kumar, et al.
Veröffentlicht: (2024)
von: Dam, Sumit Kumar, et al.
Veröffentlicht: (2024)
MOTIF: Multi-strategy Optimization via Turn-based Interactive Framework
von: Kiet, Nguyen Viet Tuan, et al.
Veröffentlicht: (2025)
von: Kiet, Nguyen Viet Tuan, et al.
Veröffentlicht: (2025)
Resource-efficient Layer-wise Federated Self-supervised Learning
von: Tun, Ye Lin, et al.
Veröffentlicht: (2024)
von: Tun, Ye Lin, et al.
Veröffentlicht: (2024)
Anti-I2V: Safeguarding your photos from malicious image-to-video generation
von: Vu, Duc, et al.
Veröffentlicht: (2026)
von: Vu, Duc, et al.
Veröffentlicht: (2026)
Learning to Stop Overthinking at Test Time
von: Bao, Hieu Tran, et al.
Veröffentlicht: (2025)
von: Bao, Hieu Tran, et al.
Veröffentlicht: (2025)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
von: Minh, Nguyen Huu Nhat, et al.
Veröffentlicht: (2025)
von: Minh, Nguyen Huu Nhat, et al.
Veröffentlicht: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
FedDAP: Domain-Aware Prototype Learning for Federated Learning under Domain Shift
von: Le, Huy Q., et al.
Veröffentlicht: (2026)
von: Le, Huy Q., et al.
Veröffentlicht: (2026)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
von: Le, Huy, et al.
Veröffentlicht: (2025)
von: Le, Huy, et al.
Veröffentlicht: (2025)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey
von: Qiao, Yu, et al.
Veröffentlicht: (2025)
von: Qiao, Yu, et al.
Veröffentlicht: (2025)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2026)
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2026)
Mitigating Domain Shift in Federated Learning via Intra- and Inter-Domain Prototypes
von: Le, Huy Q., et al.
Veröffentlicht: (2025)
von: Le, Huy Q., et al.
Veröffentlicht: (2025)
Leveraging AI for Enhanced Software Effort Estimation: A Comprehensive Study and Framework Proposal
von: Tran, Nhi, et al.
Veröffentlicht: (2024)
von: Tran, Nhi, et al.
Veröffentlicht: (2024)
An Evaluation of LLMs Inference on Popular Single-board Computers
von: Tung, et al.
Veröffentlicht: (2025)
von: Tung, et al.
Veröffentlicht: (2025)
Pairwise Calibrated Rewards for Pluralistic Alignment
von: Halpern, Daniel, et al.
Veröffentlicht: (2025)
von: Halpern, Daniel, et al.
Veröffentlicht: (2025)
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
von: Tran, Quyen, et al.
Veröffentlicht: (2024)
von: Tran, Quyen, et al.
Veröffentlicht: (2024)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang-Binh, et al.
Veröffentlicht: (2025)
Fast Stochastic Greedy Algorithm for $k$-Submodular Cover Problem
von: Nguyen, Hue T., et al.
Veröffentlicht: (2025)
von: Nguyen, Hue T., et al.
Veröffentlicht: (2025)
Preference as Reward, Maximum Preference Optimization with Importance Sampling
von: Jiang, Zaifan, et al.
Veröffentlicht: (2023)
von: Jiang, Zaifan, et al.
Veröffentlicht: (2023)
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions
von: Sajadi, Amirali, et al.
Veröffentlicht: (2025)
von: Sajadi, Amirali, et al.
Veröffentlicht: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
LICO: Large Language Models for In-Context Molecular Optimization
von: Nguyen, Tung, et al.
Veröffentlicht: (2024)
von: Nguyen, Tung, et al.
Veröffentlicht: (2024)
Robust SDE Parameter Estimation Under Missing Time Information Setting
von: Van Tran, Long, et al.
Veröffentlicht: (2026)
von: Van Tran, Long, et al.
Veröffentlicht: (2026)
ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
von: Nguyen, Hong, et al.
Veröffentlicht: (2024)
von: Nguyen, Hong, et al.
Veröffentlicht: (2024)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
von: Jian, Ai, et al.
Veröffentlicht: (2025)
von: Jian, Ai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Federated PCA on Grassmann Manifold for IoT Anomaly Detection
von: Nguyen, Tung-Anh, et al.
Veröffentlicht: (2024) -
Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
von: Le, Long Tan, et al.
Veröffentlicht: (2025) -
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024) -
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025) -
CDKT-FL: Cross-Device Knowledge Transfer using Proxy Dataset in Federated Learning
von: Le, Huy Q., et al.
Veröffentlicht: (2022)