Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Giyeong, Lee, Junghyun, Park, Jaehyun, Yu, Youngjae, Bae, Wonho, Noh, Junhyug |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalized Coverage for More Robust Low-Budget Active Learning
by: Bae, Wonho, et al.
Published: (2024)
by: Bae, Wonho, et al.
Published: (2024)
Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
by: Kim, Jeongin, et al.
Published: (2025)
by: Kim, Jeongin, et al.
Published: (2025)
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
by: Lee, Junghyun, et al.
Published: (2026)
by: Lee, Junghyun, et al.
Published: (2026)
Active Learning for Continual Learning: Keeping the Past Alive in the Present
by: Park, Jaehyun, et al.
Published: (2025)
by: Park, Jaehyun, et al.
Published: (2025)
Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
HardNet: Hard-Constrained Neural Networks with Universal Approximation Guarantees
by: Min, Youngjae, et al.
Published: (2024)
by: Min, Youngjae, et al.
Published: (2024)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
by: Oh, Yoojin, et al.
Published: (2025)
by: Oh, Yoojin, et al.
Published: (2025)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Performance Metric for Multiple Anomaly Score Distributions with Discrete Severity Levels
by: Yi, Wonjun, et al.
Published: (2024)
by: Yi, Wonjun, et al.
Published: (2024)
LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
by: Park, Taekhyun, et al.
Published: (2026)
by: Park, Taekhyun, et al.
Published: (2026)
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Querying Easily Flip-flopped Samples for Deep Active Learning
by: Cho, Seong Jin, et al.
Published: (2024)
by: Cho, Seong Jin, et al.
Published: (2024)
SCONE: A Novel Stochastic Sampling to Generate Contrastive Views and Hard Negative Samples for Recommendation
by: Lee, Chaejeong, et al.
Published: (2024)
by: Lee, Chaejeong, et al.
Published: (2024)
Tabular Transfer Learning via Prompting LLMs
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
Retrieval-Retro: Retrieval-based Inorganic Retrosynthesis with Expert Knowledge
by: Noh, Heewoong, et al.
Published: (2024)
by: Noh, Heewoong, et al.
Published: (2024)
TLDR: Unsupervised Goal-Conditioned RL via Temporal Distance-Aware Representations
by: Bae, Junik, et al.
Published: (2024)
by: Bae, Junik, et al.
Published: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
by: Zhang, Zhengze, et al.
Published: (2025)
by: Zhang, Zhengze, et al.
Published: (2025)
3D Interaction Geometric Pre-training for Molecular Relational Learning
by: Lee, Namkyeong, et al.
Published: (2024)
by: Lee, Namkyeong, et al.
Published: (2024)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks
by: Oh, Giyeong, et al.
Published: (2025)
by: Oh, Giyeong, et al.
Published: (2025)
Towards Visual Text Design Transfer Across Languages
by: Choi, Yejin, et al.
Published: (2024)
by: Choi, Yejin, et al.
Published: (2024)
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
Active Attacks: Red-teaming LLMs via Adaptive Environments
by: Yun, Taeyoung, et al.
Published: (2025)
by: Yun, Taeyoung, et al.
Published: (2025)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
SP^2DPO: An LLM-assisted Semantic Per-Pair DPO Generalization
by: He, Chaoyue, et al.
Published: (2026)
by: He, Chaoyue, et al.
Published: (2026)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Domain-Adaptive Health Indicator Learning with Degradation-Stage Synchronized Sampling and Cross-Domain Autoencoder
by: Choo, Jungho, et al.
Published: (2026)
by: Choo, Jungho, et al.
Published: (2026)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
by: Pollanen, Marco
Published: (2026)
by: Pollanen, Marco
Published: (2026)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
Variance Alignment Score: A Simple But Tough-to-Beat Data Selection Method for Multimodal Contrastive Learning
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Beating Transformers using Synthetic Cognition
by: Ibias, Alfredo, et al.
Published: (2025)
by: Ibias, Alfredo, et al.
Published: (2025)
Is Prompt Selection Necessary for Task-Free Online Continual Learning?
by: Park, Seoyoung, et al.
Published: (2026)
by: Park, Seoyoung, et al.
Published: (2026)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Similar Items
-
Generalized Coverage for More Robust Low-Budget Active Learning
by: Bae, Wonho, et al.
Published: (2024) -
Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
by: Kim, Jeongin, et al.
Published: (2025) -
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
by: Lee, Junghyun, et al.
Published: (2026) -
Active Learning for Continual Learning: Keeping the Past Alive in the Present
by: Park, Jaehyun, et al.
Published: (2025) -
Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
by: Wang, Jing, et al.
Published: (2025)