Convex Optimization for Alignment and Preference Learning on a Single GPU
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Miria, Pilanci, Mert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IGT-OMD: Implicit Gradient Transport for Decision-Focused Learning under Delayed Feedback
von: Amoh, Benjamin, et al.
Veröffentlicht: (2026)
von: Amoh, Benjamin, et al.
Veröffentlicht: (2026)
A Reduction from Delayed to Immediate Feedback for Online Convex Optimization with Improved Guarantees
von: Ryabchenko, Alexander, et al.
Veröffentlicht: (2026)
von: Ryabchenko, Alexander, et al.
Veröffentlicht: (2026)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
von: Saxena, Udit
Veröffentlicht: (2025)
von: Saxena, Udit
Veröffentlicht: (2025)
J6: Jacobian-Driven Role Attribution for Multi-Objective Prompt Optimization in LLMs
von: Wu, Yao
Veröffentlicht: (2025)
von: Wu, Yao
Veröffentlicht: (2025)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)
Simplifying Hyperparameter Tuning in Online Machine Learning -- The spotRiverGUI
von: Bartz-Beielstein, Thomas
Veröffentlicht: (2024)
von: Bartz-Beielstein, Thomas
Veröffentlicht: (2024)
Bed-Attached Vibration Sensor System: A Machine Learning Approach for Fall Detection in Nursing Homes
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2024)
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2024)
Pseudoconvex Problems in Operational Decision Systems: Algorithms for Joint Learning and Optimization
von: Li, Zijun, et al.
Veröffentlicht: (2026)
von: Li, Zijun, et al.
Veröffentlicht: (2026)
Multi-Objective Optimization and Hyperparameter Tuning With Desirability Functions
von: Bartz-Beielstein, Thomas
Veröffentlicht: (2025)
von: Bartz-Beielstein, Thomas
Veröffentlicht: (2025)
$\texttt{skwdro}$: a library for Wasserstein distributionally robust machine learning
von: Vincent, Florian, et al.
Veröffentlicht: (2024)
von: Vincent, Florian, et al.
Veröffentlicht: (2024)
Multi-Objective Optimization with Desirability and Morris-Mitchell Criterion
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2025)
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Setwise Coordinate Descent for Dual Asynchronous Decentralized Optimization
von: Costantini, Marina, et al.
Veröffentlicht: (2025)
von: Costantini, Marina, et al.
Veröffentlicht: (2025)
Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
von: Vary, Simon, et al.
Veröffentlicht: (2024)
von: Vary, Simon, et al.
Veröffentlicht: (2024)
FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection
von: Singh, Devender, et al.
Veröffentlicht: (2026)
von: Singh, Devender, et al.
Veröffentlicht: (2026)
Decentralized Optimization with Topology-Independent Communication
von: Lin, Ying, et al.
Veröffentlicht: (2025)
von: Lin, Ying, et al.
Veröffentlicht: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
von: Easley, Eric, et al.
Veröffentlicht: (2026)
von: Easley, Eric, et al.
Veröffentlicht: (2026)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
von: Du, Wenzhang
Veröffentlicht: (2025)
von: Du, Wenzhang
Veröffentlicht: (2025)
Improving ML Training Data with Gold-Standard Quality Metrics
von: Barrett, Leslie, et al.
Veröffentlicht: (2025)
von: Barrett, Leslie, et al.
Veröffentlicht: (2025)
Towards Systematic Generalization for Power Grid Optimization Problems
von: Memon, Zeeshan, et al.
Veröffentlicht: (2026)
von: Memon, Zeeshan, et al.
Veröffentlicht: (2026)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
Stable and Convexified Information Bottleneck Optimization via Symbolic Continuation and Entropy-Regularized Trajectories
von: Alpay, Faruk
Veröffentlicht: (2025)
von: Alpay, Faruk
Veröffentlicht: (2025)
Faster, Cheaper, Better: Multi-Objective Hyperparameter Optimization for LLM and RAG Systems
von: Barker, Matthew, et al.
Veröffentlicht: (2025)
von: Barker, Matthew, et al.
Veröffentlicht: (2025)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
Training Language Models to Use Prolog as a Tool
von: Mellgren, Niklas, et al.
Veröffentlicht: (2025)
von: Mellgren, Niklas, et al.
Veröffentlicht: (2025)
Adaptive Candidate Point Thompson Sampling for High-Dimensional Bayesian Optimization
von: Fan, Donney, et al.
Veröffentlicht: (2026)
von: Fan, Donney, et al.
Veröffentlicht: (2026)
A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muo
von: Gratton, S., et al.
Veröffentlicht: (2026)
von: Gratton, S., et al.
Veröffentlicht: (2026)
Social Cooperation in Conversational AI Agents
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
Robustness, Cost, and Attack-Surface Concentration in Phishing Detection
von: Allagan, Julian, et al.
Veröffentlicht: (2026)
von: Allagan, Julian, et al.
Veröffentlicht: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
von: McCann, Jordan F.
Veröffentlicht: (2026)
von: McCann, Jordan F.
Veröffentlicht: (2026)
Super Apriel: One Checkpoint, Many Speeds
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
von: Ye, Hua, et al.
Veröffentlicht: (2025)
von: Ye, Hua, et al.
Veröffentlicht: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
QuAnTS: Question Answering on Time Series
von: Divo, Felix, et al.
Veröffentlicht: (2025)
von: Divo, Felix, et al.
Veröffentlicht: (2025)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
IGT-OMD: Implicit Gradient Transport for Decision-Focused Learning under Delayed Feedback
von: Amoh, Benjamin, et al.
Veröffentlicht: (2026) -
A Reduction from Delayed to Immediate Feedback for Online Convex Optimization with Improved Guarantees
von: Ryabchenko, Alexander, et al.
Veröffentlicht: (2026) -
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
von: Saxena, Udit
Veröffentlicht: (2025) -
J6: Jacobian-Driven Role Attribution for Multi-Objective Prompt Optimization in LLMs
von: Wu, Yao
Veröffentlicht: (2025) -
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)