Learning from Preferences and Mixed Demonstrations in General Settings
Fuente:
arXiv
Saved in:
| Main Authors: | Brown, Jason R, Ek, Carl Henrik, Mullins, Robert D |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
by: Brown, Jason R, et al.
Published: (2025)
by: Brown, Jason R, et al.
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
Normalisation and Initialisation Strategies for Graph Neural Networks in Blockchain Anomaly Detection
by: Duy, Dang Sy, et al.
Published: (2026)
by: Duy, Dang Sy, et al.
Published: (2026)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025)
by: Amini, Ali
Published: (2025)
LUCAS-MEGA: A Large-Scale Multimodal Dataset for Representation Learning in Soil-Environment Systems
by: Leng, Kuangdai, et al.
Published: (2026)
by: Leng, Kuangdai, et al.
Published: (2026)
Multimodal Generative AI for Story Point Estimation in Software Development
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
Effective High-order Graph Representation Learning for Credit Card Fraud Detection
by: Zou, Yao, et al.
Published: (2025)
by: Zou, Yao, et al.
Published: (2025)
Recent Advances in Data-Driven Business Process Management
by: Ackermann, Lars, et al.
Published: (2024)
by: Ackermann, Lars, et al.
Published: (2024)
When Does Global Attention Help? A Unified Empirical Study on Atomistic Graph Learning
by: Chowdhury, Arindam, et al.
Published: (2025)
by: Chowdhury, Arindam, et al.
Published: (2025)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
by: Akella, Aditya
Published: (2025)
by: Akella, Aditya
Published: (2025)
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery
by: Su, Jiarui, et al.
Published: (2026)
by: Su, Jiarui, et al.
Published: (2026)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026)
by: Klačan, Ján, et al.
Published: (2026)
Analyzing the Impact of Release Season and Production Budget on Movie Revenue and Profitability
by: Torkamani, Mohammad Jalili, et al.
Published: (2026)
by: Torkamani, Mohammad Jalili, et al.
Published: (2026)
iLTM: Integrated Large Tabular Model
by: Bonet, David, et al.
Published: (2025)
by: Bonet, David, et al.
Published: (2025)
Evolving machine learning workflows through interactive AutoML
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Modeling Hypergraph Using Large Language Models
by: Gu, Bingqiao, et al.
Published: (2025)
by: Gu, Bingqiao, et al.
Published: (2025)
Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting
by: Phungtua-eng, Thanapol, et al.
Published: (2026)
by: Phungtua-eng, Thanapol, et al.
Published: (2026)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
Proving Olympiad Algebraic Inequalities without Human Demonstrations
by: Wei, Chenrui, et al.
Published: (2024)
by: Wei, Chenrui, et al.
Published: (2024)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Pattern Recognition Tasks with Personalized Federated Learning
by: Rahman, Md. Arifur, et al.
Published: (2026)
by: Rahman, Md. Arifur, et al.
Published: (2026)
AI Agents: Evolution, Architecture, and Real-World Applications
by: Krishnan, Naveen
Published: (2025)
by: Krishnan, Naveen
Published: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
by: Shekar, Pavan C, et al.
Published: (2025)
by: Shekar, Pavan C, et al.
Published: (2025)
UWM-JEPA: Predictive World Models That Imagine in Belief Space
by: Radha, Santosh Kumar, et al.
Published: (2026)
by: Radha, Santosh Kumar, et al.
Published: (2026)
Unified Interaction Foundational Model (UIFM) for Predicting Complex User and System Behavior
by: Ethiraj, Vignesh, et al.
Published: (2025)
by: Ethiraj, Vignesh, et al.
Published: (2025)
Generative AI and the Transformation of Software Development Practices
by: Acharya, Vivek
Published: (2025)
by: Acharya, Vivek
Published: (2025)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
by: Raman, Vishal, et al.
Published: (2025)
by: Raman, Vishal, et al.
Published: (2025)
Simple yet Effective Node Property Prediction on Edge Streams under Distribution Shifts
by: Lee, Jongha, et al.
Published: (2025)
by: Lee, Jongha, et al.
Published: (2025)
COMET: Codebook-based Online-adaptive Multi-scale Embedding for Time-series Anomaly Detection
by: Park, Jinwoo, et al.
Published: (2026)
by: Park, Jinwoo, et al.
Published: (2026)
Similar Items
-
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
by: Brown, Jason R, et al.
Published: (2025) -
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026) -
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026) -
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025) -
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)