The Trajectory Alignment Coefficient in Two Acts: From Reward Tuning to Reward Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Muslimani, Calarina, Du, Yunshu, Kawamoto, Kenta, Subramanian, Kaushik, Stone, Peter, Wurman, Peter |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Super-human Vision-based Reinforcement Learning Agent for Autonomous Racing in Gran Turismo
di: Vasco, Miguel, et al.
Pubblicazione: (2024)
di: Vasco, Miguel, et al.
Pubblicazione: (2024)
Reward Learning through Ranking Mean Squared Error
di: Kharyal, Chaitanya, et al.
Pubblicazione: (2026)
di: Kharyal, Chaitanya, et al.
Pubblicazione: (2026)
GFlowState: Visualizing the Training of Generative Flow Networks Beyond the Reward
di: Holeczek, Florian, et al.
Pubblicazione: (2026)
di: Holeczek, Florian, et al.
Pubblicazione: (2026)
Hindsight PRIORs for Reward Learning from Human Preferences
di: Verma, Mudit, et al.
Pubblicazione: (2024)
di: Verma, Mudit, et al.
Pubblicazione: (2024)
Training AI Co-Scientists Using Rubric Rewards
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
Learning to Assist Humans without Inferring Rewards
di: Myers, Vivek, et al.
Pubblicazione: (2024)
di: Myers, Vivek, et al.
Pubblicazione: (2024)
Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents
di: Septon, Yael, et al.
Pubblicazione: (2022)
di: Septon, Yael, et al.
Pubblicazione: (2022)
Crowd-PrefRL: Preference-Based Reward Learning from Crowds
di: Chhan, David, et al.
Pubblicazione: (2024)
di: Chhan, David, et al.
Pubblicazione: (2024)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
di: Dai, Xiangxiang, et al.
Pubblicazione: (2024)
Explanation through Reward Model Reconciliation using POMDP Tree Search
di: Kraske, Benjamin D., et al.
Pubblicazione: (2023)
di: Kraske, Benjamin D., et al.
Pubblicazione: (2023)
Interaction Dynamics as a Reward Signal for LLMs
di: Gooding, Sian, et al.
Pubblicazione: (2025)
di: Gooding, Sian, et al.
Pubblicazione: (2025)
A Champion-level Vision-based Reinforcement Learning Agent for Competitive Racing in Gran Turismo 7
di: Lee, Hojoon, et al.
Pubblicazione: (2025)
di: Lee, Hojoon, et al.
Pubblicazione: (2025)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
di: Muslimani, Calarina, et al.
Pubblicazione: (2025)
di: Muslimani, Calarina, et al.
Pubblicazione: (2025)
Unsupervised Reward-Driven Image Segmentation in Automated Scanning Transmission Electron Microscopy Experiments
di: Barakati, Kamyar, et al.
Pubblicazione: (2024)
di: Barakati, Kamyar, et al.
Pubblicazione: (2024)
A Koopman-Bayesian Framework for High-Fidelity, Perceptually Optimized Haptic Surgical Simulation
di: Kaushik, Rohit, et al.
Pubblicazione: (2026)
di: Kaushik, Rohit, et al.
Pubblicazione: (2026)
HappyRouting: Learning Emotion-Aware Route Trajectories for Scalable In-The-Wild Navigation
di: Bethge, David, et al.
Pubblicazione: (2024)
di: Bethge, David, et al.
Pubblicazione: (2024)
Hierarchical Reward Design from Language: Enhancing Alignment of Agent Behavior with Human Specifications
di: Qian, Zhiqin, et al.
Pubblicazione: (2026)
di: Qian, Zhiqin, et al.
Pubblicazione: (2026)
Influencing Humans to Conform to Preference Models for RLHF
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
di: Merker, Helena, et al.
Pubblicazione: (2026)
di: Merker, Helena, et al.
Pubblicazione: (2026)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
di: Dennler, Nathaniel, et al.
Pubblicazione: (2024)
Reward driven workflows for unsupervised explainable analysis of phases and ferroic variants from atomically resolved imaging data
di: Barakati, Kamyar, et al.
Pubblicazione: (2024)
di: Barakati, Kamyar, et al.
Pubblicazione: (2024)
Revisiting Euclidean Alignment for Transfer Learning in EEG-Based Brain-Computer Interfaces
di: Wu, Dongrui
Pubblicazione: (2025)
di: Wu, Dongrui
Pubblicazione: (2025)
A Two-Stage Learning-to-Defer Approach for Multi-Task Learning
di: Montreuil, Yannis, et al.
Pubblicazione: (2024)
di: Montreuil, Yannis, et al.
Pubblicazione: (2024)
Uncertainty Tube Visualization of Particle Trajectories
di: Li, Jixian, et al.
Pubblicazione: (2025)
di: Li, Jixian, et al.
Pubblicazione: (2025)
Graph-Based Learning of Spectro-Topographical EEG Representations with Gradient Alignment for Brain-Computer Interfaces
di: Angkan, Prithila, et al.
Pubblicazione: (2025)
di: Angkan, Prithila, et al.
Pubblicazione: (2025)
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
di: Pandey, Gaurav, et al.
Pubblicazione: (2024)
di: Pandey, Gaurav, et al.
Pubblicazione: (2024)
Designing Rewards for Rewarding Designs: Demonstrating the Impact of Rewards on the Creative Design Process
di: Nath, Surabhi S, et al.
Pubblicazione: (2026)
di: Nath, Surabhi S, et al.
Pubblicazione: (2026)
Discovering Cyclists' Visual Preferences Through Shared Bike Trajectories and Street View Images Using Inverse Reinforcement Learning
di: Ren, Kezhou, et al.
Pubblicazione: (2024)
di: Ren, Kezhou, et al.
Pubblicazione: (2024)
StackGenVis: Alignment of Data, Algorithms, and Models for Stacking Ensemble Learning Using Performance Metrics
di: Chatzimparmpas, Angelos, et al.
Pubblicazione: (2020)
di: Chatzimparmpas, Angelos, et al.
Pubblicazione: (2020)
An Analysis of Human Alignment of Latent Diffusion Models
di: Linhardt, Lorenz, et al.
Pubblicazione: (2024)
di: Linhardt, Lorenz, et al.
Pubblicazione: (2024)
Human-in-the-Loop Systems for Adaptive Learning Using Generative AI
di: Tarun, Bhavishya, et al.
Pubblicazione: (2025)
di: Tarun, Bhavishya, et al.
Pubblicazione: (2025)
Driving Safety Prediction and Safe Route Mapping Using In-vehicle and Roadside Data
di: Huang, Yufei, et al.
Pubblicazione: (2022)
di: Huang, Yufei, et al.
Pubblicazione: (2022)
Abstracted Trajectory Visualization for Explainability in Reinforcement Learning
di: Takagi, Yoshiki, et al.
Pubblicazione: (2024)
di: Takagi, Yoshiki, et al.
Pubblicazione: (2024)
Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining
di: Li, Aaron J., et al.
Pubblicazione: (2023)
di: Li, Aaron J., et al.
Pubblicazione: (2023)
CataractBot: An LLM-Powered Expert-in-the-Loop Chatbot for Cataract Patients
di: Ramjee, Pragnya, et al.
Pubblicazione: (2024)
di: Ramjee, Pragnya, et al.
Pubblicazione: (2024)
Realistic Adversarial Attacks for Robustness Evaluation of Trajectory Prediction Models via Future State Perturbation
di: Schumann, Julian F., et al.
Pubblicazione: (2025)
di: Schumann, Julian F., et al.
Pubblicazione: (2025)
Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education
di: Lekshmi-Narayanan, Arun-Balajiee, et al.
Pubblicazione: (2026)
di: Lekshmi-Narayanan, Arun-Balajiee, et al.
Pubblicazione: (2026)
A Population-to-individual Tuning Framework for Adapting Pretrained LM to On-device User Intent Prediction
di: Gong, Jiahui, et al.
Pubblicazione: (2024)
di: Gong, Jiahui, et al.
Pubblicazione: (2024)
Learning to Decide with AI Assistance under Human-Alignment
di: Benz, Nina Corvelo, et al.
Pubblicazione: (2026)
di: Benz, Nina Corvelo, et al.
Pubblicazione: (2026)
Reducing Label Dependency in Human Activity Recognition with Wearables: From Supervised Learning to Novel Weakly Self-Supervised Approaches
di: Sheng, Taoran, et al.
Pubblicazione: (2025)
di: Sheng, Taoran, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Super-human Vision-based Reinforcement Learning Agent for Autonomous Racing in Gran Turismo
di: Vasco, Miguel, et al.
Pubblicazione: (2024) -
Reward Learning through Ranking Mean Squared Error
di: Kharyal, Chaitanya, et al.
Pubblicazione: (2026) -
GFlowState: Visualizing the Training of Generative Flow Networks Beyond the Reward
di: Holeczek, Florian, et al.
Pubblicazione: (2026) -
Hindsight PRIORs for Reward Learning from Human Preferences
di: Verma, Mudit, et al.
Pubblicazione: (2024) -
Training AI Co-Scientists Using Rubric Rewards
di: Goel, Shashwat, et al.
Pubblicazione: (2025)