Gespeichert in:
| Hauptverfasser: | Skalse, Joar, Farnik, Lucy, Motwani, Sumeet Ramesh, Jenner, Erik, Gleave, Adam, Abate, Alessandro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2309.15257 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
Temporal-Difference Variational Continual Learning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
von: Farnik, Lucy, et al.
Veröffentlicht: (2025)
Defining and Characterizing Reward Hacking
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2025)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2025)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
von: Putta, Pranav, et al.
Veröffentlicht: (2024)
von: Putta, Pranav, et al.
Veröffentlicht: (2024)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2024)
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
von: Struppek, Lukas, et al.
Veröffentlicht: (2026)
MALT: Improving Reasoning with Multi-Agent LLM Training
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2024)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2024)
Neural Proofs for Sound Verification and Control of Complex Systems
von: Abate, Alessandro
Veröffentlicht: (2025)
von: Abate, Alessandro
Veröffentlicht: (2025)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2025)
Networked Communication for Mean-Field Games with Function Approximation and Empirical Mean-Field Estimation
von: Benjamin, Patrick, et al.
Veröffentlicht: (2024)
von: Benjamin, Patrick, et al.
Veröffentlicht: (2024)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
Can Go AIs be adversarially robust?
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
Zero-Shot Instruction Following in RL via Structured LTL Representations
von: Giuri, Mattia, et al.
Veröffentlicht: (2025)
von: Giuri, Mattia, et al.
Veröffentlicht: (2025)
Efficient Solution and Learning of Robust Factored MDPs
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
Inducing Human-like Biases in Moral Reasoning Language Models
von: Karpov, Artem, et al.
Veröffentlicht: (2024)
von: Karpov, Artem, et al.
Veröffentlicht: (2024)
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
von: Dalrymple, David "davidad", et al.
Veröffentlicht: (2024)
von: Dalrymple, David "davidad", et al.
Veröffentlicht: (2024)
Zero-Shot Visual Generalization in Robot Manipulation
von: Batra, Sumeet, et al.
Veröffentlicht: (2025)
von: Batra, Sumeet, et al.
Veröffentlicht: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
Scaling Trends for Data Poisoning in LLMs
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
Certifiably Robust Policies for Uncertain Parametric Environments
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2024)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2024)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
Zero-Shot Instruction Following in RL via Structured LTL Representations
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2026)
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2026)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025)
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
von: Lang, Leon, et al.
Veröffentlicht: (2024)
von: Lang, Leon, et al.
Veröffentlicht: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
von: Jenner, Erik, et al.
Veröffentlicht: (2024)
von: Jenner, Erik, et al.
Veröffentlicht: (2024)
Networked Communication for Decentralised Agents in Mean-Field Games
von: Benjamin, Patrick, et al.
Veröffentlicht: (2023)
von: Benjamin, Patrick, et al.
Veröffentlicht: (2023)
Reinforcement Learning with Quasi-Hyperbolic Discounting
von: Eshwar, S. R., et al.
Veröffentlicht: (2024)
von: Eshwar, S. R., et al.
Veröffentlicht: (2024)
Planning in a recurrent neural network that plays Sokoban
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2024)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Studying Cross-cluster Modularity in Neural Networks
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
von: Fluri, Lukas, et al.
Veröffentlicht: (2024)