Saved in:
| Main Authors: | Dalrymple, David "davidad", Skalse, Joar, Bengio, Yoshua, Russell, Stuart, Tegmark, Max, Seshia, Sanjit, Omohundro, Steve, Szegedy, Christian, Goldhaber, Ben, Ammann, Nora, Abate, Alessandro, Halpern, Joe, Barrett, Clark, Zhao, Ding, Zhi-Xuan, Tan, Wing, Jeannette, Tenenbaum, Joshua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2405.06624 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Training Safe Neural Networks with Global SDP Bounds
by: Soletskyi, Roman, et al.
Published: (2024)
by: Soletskyi, Roman, et al.
Published: (2024)
Flexible Hardware-Enabled Guarantees for AI Compute
by: Petrie, James, et al.
Published: (2025)
by: Petrie, James, et al.
Published: (2025)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
by: Fluri, Lukas, et al.
Published: (2024)
by: Fluri, Lukas, et al.
Published: (2024)
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
STARC: A General Framework For Quantifying Differences Between Reward Functions
by: Skalse, Joar, et al.
Published: (2023)
by: Skalse, Joar, et al.
Published: (2023)
Learning Contextual Runtime Monitors for Safe AI-Based Autonomy
by: Luque-Cerpa, Alejandro, et al.
Published: (2026)
by: Luque-Cerpa, Alejandro, et al.
Published: (2026)
Defining and Characterizing Reward Hacking
by: Skalse, Joar, et al.
Published: (2022)
by: Skalse, Joar, et al.
Published: (2022)
On The Expressivity of Objective-Specification Formalisms in Reinforcement Learning
by: Subramani, Rohan, et al.
Published: (2023)
by: Subramani, Rohan, et al.
Published: (2023)
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
SemPat: Using Hyperproperty-based Semantic Analysis to Generate Microarchitectural Attack Patterns
by: Godbole, Adwait, et al.
Published: (2024)
by: Godbole, Adwait, et al.
Published: (2024)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization
by: Liu, Shengchao, et al.
Published: (2025)
by: Liu, Shengchao, et al.
Published: (2025)
Do Two AI Scientists Agree?
by: Fu, Xinghong, et al.
Published: (2025)
by: Fu, Xinghong, et al.
Published: (2025)
Artifact: PyCaliper: Python-embedded Infrastructure for RTL Verification and Specification Synthesis
by: Godbole, Adwait, et al.
Published: (2025)
by: Godbole, Adwait, et al.
Published: (2025)
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
by: Yalcinkaya, Beyazit, et al.
Published: (2024)
by: Yalcinkaya, Beyazit, et al.
Published: (2024)
Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning
by: Yalcinkaya, Beyazit, et al.
Published: (2025)
by: Yalcinkaya, Beyazit, et al.
Published: (2025)
Learning Formal Specifications from Membership and Preference Queries
by: Shah, Ameesh, et al.
Published: (2023)
by: Shah, Ameesh, et al.
Published: (2023)
On Generalization for Generative Flow Networks
by: Krichel, Anas, et al.
Published: (2024)
by: Krichel, Anas, et al.
Published: (2024)
Interventional Causal Representation Learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Relative Trajectory Balance is equivalent to Trust-PCL
by: Deleu, Tristan, et al.
Published: (2025)
by: Deleu, Tristan, et al.
Published: (2025)
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
by: Yoon, Jaesik, et al.
Published: (2025)
by: Yoon, Jaesik, et al.
Published: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
GFlowNet Foundations
by: Bengio, Yoshua, et al.
Published: (2021)
by: Bengio, Yoshua, et al.
Published: (2021)
Verified Code Transpilation with LLMs
by: Bhatia, Sahil, et al.
Published: (2024)
by: Bhatia, Sahil, et al.
Published: (2024)
Syzygy: Dual Code-Test C to (safe) Rust Translation using LLMs and Dynamic Analysis
by: Shetty, Manish, et al.
Published: (2024)
by: Shetty, Manish, et al.
Published: (2024)
Locally Pareto-Optimal Interpretations for Black-Box Machine Learning Models
by: Joshi, Aniruddha, et al.
Published: (2025)
by: Joshi, Aniruddha, et al.
Published: (2025)
Learning Symbolic Task Decompositions for Multi-Agent Teams
by: Shah, Ameesh, et al.
Published: (2025)
by: Shah, Ameesh, et al.
Published: (2025)
Star observations in bounded-degree graphs
by: Szegedy, Balazs
Published: (2026)
by: Szegedy, Balazs
Published: (2026)
Holographic functions and neural networks
by: Szegedy, Balazs
Published: (2026)
by: Szegedy, Balazs
Published: (2026)
A higher-order generalization of group theory
by: Szegedy, Balazs
Published: (2024)
by: Szegedy, Balazs
Published: (2024)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
RL, but don't do anything I wouldn't do
by: Cohen, Michael K., et al.
Published: (2024)
by: Cohen, Michael K., et al.
Published: (2024)
Local Search GFlowNets
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
Similar Items
-
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
by: Skalse, Joar, et al.
Published: (2024) -
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
by: Skalse, Joar, et al.
Published: (2024) -
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
by: Skalse, Joar, et al.
Published: (2024) -
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
by: Skalse, Joar, et al.
Published: (2024) -
Training Safe Neural Networks with Global SDP Bounds
by: Soletskyi, Roman, et al.
Published: (2024)