Frontier Models Can Take Actions at Low Probabilities
Fuente:
arXiv
Salvato in:
| Autori principali: | Serrano, Alex, Xing, Wen, Lindner, David, Jenner, Erik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
di: Gupta, Rohan, et al.
Pubblicazione: (2025)
di: Gupta, Rohan, et al.
Pubblicazione: (2025)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
di: Fronsdal, Kai, et al.
Pubblicazione: (2024)
di: Fronsdal, Kai, et al.
Pubblicazione: (2024)
Evaluating Frontier Models for Stealth and Situational Awareness
di: Phuong, Mary, et al.
Pubblicazione: (2025)
di: Phuong, Mary, et al.
Pubblicazione: (2025)
Obfuscated Activations Bypass LLM Latent-Space Defenses
di: Bailey, Luke, et al.
Pubblicazione: (2024)
di: Bailey, Luke, et al.
Pubblicazione: (2024)
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
di: McGuinness, Max, et al.
Pubblicazione: (2025)
di: McGuinness, Max, et al.
Pubblicazione: (2025)
Evaluating Frontier Models for Dangerous Capabilities
di: Phuong, Mary, et al.
Pubblicazione: (2024)
di: Phuong, Mary, et al.
Pubblicazione: (2024)
Early Signs of Steganographic Capabilities in Frontier LLMs
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
Can Machines Learn the True Probabilities?
di: Kim, Jinsook
Pubblicazione: (2024)
di: Kim, Jinsook
Pubblicazione: (2024)
Predicting Fault-Ride-Through Probability of Inverter-Dominated Power Grids using Machine Learning
di: Nauck, Christian, et al.
Pubblicazione: (2024)
di: Nauck, Christian, et al.
Pubblicazione: (2024)
Pragmatist Intelligence: Where the Principle of Usefulness Can Take ANNs
di: Bikić, Antonio, et al.
Pubblicazione: (2024)
di: Bikić, Antonio, et al.
Pubblicazione: (2024)
An Investigation of Offline Reinforcement Learning in Factorisable Action Spaces
di: Beeson, Alex, et al.
Pubblicazione: (2024)
di: Beeson, Alex, et al.
Pubblicazione: (2024)
Learning to Represent Surroundings, Anticipate Motion and Take Informed Actions in Unstructured Environments
di: Zhi, Weiming
Pubblicazione: (2024)
di: Zhi, Weiming
Pubblicazione: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
di: Lang, Leon, et al.
Pubblicazione: (2024)
di: Lang, Leon, et al.
Pubblicazione: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
di: Jenner, Erik, et al.
Pubblicazione: (2024)
di: Jenner, Erik, et al.
Pubblicazione: (2024)
Does Spatial Cognition Emerge in Frontier Models?
di: Ramakrishnan, Santhosh Kumar, et al.
Pubblicazione: (2024)
di: Ramakrishnan, Santhosh Kumar, et al.
Pubblicazione: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
di: Lee, Heekyung, et al.
Pubblicazione: (2025)
di: Lee, Heekyung, et al.
Pubblicazione: (2025)
Analysis of Value Iteration Through Absolute Probability Sequences
di: Mustafin, Arsenii, et al.
Pubblicazione: (2025)
di: Mustafin, Arsenii, et al.
Pubblicazione: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
STARC: A General Framework For Quantifying Differences Between Reward Functions
di: Skalse, Joar, et al.
Pubblicazione: (2023)
di: Skalse, Joar, et al.
Pubblicazione: (2023)
PPGF: Probability Pattern-Guided Time Series Forecasting
di: Sun, Yanru, et al.
Pubblicazione: (2025)
di: Sun, Yanru, et al.
Pubblicazione: (2025)
Probability-density-aware Semi-supervised Learning
di: Liu, Shuyang, et al.
Pubblicazione: (2024)
di: Liu, Shuyang, et al.
Pubblicazione: (2024)
It Takes a Good Model to Train a Good Model: Generalized Gaussian Priors for Optimized LLMs
di: Wu, Jun, et al.
Pubblicazione: (2025)
di: Wu, Jun, et al.
Pubblicazione: (2025)
Unveiling High-Probability Generalization in Decentralized SGD
di: Wang, Jiahuan, et al.
Pubblicazione: (2026)
di: Wang, Jiahuan, et al.
Pubblicazione: (2026)
Taking the GP Out of the Loop
di: Bafna, Mehul, et al.
Pubblicazione: (2025)
di: Bafna, Mehul, et al.
Pubblicazione: (2025)
Joint Bayesian Parameter and Model Order Estimation for Low-Rank Probability Mass Tensors
di: Chege, Joseph K., et al.
Pubblicazione: (2024)
di: Chege, Joseph K., et al.
Pubblicazione: (2024)
Adversaries Can Misuse Combinations of Safe Models
di: Jones, Erik, et al.
Pubblicazione: (2024)
di: Jones, Erik, et al.
Pubblicazione: (2024)
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
di: Zhu, Yuanyang, et al.
Pubblicazione: (2024)
di: Zhu, Yuanyang, et al.
Pubblicazione: (2024)
Probably Approximately Correct Causal Discovery
di: Wei, Mian, et al.
Pubblicazione: (2025)
di: Wei, Mian, et al.
Pubblicazione: (2025)
Stabilized Inverse Probability Weighting via Isotonic Calibration
di: van der Laan, Lars, et al.
Pubblicazione: (2024)
di: van der Laan, Lars, et al.
Pubblicazione: (2024)
Density-Informed VAE (DiVAE): Reliable Log-Prior Probability via Density Alignment Regularization
di: Alessi, Michele, et al.
Pubblicazione: (2025)
di: Alessi, Michele, et al.
Pubblicazione: (2025)
Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density
di: Fei, Jingru, et al.
Pubblicazione: (2026)
di: Fei, Jingru, et al.
Pubblicazione: (2026)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
di: Jang, Eyon, et al.
Pubblicazione: (2026)
di: Jang, Eyon, et al.
Pubblicazione: (2026)
Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language Models
di: Nguyen, Manh, et al.
Pubblicazione: (2025)
di: Nguyen, Manh, et al.
Pubblicazione: (2025)
Interpretable Probability Estimation with LLMs via Shapley Reconstruction
di: Nan, Yang, et al.
Pubblicazione: (2026)
di: Nan, Yang, et al.
Pubblicazione: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
Multiclass Calibration Assessment and Recalibration of Probability Predictions via the Linear Log Odds Calibration Function
di: Vennos, Amy, et al.
Pubblicazione: (2026)
di: Vennos, Amy, et al.
Pubblicazione: (2026)
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
di: Rocamonde, Juan, et al.
Pubblicazione: (2023)
di: Rocamonde, Juan, et al.
Pubblicazione: (2023)
Short-Long Policy Evaluation with Novel Actions
di: Nam, Hyunji Alex, et al.
Pubblicazione: (2024)
di: Nam, Hyunji Alex, et al.
Pubblicazione: (2024)
Sabotage Evaluations for Frontier Models
di: Benton, Joe, et al.
Pubblicazione: (2024)
di: Benton, Joe, et al.
Pubblicazione: (2024)
Documenti analoghi
-
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
di: Gupta, Rohan, et al.
Pubblicazione: (2025) -
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
di: Zolkowski, Artur, et al.
Pubblicazione: (2025) -
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
di: Fronsdal, Kai, et al.
Pubblicazione: (2024) -
Evaluating Frontier Models for Stealth and Situational Awareness
di: Phuong, Mary, et al.
Pubblicazione: (2025) -
Obfuscated Activations Bypass LLM Latent-Space Defenses
di: Bailey, Luke, et al.
Pubblicazione: (2024)