Frontier Models Can Take Actions at Low Probabilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Serrano, Alex, Xing, Wen, Lindner, David, Jenner, Erik |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
por: Gupta, Rohan, et al.
Publicado: (2025)
por: Gupta, Rohan, et al.
Publicado: (2025)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
por: Zolkowski, Artur, et al.
Publicado: (2025)
por: Zolkowski, Artur, et al.
Publicado: (2025)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
por: Fronsdal, Kai, et al.
Publicado: (2024)
por: Fronsdal, Kai, et al.
Publicado: (2024)
Evaluating Frontier Models for Stealth and Situational Awareness
por: Phuong, Mary, et al.
Publicado: (2025)
por: Phuong, Mary, et al.
Publicado: (2025)
Obfuscated Activations Bypass LLM Latent-Space Defenses
por: Bailey, Luke, et al.
Publicado: (2024)
por: Bailey, Luke, et al.
Publicado: (2024)
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
por: McGuinness, Max, et al.
Publicado: (2025)
por: McGuinness, Max, et al.
Publicado: (2025)
Evaluating Frontier Models for Dangerous Capabilities
por: Phuong, Mary, et al.
Publicado: (2024)
por: Phuong, Mary, et al.
Publicado: (2024)
Early Signs of Steganographic Capabilities in Frontier LLMs
por: Zolkowski, Artur, et al.
Publicado: (2025)
por: Zolkowski, Artur, et al.
Publicado: (2025)
Can Machines Learn the True Probabilities?
por: Kim, Jinsook
Publicado: (2024)
por: Kim, Jinsook
Publicado: (2024)
Predicting Fault-Ride-Through Probability of Inverter-Dominated Power Grids using Machine Learning
por: Nauck, Christian, et al.
Publicado: (2024)
por: Nauck, Christian, et al.
Publicado: (2024)
Pragmatist Intelligence: Where the Principle of Usefulness Can Take ANNs
por: Bikić, Antonio, et al.
Publicado: (2024)
por: Bikić, Antonio, et al.
Publicado: (2024)
An Investigation of Offline Reinforcement Learning in Factorisable Action Spaces
por: Beeson, Alex, et al.
Publicado: (2024)
por: Beeson, Alex, et al.
Publicado: (2024)
Learning to Represent Surroundings, Anticipate Motion and Take Informed Actions in Unstructured Environments
por: Zhi, Weiming
Publicado: (2024)
por: Zhi, Weiming
Publicado: (2024)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
por: Lang, Leon, et al.
Publicado: (2024)
por: Lang, Leon, et al.
Publicado: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
por: Jenner, Erik, et al.
Publicado: (2024)
por: Jenner, Erik, et al.
Publicado: (2024)
Does Spatial Cognition Emerge in Frontier Models?
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2024)
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
por: Lee, Heekyung, et al.
Publicado: (2025)
por: Lee, Heekyung, et al.
Publicado: (2025)
Analysis of Value Iteration Through Absolute Probability Sequences
por: Mustafin, Arsenii, et al.
Publicado: (2025)
por: Mustafin, Arsenii, et al.
Publicado: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
por: Farquhar, Sebastian, et al.
Publicado: (2025)
por: Farquhar, Sebastian, et al.
Publicado: (2025)
STARC: A General Framework For Quantifying Differences Between Reward Functions
por: Skalse, Joar, et al.
Publicado: (2023)
por: Skalse, Joar, et al.
Publicado: (2023)
PPGF: Probability Pattern-Guided Time Series Forecasting
por: Sun, Yanru, et al.
Publicado: (2025)
por: Sun, Yanru, et al.
Publicado: (2025)
Probability-density-aware Semi-supervised Learning
por: Liu, Shuyang, et al.
Publicado: (2024)
por: Liu, Shuyang, et al.
Publicado: (2024)
It Takes a Good Model to Train a Good Model: Generalized Gaussian Priors for Optimized LLMs
por: Wu, Jun, et al.
Publicado: (2025)
por: Wu, Jun, et al.
Publicado: (2025)
Unveiling High-Probability Generalization in Decentralized SGD
por: Wang, Jiahuan, et al.
Publicado: (2026)
por: Wang, Jiahuan, et al.
Publicado: (2026)
Taking the GP Out of the Loop
por: Bafna, Mehul, et al.
Publicado: (2025)
por: Bafna, Mehul, et al.
Publicado: (2025)
Joint Bayesian Parameter and Model Order Estimation for Low-Rank Probability Mass Tensors
por: Chege, Joseph K., et al.
Publicado: (2024)
por: Chege, Joseph K., et al.
Publicado: (2024)
Adversaries Can Misuse Combinations of Safe Models
por: Jones, Erik, et al.
Publicado: (2024)
por: Jones, Erik, et al.
Publicado: (2024)
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
por: Zhu, Yuanyang, et al.
Publicado: (2024)
por: Zhu, Yuanyang, et al.
Publicado: (2024)
Probably Approximately Correct Causal Discovery
por: Wei, Mian, et al.
Publicado: (2025)
por: Wei, Mian, et al.
Publicado: (2025)
Stabilized Inverse Probability Weighting via Isotonic Calibration
por: van der Laan, Lars, et al.
Publicado: (2024)
por: van der Laan, Lars, et al.
Publicado: (2024)
Density-Informed VAE (DiVAE): Reliable Log-Prior Probability via Density Alignment Regularization
por: Alessi, Michele, et al.
Publicado: (2025)
por: Alessi, Michele, et al.
Publicado: (2025)
Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density
por: Fei, Jingru, et al.
Publicado: (2026)
por: Fei, Jingru, et al.
Publicado: (2026)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
por: Jang, Eyon, et al.
Publicado: (2026)
por: Jang, Eyon, et al.
Publicado: (2026)
Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language Models
por: Nguyen, Manh, et al.
Publicado: (2025)
por: Nguyen, Manh, et al.
Publicado: (2025)
Interpretable Probability Estimation with LLMs via Shapley Reconstruction
por: Nan, Yang, et al.
Publicado: (2026)
por: Nan, Yang, et al.
Publicado: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
Multiclass Calibration Assessment and Recalibration of Probability Predictions via the Linear Log Odds Calibration Function
por: Vennos, Amy, et al.
Publicado: (2026)
por: Vennos, Amy, et al.
Publicado: (2026)
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
por: Rocamonde, Juan, et al.
Publicado: (2023)
por: Rocamonde, Juan, et al.
Publicado: (2023)
Short-Long Policy Evaluation with Novel Actions
por: Nam, Hyunji Alex, et al.
Publicado: (2024)
por: Nam, Hyunji Alex, et al.
Publicado: (2024)
Sabotage Evaluations for Frontier Models
por: Benton, Joe, et al.
Publicado: (2024)
por: Benton, Joe, et al.
Publicado: (2024)
Ejemplares similares
-
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
por: Gupta, Rohan, et al.
Publicado: (2025) -
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
por: Zolkowski, Artur, et al.
Publicado: (2025) -
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
por: Fronsdal, Kai, et al.
Publicado: (2024) -
Evaluating Frontier Models for Stealth and Situational Awareness
por: Phuong, Mary, et al.
Publicado: (2025) -
Obfuscated Activations Bypass LLM Latent-Space Defenses
por: Bailey, Luke, et al.
Publicado: (2024)