Gespeichert in:
| Hauptverfasser: | Greenblatt, Ryan, Shlegeris, Buck, Sachan, Kshitij, Roger, Fabien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2312.06942 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Polysemanticity and Capacity in Neural Networks
von: Scherlis, Adam, et al.
Veröffentlicht: (2022)
von: Scherlis, Adam, et al.
Veröffentlicht: (2022)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
Stress-Testing Capability Elicitation With Password-Locked Models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Ctrl-Z: Controlling AI Agents via Resampling
von: Bhatt, Aryan, et al.
Veröffentlicht: (2025)
von: Bhatt, Aryan, et al.
Veröffentlicht: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
Alignment faking in large language models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
von: Hubinger, Evan, et al.
Veröffentlicht: (2024)
von: Hubinger, Evan, et al.
Veröffentlicht: (2024)
Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation
von: Canavan, Callum, et al.
Veröffentlicht: (2026)
von: Canavan, Callum, et al.
Veröffentlicht: (2026)
Do Unlearning Methods Remove Information from Language Model Weights?
von: Deeb, Aghyad, et al.
Veröffentlicht: (2024)
von: Deeb, Aghyad, et al.
Veröffentlicht: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
von: Khullar, Dipika, et al.
Veröffentlicht: (2026)
von: Khullar, Dipika, et al.
Veröffentlicht: (2026)
Steering Language Models with Weight Arithmetic
von: Fierro, Constanza, et al.
Veröffentlicht: (2025)
von: Fierro, Constanza, et al.
Veröffentlicht: (2025)
Safety-Critical Traffic Simulation with Adversarial Transfer of Driving Intentions
von: Huang, Zherui, et al.
Veröffentlicht: (2025)
von: Huang, Zherui, et al.
Veröffentlicht: (2025)
VRU-CIPI: Crossing Intention Prediction at Intersections for Improving Vulnerable Road Users Safety
von: Abdelrahman, Ahmed S., et al.
Veröffentlicht: (2025)
von: Abdelrahman, Ahmed S., et al.
Veröffentlicht: (2025)
The Persistence of Neural Collapse Despite Low-Rank Bias
von: Garrod, Connall, et al.
Veröffentlicht: (2024)
von: Garrod, Connall, et al.
Veröffentlicht: (2024)
Energy-Guided Data Sampling for Traffic Prediction with Mini Training Datasets
von: Yang, Zhaohui, et al.
Veröffentlicht: (2024)
von: Yang, Zhaohui, et al.
Veröffentlicht: (2024)
DiFR: Inference Verification Despite Nondeterminism
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
von: Xue, Anton, et al.
Veröffentlicht: (2024)
von: Xue, Anton, et al.
Veröffentlicht: (2024)
ProPACT: A Proactive AI-Driven Adaptive Collaborative Tutor for Pair Programming
von: Golrang, Anahita, et al.
Veröffentlicht: (2026)
von: Golrang, Anahita, et al.
Veröffentlicht: (2026)
A physics-informed U-Net-LSTM network for nonlinear structural response under seismic excitation
von: Biswas, Sutirtha, et al.
Veröffentlicht: (2025)
von: Biswas, Sutirtha, et al.
Veröffentlicht: (2025)
Pretrained LLMs as Real-Time Controllers for Robot Operated Serial Production Line
von: Waseem, Muhammad, et al.
Veröffentlicht: (2025)
von: Waseem, Muhammad, et al.
Veröffentlicht: (2025)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
von: Ozyurt, Yilmazcan, et al.
Veröffentlicht: (2024)
von: Ozyurt, Yilmazcan, et al.
Veröffentlicht: (2024)
Learning-Augmented Robust Algorithmic Recourse
von: Kayastha, Kshitij, et al.
Veröffentlicht: (2024)
von: Kayastha, Kshitij, et al.
Veröffentlicht: (2024)
Optimal Robust Recourse with $L^p$-Bounded Model Change
von: Kyaw, Phone, et al.
Veröffentlicht: (2025)
von: Kyaw, Phone, et al.
Veröffentlicht: (2025)
Sabotage Evaluations for Frontier Models
von: Benton, Joe, et al.
Veröffentlicht: (2024)
von: Benton, Joe, et al.
Veröffentlicht: (2024)
Improving Language Models with Intentional Analysis
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness
von: Patel, Kumar Kshitij
Veröffentlicht: (2025)
von: Patel, Kumar Kshitij
Veröffentlicht: (2025)
Excess Description Length of Learning Generalizable Predictors
von: Donoway, Elizabeth, et al.
Veröffentlicht: (2026)
von: Donoway, Elizabeth, et al.
Veröffentlicht: (2026)
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
von: Choi, Juheon, et al.
Veröffentlicht: (2025)
von: Choi, Juheon, et al.
Veröffentlicht: (2025)
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Tabular and Deep Reinforcement Learning for Gittins Index
von: Dhankhar, Harshit, et al.
Veröffentlicht: (2024)
von: Dhankhar, Harshit, et al.
Veröffentlicht: (2024)
Optimality of Sub-network Laplace Approximations: New Results and Methods
von: Raha, Swarnali, et al.
Veröffentlicht: (2026)
von: Raha, Swarnali, et al.
Veröffentlicht: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
A Framework for Evaluating Faithfulness in Explainable AI for Machine Anomalous Sound Detection Using Frequency-Band Perturbation
von: Buck, Alexander, et al.
Veröffentlicht: (2026)
von: Buck, Alexander, et al.
Veröffentlicht: (2026)
Structured Sparsity and Weight-adaptive Pruning for Memory and Compute efficient Whisper models
von: Mudi, Prasenjit K, et al.
Veröffentlicht: (2025)
von: Mudi, Prasenjit K, et al.
Veröffentlicht: (2025)
Auto-Adaptive PINNs with Applications to Phase Transitions
von: Buck, Kevin, et al.
Veröffentlicht: (2025)
von: Buck, Kevin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Polysemanticity and Capacity in Neural Networks
von: Scherlis, Adam, et al.
Veröffentlicht: (2022) -
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
von: Griffin, Charlie, et al.
Veröffentlicht: (2024) -
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
von: Mallen, Alex, et al.
Veröffentlicht: (2024) -
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022) -
Stress-Testing Capability Elicitation With Password-Locked Models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)