Saved in:
| Main Author: | Potham, Ram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.02357 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
by: Potham, Ram, et al.
Published: (2025)
by: Potham, Ram, et al.
Published: (2025)
MAEBE: Multi-Agent Emergent Behavior Framework
by: Erisken, Sinem, et al.
Published: (2025)
by: Erisken, Sinem, et al.
Published: (2025)
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
by: Heitzig, Jobst, et al.
Published: (2025)
by: Heitzig, Jobst, et al.
Published: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
by: Kelly, Christopher, et al.
Published: (2026)
by: Kelly, Christopher, et al.
Published: (2026)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2025)
by: Liu, Haokun, et al.
Published: (2025)
Moral Alignment for LLM Agents
by: Tennant, Elizaveta, et al.
Published: (2024)
by: Tennant, Elizaveta, et al.
Published: (2024)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Defining and Evaluating Physical Safety for Large Language Models
by: Tang, Yung-Chen, et al.
Published: (2024)
by: Tang, Yung-Chen, et al.
Published: (2024)
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
by: Raman, Naveen, et al.
Published: (2026)
by: Raman, Naveen, et al.
Published: (2026)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025)
by: Bowen, Dillon, et al.
Published: (2025)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital
by: Kothari, Avni, et al.
Published: (2025)
by: Kothari, Avni, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
by: Wilming, Rick, et al.
Published: (2024)
by: Wilming, Rick, et al.
Published: (2024)
A Justice Lens on Fairness and Ethics Courses in Computing Education: LLM-Assisted Multi-Perspective and Thematic Evaluation
by: Andrews, Kenya S., et al.
Published: (2025)
by: Andrews, Kenya S., et al.
Published: (2025)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
by: Arghal, Raghu, et al.
Published: (2026)
by: Arghal, Raghu, et al.
Published: (2026)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
The 2025 Foundation Model Transparency Index
by: Wan, Alexander, et al.
Published: (2025)
by: Wan, Alexander, et al.
Published: (2025)
On Catastrophic Inheritance of Large Foundation Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
On the Societal Impact of Open Foundation Models
by: Kapoor, Sayash, et al.
Published: (2024)
by: Kapoor, Sayash, et al.
Published: (2024)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Procedural Fairness Through Decoupling Objectionable Data Generating Components
by: Tang, Zeyu, et al.
Published: (2023)
by: Tang, Zeyu, et al.
Published: (2023)
Welfare as a Guiding Principle for Machine Learning -- From Compass, to Lens, to Roadmap
by: Rosenfeld, Nir, et al.
Published: (2025)
by: Rosenfeld, Nir, et al.
Published: (2025)
Ecosystem Graphs: The Social Footprint of Foundation Models
by: Bommasani, Rishi, et al.
Published: (2023)
by: Bommasani, Rishi, et al.
Published: (2023)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
Towards Urban General Intelligence: A Review and Outlook of Urban Foundation Models
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Can Machines Think Like Humans? A Behavioral Evaluation of LLM Agents in Dictator Games
by: Ma, Ji
Published: (2024)
by: Ma, Ji
Published: (2024)
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models II: Benchmark Generation Process
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Deprecating Benchmarks: Criteria and Framework
by: Joaquin, Ayrton San, et al.
Published: (2025)
by: Joaquin, Ayrton San, et al.
Published: (2025)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
Similar Items
-
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
by: Potham, Ram, et al.
Published: (2025) -
MAEBE: Multi-Agent Emergent Behavior Framework
by: Erisken, Sinem, et al.
Published: (2025) -
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
by: Heitzig, Jobst, et al.
Published: (2025) -
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024) -
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
by: Kelly, Christopher, et al.
Published: (2026)