Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Potham, Ram, Harms, Max |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
von: Potham, Ram
Veröffentlicht: (2025)
von: Potham, Ram
Veröffentlicht: (2025)
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
von: Heitzig, Jobst, et al.
Veröffentlicht: (2025)
von: Heitzig, Jobst, et al.
Veröffentlicht: (2025)
MAEBE: Multi-Agent Emergent Behavior Framework
von: Erisken, Sinem, et al.
Veröffentlicht: (2025)
von: Erisken, Sinem, et al.
Veröffentlicht: (2025)
Reliable and Responsible Foundation Models: A Comprehensive Survey
von: Yang, Xinyu, et al.
Veröffentlicht: (2026)
von: Yang, Xinyu, et al.
Veröffentlicht: (2026)
Foundation Model Transparency Reports
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
The 2025 Foundation Model Transparency Index
von: Wan, Alexander, et al.
Veröffentlicht: (2025)
von: Wan, Alexander, et al.
Veröffentlicht: (2025)
On Catastrophic Inheritance of Large Foundation Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
The 2024 Foundation Model Transparency Index
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
On the Societal Impact of Open Foundation Models
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Fair and Reliable Organ Allocation
von: Kim, Brian Hyeongseok, et al.
Veröffentlicht: (2025)
von: Kim, Brian Hyeongseok, et al.
Veröffentlicht: (2025)
Ecosystem Graphs: The Social Footprint of Foundation Models
von: Bommasani, Rishi, et al.
Veröffentlicht: (2023)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2023)
Towards a Science of AI Agent Reliability
von: Rabanser, Stephan, et al.
Veröffentlicht: (2026)
von: Rabanser, Stephan, et al.
Veröffentlicht: (2026)
Towards Urban General Intelligence: A Review and Outlook of Urban Foundation Models
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
Training Foundation Models as Data Compression: On Information, Model Weights and Copyright Law
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024)
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction
von: Li, Zhonghang, et al.
Veröffentlicht: (2024)
von: Li, Zhonghang, et al.
Veröffentlicht: (2024)
Demographic Bias of Expert-Level Vision-Language Foundation Models in Medical Imaging
von: Yang, Yuzhe, et al.
Veröffentlicht: (2024)
von: Yang, Yuzhe, et al.
Veröffentlicht: (2024)
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2024)
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2024)
FairTargetSim: An Interactive Simulator for Understanding and Explaining the Fairness Effects of Target Variable Definition
von: Gala, Dalia, et al.
Veröffentlicht: (2024)
von: Gala, Dalia, et al.
Veröffentlicht: (2024)
Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale
von: Cooper, A. Feder
Veröffentlicht: (2024)
von: Cooper, A. Feder
Veröffentlicht: (2024)
RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning
von: Liu, Junyuan, et al.
Veröffentlicht: (2025)
von: Liu, Junyuan, et al.
Veröffentlicht: (2025)
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
von: Rabanser, Stephan
Veröffentlicht: (2025)
von: Rabanser, Stephan
Veröffentlicht: (2025)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
Core Safety Values for Provably Corrigible Agents
von: Nayebi, Aran
Veröffentlicht: (2025)
von: Nayebi, Aran
Veröffentlicht: (2025)
Scaling Laws For Scalable Oversight
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
The Quest for Reliable Metrics of Responsible AI
von: Rampisela, Theresia Veronika, et al.
Veröffentlicht: (2025)
von: Rampisela, Theresia Veronika, et al.
Veröffentlicht: (2025)
Vision Paper: Designing Graph Neural Networks in Compliance with the European Artificial Intelligence Act
von: Hoffmann, Barbara, et al.
Veröffentlicht: (2024)
von: Hoffmann, Barbara, et al.
Veröffentlicht: (2024)
Statistical Validation in Cultural Adaptations of Cognitive Tests: A Multi- Regional Systematic Review
von: Daga, Miit, et al.
Veröffentlicht: (2025)
von: Daga, Miit, et al.
Veröffentlicht: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
Empowering Cognitive Digital Twins with Generative Foundation Models: Developing a Low-Carbon Integrated Freight Transportation System
von: Li, Xueping, et al.
Veröffentlicht: (2024)
von: Li, Xueping, et al.
Veröffentlicht: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
von: Barrett, Anthony M., et al.
Veröffentlicht: (2024)
von: Barrett, Anthony M., et al.
Veröffentlicht: (2024)
Social Perception of Faces in a Vision-Language Model
von: Hausladen, Carina I., et al.
Veröffentlicht: (2024)
von: Hausladen, Carina I., et al.
Veröffentlicht: (2024)
Explaining How Quantization Disparately Skews a Model
von: Bellam, Abhimanyu, et al.
Veröffentlicht: (2025)
von: Bellam, Abhimanyu, et al.
Veröffentlicht: (2025)
Building a Domain-specific Guardrail Model in Production
von: Niknazar, Mohammad, et al.
Veröffentlicht: (2024)
von: Niknazar, Mohammad, et al.
Veröffentlicht: (2024)
The Pursuit of Fairness in Artificial Intelligence Models: A Survey
von: Kheya, Tahsin Alamgir, et al.
Veröffentlicht: (2024)
von: Kheya, Tahsin Alamgir, et al.
Veröffentlicht: (2024)
Say My Name: a Model's Bias Discovery Framework
von: Ciranni, Massimiliano, et al.
Veröffentlicht: (2024)
von: Ciranni, Massimiliano, et al.
Veröffentlicht: (2024)
VirnyFlow: A Design Space for Responsible Model Development
von: Herasymuk, Denys, et al.
Veröffentlicht: (2025)
von: Herasymuk, Denys, et al.
Veröffentlicht: (2025)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
von: Potham, Ram
Veröffentlicht: (2025) -
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
von: Heitzig, Jobst, et al.
Veröffentlicht: (2025) -
MAEBE: Multi-Agent Emergent Behavior Framework
von: Erisken, Sinem, et al.
Veröffentlicht: (2025) -
Reliable and Responsible Foundation Models: A Comprehensive Survey
von: Yang, Xinyu, et al.
Veröffentlicht: (2026) -
Foundation Model Transparency Reports
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)