Mitigating loss of control in advanced AI systems through instrumental goal trajectories
Fuente:
arXiv
Salvato in:
| Autore principale: | Fourie, Willem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Aristotelian ontology of instrumental goals: Structural features to be managed and not failures to be eliminated
di: Fourie, Willem
Pubblicazione: (2025)
di: Fourie, Willem
Pubblicazione: (2025)
Deciding how to respond: A deliberative framework to guide policymaker responses to AI systems
di: Fourie, Willem
Pubblicazione: (2025)
di: Fourie, Willem
Pubblicazione: (2025)
How will advanced AI systems impact democracy?
di: Summerfield, Christopher, et al.
Pubblicazione: (2024)
di: Summerfield, Christopher, et al.
Pubblicazione: (2024)
Mapping AI Risk Mitigations: Evidence Scan and Preliminary AI Risk Mitigation Taxonomy
di: Saeri, Alexander K., et al.
Pubblicazione: (2025)
di: Saeri, Alexander K., et al.
Pubblicazione: (2025)
AI Identity, Empowerment, and Mindfulness in Mitigating Unethical AI Use
di: Shaayesteh, Mayssam Tarighi, et al.
Pubblicazione: (2025)
di: Shaayesteh, Mayssam Tarighi, et al.
Pubblicazione: (2025)
Effective Mitigations for Systemic Risks from General-Purpose AI
di: Uuk, Risto, et al.
Pubblicazione: (2024)
di: Uuk, Risto, et al.
Pubblicazione: (2024)
Mitigating Societal Cognitive Overload in the Age of AI: Challenges and Directions
di: Lahlou, Salem
Pubblicazione: (2025)
di: Lahlou, Salem
Pubblicazione: (2025)
AI Mimicry and Human Dignity: Chatbot Use as a Violation of Self-Respect
di: van der Rijt, Jan-Willem, et al.
Pubblicazione: (2025)
di: van der Rijt, Jan-Willem, et al.
Pubblicazione: (2025)
Safeguarding Marketing Research: The Generation, Identification, and Mitigation of AI-Fabricated Disinformation
di: Mukherjee, Anirban
Pubblicazione: (2024)
di: Mukherjee, Anirban
Pubblicazione: (2024)
Generative AI and Power Imbalances in Global Education: Frameworks for Bias Mitigation
di: Nyaaba, Matthew, et al.
Pubblicazione: (2024)
di: Nyaaba, Matthew, et al.
Pubblicazione: (2024)
Ethical Concerns of Generative AI and Mitigation Strategies: A Systematic Mapping Study
di: Huang, Yutan, et al.
Pubblicazione: (2025)
di: Huang, Yutan, et al.
Pubblicazione: (2025)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
di: El-Sayed, Seliem, et al.
Pubblicazione: (2024)
di: El-Sayed, Seliem, et al.
Pubblicazione: (2024)
Mapping data literacy trajectories in K-12 education
di: Whyte, Robert, et al.
Pubblicazione: (2026)
di: Whyte, Robert, et al.
Pubblicazione: (2026)
When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies
di: Popchanovska, Evgenija, et al.
Pubblicazione: (2026)
di: Popchanovska, Evgenija, et al.
Pubblicazione: (2026)
From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI
di: Dunivin, Zackary Okun, et al.
Pubblicazione: (2026)
di: Dunivin, Zackary Okun, et al.
Pubblicazione: (2026)
Healthy Distrust in AI systems
di: Paaßen, Benjamin, et al.
Pubblicazione: (2025)
di: Paaßen, Benjamin, et al.
Pubblicazione: (2025)
What if AI systems weren't chatbots?
di: Ghosh, Sourojit, et al.
Pubblicazione: (2026)
di: Ghosh, Sourojit, et al.
Pubblicazione: (2026)
Deceptive AI systems that give explanations are more convincing than honest AI systems and can amplify belief in misinformation
di: Danry, Valdemar, et al.
Pubblicazione: (2024)
di: Danry, Valdemar, et al.
Pubblicazione: (2024)
Implications for Governance in Public Perceptions of Societal-scale AI Risks
di: Gruetzemacher, Ross, et al.
Pubblicazione: (2024)
di: Gruetzemacher, Ross, et al.
Pubblicazione: (2024)
Deconstructing Student Perceptions of Generative AI (GenAI) through an Expectancy Value Theory (EVT)-based Instrument
di: Chan, Cecilia Ka Yuk, et al.
Pubblicazione: (2023)
di: Chan, Cecilia Ka Yuk, et al.
Pubblicazione: (2023)
When Discourse Stalls: Moving Past Five Semantic Stopsigns about Generative AI in Design Research
di: van der Maden, Willem, et al.
Pubblicazione: (2025)
di: van der Maden, Willem, et al.
Pubblicazione: (2025)
Distributed agency in second language learning and teaching through generative AI
di: Godwin-Jones, Robert
Pubblicazione: (2024)
di: Godwin-Jones, Robert
Pubblicazione: (2024)
AI threats to national security can be countered through an incident regime
di: Ortega, Alejandro
Pubblicazione: (2025)
di: Ortega, Alejandro
Pubblicazione: (2025)
The implicated scientist: on the role of AI researchers in the development of weapons systems
di: Volokhova, Alexandra, et al.
Pubblicazione: (2026)
di: Volokhova, Alexandra, et al.
Pubblicazione: (2026)
Positive AI: Key Challenges in Designing Artificial Intelligence for Wellbeing
di: van der Maden, Willem, et al.
Pubblicazione: (2023)
di: van der Maden, Willem, et al.
Pubblicazione: (2023)
Tracing GenAI Literacy: Uncovering Student-AI Interaction Patterns in Academic Writing through Epistemic Network Analysis
di: Chen, Angxuan, et al.
Pubblicazione: (2026)
di: Chen, Angxuan, et al.
Pubblicazione: (2026)
Achieving Responsible AI through ESG: Insights and Recommendations from Industry Engagement
di: Perera, Harsha, et al.
Pubblicazione: (2024)
di: Perera, Harsha, et al.
Pubblicazione: (2024)
Developmental trajectories of decision making and affective dynamics in large language models
di: Wang, Zhihao, et al.
Pubblicazione: (2025)
di: Wang, Zhihao, et al.
Pubblicazione: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
di: Bowen, Dillon, et al.
Pubblicazione: (2025)
di: Bowen, Dillon, et al.
Pubblicazione: (2025)
Understanding and Mitigating Risks of Generative AI in Financial Services
di: Gehrmann, Sebastian, et al.
Pubblicazione: (2025)
di: Gehrmann, Sebastian, et al.
Pubblicazione: (2025)
"My Kind of Woman": Analysing Gender Stereotypes in AI through The Averageness Theory and EU Law
di: Doh, Miriam, et al.
Pubblicazione: (2024)
di: Doh, Miriam, et al.
Pubblicazione: (2024)
Exploring Societal Concerns and Perceptions of AI: A Thematic Analysis through the Lens of Problem-Seeking
di: Kayembe, Naomi Omeonga wa
Pubblicazione: (2025)
di: Kayembe, Naomi Omeonga wa
Pubblicazione: (2025)
Between Fear and Desire, the Monster Artificial Intelligence (AI): Analysis through the Lenses of Monster Theory
di: Tlili, Ahmed
Pubblicazione: (2025)
di: Tlili, Ahmed
Pubblicazione: (2025)
Personalized Knowledge Tracing through Student Representation Reconstruction and Class Imbalance Mitigation
di: Chen, Zhiyu, et al.
Pubblicazione: (2024)
di: Chen, Zhiyu, et al.
Pubblicazione: (2024)
Business and ethical concerns in domestic Conversational Generative AI-empowered multi-robot systems
di: Rousi, Rebekah, et al.
Pubblicazione: (2024)
di: Rousi, Rebekah, et al.
Pubblicazione: (2024)
Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety
di: Du, Y.
Pubblicazione: (2025)
di: Du, Y.
Pubblicazione: (2025)
Exploring Moral Exercises for Human Oversight of AI systems: Insights from Three Pilot Studies
di: Crafa, Silvia, et al.
Pubblicazione: (2025)
di: Crafa, Silvia, et al.
Pubblicazione: (2025)
Measuring and Mitigating Bias in Code Generated by Large Language Models
di: Chen, Yuxi, et al.
Pubblicazione: (2026)
di: Chen, Yuxi, et al.
Pubblicazione: (2026)
Recognising, Anticipating, and Mitigating LLM Pollution of Online Behavioural Research
di: Rilla, Raluca, et al.
Pubblicazione: (2025)
di: Rilla, Raluca, et al.
Pubblicazione: (2025)
Defining bias in AI-systems: Biased models are fair models
di: Lindloff, Chiara, et al.
Pubblicazione: (2025)
di: Lindloff, Chiara, et al.
Pubblicazione: (2025)
Documenti analoghi
-
An Aristotelian ontology of instrumental goals: Structural features to be managed and not failures to be eliminated
di: Fourie, Willem
Pubblicazione: (2025) -
Deciding how to respond: A deliberative framework to guide policymaker responses to AI systems
di: Fourie, Willem
Pubblicazione: (2025) -
How will advanced AI systems impact democracy?
di: Summerfield, Christopher, et al.
Pubblicazione: (2024) -
Mapping AI Risk Mitigations: Evidence Scan and Preliminary AI Risk Mitigation Taxonomy
di: Saeri, Alexander K., et al.
Pubblicazione: (2025) -
AI Identity, Empowerment, and Mindfulness in Mitigating Unethical AI Use
di: Shaayesteh, Mayssam Tarighi, et al.
Pubblicazione: (2025)