Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence
Fuente:
arXiv
Salvato in:
| Autori principali: | Shane, Tommy Shaffer, Mylius, Simon, Hobbs, Hamish |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AI incidents and 'networked trouble': The case for a research agenda
di: Shane, Tommy Shaffer
Pubblicazione: (2024)
di: Shane, Tommy Shaffer
Pubblicazione: (2024)
Systematic Hazard Analysis for Frontier AI using STPA
di: Mylius, Simon
Pubblicazione: (2025)
di: Mylius, Simon
Pubblicazione: (2025)
Berta: an open-source, modular tool for AI-enabled clinical documentation
di: Vaid, Samridhi, et al.
Pubblicazione: (2026)
di: Vaid, Samridhi, et al.
Pubblicazione: (2026)
Can an AI-Powered Presentation Platform Based On The Game "Just a Minute" Be Used To Improve Students' Public Speaking Skills?
di: Higham, Frederic, et al.
Pubblicazione: (2025)
di: Higham, Frederic, et al.
Pubblicazione: (2025)
Introducing the A2AJ's Canadian Legal Data: An open-source alternative to CanLII for the era of computational law
di: Wallace, Simon, et al.
Pubblicazione: (2025)
di: Wallace, Simon, et al.
Pubblicazione: (2025)
Navigating the EU AI Act: Foreseeable Challenges in Qualifying Deep Learning-Based Automated Inspections of Class III Medical Devices
di: Diaz, Julio Zanon, et al.
Pubblicazione: (2025)
di: Diaz, Julio Zanon, et al.
Pubblicazione: (2025)
AI threats to national security can be countered through an incident regime
di: Ortega, Alejandro
Pubblicazione: (2025)
di: Ortega, Alejandro
Pubblicazione: (2025)
A pragmatic classification framework for AI incident monitoring
di: Mengesha, Isaak, et al.
Pubblicazione: (2026)
di: Mengesha, Isaak, et al.
Pubblicazione: (2026)
Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds
di: Gomez, Francesca, et al.
Pubblicazione: (2026)
di: Gomez, Francesca, et al.
Pubblicazione: (2026)
AI for All: Identifying AI incidents Related to Diversity and Inclusion
di: Shams, Rifat Ara, et al.
Pubblicazione: (2024)
di: Shams, Rifat Ara, et al.
Pubblicazione: (2024)
Teacher agency in the age of generative AI: towards a framework of hybrid intelligence for learning design
di: Frøsig, Thomas B, et al.
Pubblicazione: (2024)
di: Frøsig, Thomas B, et al.
Pubblicazione: (2024)
Standardised schema and taxonomy for AI incident databases in critical digital infrastructure
di: Agarwal, Avinash, et al.
Pubblicazione: (2025)
di: Agarwal, Avinash, et al.
Pubblicazione: (2025)
Incorporating AI incident reporting into telecommunications law and policy: Insights from India
di: Agarwal, Avinash, et al.
Pubblicazione: (2025)
di: Agarwal, Avinash, et al.
Pubblicazione: (2025)
AI Thinking: A framework for rethinking artificial intelligence in practice
di: Newman-Griffis, Denis
Pubblicazione: (2024)
di: Newman-Griffis, Denis
Pubblicazione: (2024)
Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
Towards a Healthy AI Tradition: Lessons from Biology and Biomedical Science
di: Kasif, Simon
Pubblicazione: (2024)
di: Kasif, Simon
Pubblicazione: (2024)
Comparative analysis of privacy-preserving open-source LLMs regarding extraction of diagnostic information from clinical CMR imaging reports
di: Amirrajab, Sina, et al.
Pubblicazione: (2025)
di: Amirrajab, Sina, et al.
Pubblicazione: (2025)
Fundamentals of legislation for autonomous artificial intelligence systems
di: Romanova, Anna
Pubblicazione: (2024)
di: Romanova, Anna
Pubblicazione: (2024)
Promises and pitfalls of artificial intelligence for legal applications
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)
di: Kapoor, Sayash, et al.
Pubblicazione: (2024)
The environmental impact of ICT in the era of data and artificial intelligence
di: Rottenberg, François, et al.
Pubblicazione: (2026)
di: Rottenberg, François, et al.
Pubblicazione: (2026)
Artificial intelligence and machine learning applications for cultured meat
di: Todhunter, Michael E., et al.
Pubblicazione: (2024)
di: Todhunter, Michael E., et al.
Pubblicazione: (2024)
Academ-AI: documenting the undisclosed use of generative artificial intelligence in academic publishing
di: Glynn, Alex
Pubblicazione: (2024)
di: Glynn, Alex
Pubblicazione: (2024)
Future progress in artificial intelligence: A survey of expert opinion
di: Müller, Vincent C., et al.
Pubblicazione: (2025)
di: Müller, Vincent C., et al.
Pubblicazione: (2025)
An AI-driven framework for rapid and localized optimizations of urban open spaces
di: Eshraghi, Pegah, et al.
Pubblicazione: (2025)
di: Eshraghi, Pegah, et al.
Pubblicazione: (2025)
Artificially intelligent agents in the social and behavioral sciences: A history and outlook
di: Holme, Petter, et al.
Pubblicazione: (2025)
di: Holme, Petter, et al.
Pubblicazione: (2025)
What we learned while automating bias detection in AI hiring systems for compliance with NYC Local Law 144
di: Clavell, Gemma Galdon, et al.
Pubblicazione: (2024)
di: Clavell, Gemma Galdon, et al.
Pubblicazione: (2024)
A governance horizon for ethical-use constraints in open-weight AI models
di: Xu, Weiwei, et al.
Pubblicazione: (2026)
di: Xu, Weiwei, et al.
Pubblicazione: (2026)
How well are open sourced AI-generated image detection models out-of-the-box: A comprehensive benchmark study
di: Ren, Simiao, et al.
Pubblicazione: (2026)
di: Ren, Simiao, et al.
Pubblicazione: (2026)
What Makes AI Applications Acceptable or Unacceptable? A Predictive Moral Framework
di: Eriksson, Kimmo, et al.
Pubblicazione: (2025)
di: Eriksson, Kimmo, et al.
Pubblicazione: (2025)
Generative artificial intelligence and the marginalization of minoritized knowledges in higher education: the case of disability
di: Tali-Otmani, Fatiha
Pubblicazione: (2026)
di: Tali-Otmani, Fatiha
Pubblicazione: (2026)
Artificial intelligence, rationalization, and the limits of control in the public sector: the case of tax policy optimization
di: Mokander, Jakob, et al.
Pubblicazione: (2024)
di: Mokander, Jakob, et al.
Pubblicazione: (2024)
ELIZA Reanimated: The world's first chatbot restored on the world's first time sharing system
di: Lane, Rupert, et al.
Pubblicazione: (2025)
di: Lane, Rupert, et al.
Pubblicazione: (2025)
LLM hallucinations in the wild: Large-scale evidence from non-existent citations
di: Zhao, Zhenyue, et al.
Pubblicazione: (2026)
di: Zhao, Zhenyue, et al.
Pubblicazione: (2026)
Increasing intelligence in AI agents can worsen collective outcomes
di: Johnson, Neil F.
Pubblicazione: (2026)
di: Johnson, Neil F.
Pubblicazione: (2026)
The data heat island effect: quantifying the impact of AI data centers in a warming world
di: Marinoni, Andrea, et al.
Pubblicazione: (2026)
di: Marinoni, Andrea, et al.
Pubblicazione: (2026)
Improving prediction of students' performance in intelligent tutoring systems using attribute selection and ensembles of different multimodal data sources
di: Chango, W., et al.
Pubblicazione: (2024)
di: Chango, W., et al.
Pubblicazione: (2024)
Artificial intelligence-driven improvement of hospital logistics management resilience: a practical exploration based on H Hospital
di: Huang, Lu, et al.
Pubblicazione: (2026)
di: Huang, Lu, et al.
Pubblicazione: (2026)
Artificial intelligence and the internal processes of creativity
di: Aru, Jaan
Pubblicazione: (2024)
di: Aru, Jaan
Pubblicazione: (2024)
The changing surface of the world's roads
di: Randhawa, Sukanya, et al.
Pubblicazione: (2025)
di: Randhawa, Sukanya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AI incidents and 'networked trouble': The case for a research agenda
di: Shane, Tommy Shaffer
Pubblicazione: (2024) -
Systematic Hazard Analysis for Frontier AI using STPA
di: Mylius, Simon
Pubblicazione: (2025) -
Berta: an open-source, modular tool for AI-enabled clinical documentation
di: Vaid, Samridhi, et al.
Pubblicazione: (2026) -
Can an AI-Powered Presentation Platform Based On The Game "Just a Minute" Be Used To Improve Students' Public Speaking Skills?
di: Higham, Frederic, et al.
Pubblicazione: (2025) -
Introducing the A2AJ's Canadian Legal Data: An open-source alternative to CanLII for the era of computational law
di: Wallace, Simon, et al.
Pubblicazione: (2025)