When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Khadangi, Afshin, Marxen, Hanna, Sartipi, Amir, Tchappi, Igor, Fridgen, Gilbert |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CognArtive: Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
Noise Augmented Fine Tuning for Mitigating Hallucinations in Large Language Models
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
Towards Effective E-Participation of Citizens in the European Union: The Development of AskThePublic
di: Messerschmidt, Nils, et al.
Pubblicazione: (2025)
di: Messerschmidt, Nils, et al.
Pubblicazione: (2025)
Benchmarking Pre-Trained Time Series Models for Electricity Price Forecasting
di: Sartipi, Timothée Hornek Amir, et al.
Pubblicazione: (2025)
di: Sartipi, Timothée Hornek Amir, et al.
Pubblicazione: (2025)
Evaluating General-Purpose AI with Psychometrics
di: Wang, Xiting, et al.
Pubblicazione: (2023)
di: Wang, Xiting, et al.
Pubblicazione: (2023)
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
di: Vishwarupe, Varad, et al.
Pubblicazione: (2026)
di: Vishwarupe, Varad, et al.
Pubblicazione: (2026)
KG-HTC: Integrating Knowledge Graphs into LLMs for Effective Zero-shot Hierarchical Text Classification
di: Zang, Qianbo, et al.
Pubblicazione: (2025)
di: Zang, Qianbo, et al.
Pubblicazione: (2025)
Trends in Frontier AI Model Count: A Forecast to 2028
di: Kumar, Iyngkarran, et al.
Pubblicazione: (2025)
di: Kumar, Iyngkarran, et al.
Pubblicazione: (2025)
Risk Reporting for Developers' Internal AI Model Use
di: Delaney, Oscar, et al.
Pubblicazione: (2026)
di: Delaney, Oscar, et al.
Pubblicazione: (2026)
Designing AI-Agents with Personalities: A Psychometric Approach
di: Huang, Muhua, et al.
Pubblicazione: (2024)
di: Huang, Muhua, et al.
Pubblicazione: (2024)
Responsible Reporting for Frontier AI Development
di: Kolt, Noam, et al.
Pubblicazione: (2024)
di: Kolt, Noam, et al.
Pubblicazione: (2024)
Governing AI Beyond the Pretraining Frontier
di: Caputo, Nicholas A.
Pubblicazione: (2025)
di: Caputo, Nicholas A.
Pubblicazione: (2025)
Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
di: Creo, Aldan, et al.
Pubblicazione: (2025)
di: Creo, Aldan, et al.
Pubblicazione: (2025)
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents
di: Lazar, Seth
Pubblicazione: (2024)
di: Lazar, Seth
Pubblicazione: (2024)
Taking AI Welfare Seriously
di: Long, Robert, et al.
Pubblicazione: (2024)
di: Long, Robert, et al.
Pubblicazione: (2024)
Towards Safe Multilingual Frontier AI
di: Kanepajs, Artūrs, et al.
Pubblicazione: (2024)
di: Kanepajs, Artūrs, et al.
Pubblicazione: (2024)
The Architecture of AI Transformation: Four Strategic Patterns and an Emerging Frontier
di: Wolfe, Diana A., et al.
Pubblicazione: (2025)
di: Wolfe, Diana A., et al.
Pubblicazione: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI
di: Wilson, Sophia N., et al.
Pubblicazione: (2026)
di: Wilson, Sophia N., et al.
Pubblicazione: (2026)
Exploring the Potential of Machine Translation for Generating Named Entity Datasets: A Case Study between Persian and English
di: Sartipi, Amir, et al.
Pubblicazione: (2023)
di: Sartipi, Amir, et al.
Pubblicazione: (2023)
Alignment, Agency and Autonomy in Frontier AI: A Systems Engineering Perspective
di: Tallam, Krti
Pubblicazione: (2025)
di: Tallam, Krti
Pubblicazione: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
Will AI Take My Job? Evolving Perceptions of Automation and Labor Risk in Latin America
di: Cremaschi, Andrea, et al.
Pubblicazione: (2025)
di: Cremaschi, Andrea, et al.
Pubblicazione: (2025)
Dark Speculation: Combining Qualitative and Quantitative Understanding in Frontier AI Risk Analysis
di: Carpenter, Daniel, et al.
Pubblicazione: (2025)
di: Carpenter, Daniel, et al.
Pubblicazione: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
di: Potter, Yujin, et al.
Pubblicazione: (2025)
di: Potter, Yujin, et al.
Pubblicazione: (2025)
Exploring the Psychometric Validity of AI-Generated Student Responses: A Study on Virtual Personas' Learning Motivation
di: Wang, Huanxiao
Pubblicazione: (2025)
di: Wang, Huanxiao
Pubblicazione: (2025)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
di: Wu, Addison J., et al.
Pubblicazione: (2026)
di: Wu, Addison J., et al.
Pubblicazione: (2026)
Are Companies Taking AI Risks Seriously? A Systematic Analysis of Companies' AI Risk Disclosures in SEC 10-K forms
di: Marin, Lucas G. Uberti-Bona, et al.
Pubblicazione: (2025)
di: Marin, Lucas G. Uberti-Bona, et al.
Pubblicazione: (2025)
When is using AI the rational choice? The importance of counterfactuals in AI deployment decisions
di: Lehner, Paul, et al.
Pubblicazione: (2025)
di: Lehner, Paul, et al.
Pubblicazione: (2025)
Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety
di: Du, Y.
Pubblicazione: (2025)
di: Du, Y.
Pubblicazione: (2025)
Between Innovation and Oversight: A Cross-Regional Study of AI Risk Management Frameworks in the EU, U.S., UK, and China
di: Al-Maamari, Amir
Pubblicazione: (2025)
di: Al-Maamari, Amir
Pubblicazione: (2025)
When AI Navigates the Fog of War
di: Li, Ming, et al.
Pubblicazione: (2026)
di: Li, Ming, et al.
Pubblicazione: (2026)
Systematic Hazard Analysis for Frontier AI using STPA
di: Mylius, Simon
Pubblicazione: (2025)
di: Mylius, Simon
Pubblicazione: (2025)
When Should Algorithms Resign? A Proposal for AI Governance
di: Bhatt, Umang, et al.
Pubblicazione: (2024)
di: Bhatt, Umang, et al.
Pubblicazione: (2024)
AcademiClaw: When Students Set Challenges for AI Agents
di: Yu, Junjie, et al.
Pubblicazione: (2026)
di: Yu, Junjie, et al.
Pubblicazione: (2026)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
di: Gringras, David, et al.
Pubblicazione: (2026)
di: Gringras, David, et al.
Pubblicazione: (2026)
Sabotage Evaluations for Frontier Models
di: Benton, Joe, et al.
Pubblicazione: (2024)
di: Benton, Joe, et al.
Pubblicazione: (2024)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
di: Ackerman, Gary, et al.
Pubblicazione: (2025)
di: Ackerman, Gary, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CognArtive: Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements
di: Khadangi, Afshin, et al.
Pubblicazione: (2025) -
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
di: Khadangi, Afshin, et al.
Pubblicazione: (2025) -
Noise Augmented Fine Tuning for Mitigating Hallucinations in Large Language Models
di: Khadangi, Afshin, et al.
Pubblicazione: (2025) -
Towards Effective E-Participation of Citizens in the European Union: The Development of AskThePublic
di: Messerschmidt, Nils, et al.
Pubblicazione: (2025) -
Benchmarking Pre-Trained Time Series Models for Electricity Price Forecasting
di: Sartipi, Timothée Hornek Amir, et al.
Pubblicazione: (2025)