The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
Fuente:
arXiv
Salvato in:
| Autori principali: | Hägele, Alexander, Gema, Aryo Pradipta, Sleight, Henry, Perez, Ethan, Sohl-Dickstein, Jascha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AI Organizations are More Effective but Less Aligned than Individual Agents
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2026)
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2026)
Inverse Scaling in Test-Time Compute
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025)
The boundary of neural network trainability is fractal
di: Sohl-Dickstein, Jascha
Pubblicazione: (2024)
di: Sohl-Dickstein, Jascha
Pubblicazione: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
di: Kirsch, Louis, et al.
Pubblicazione: (2022)
di: Kirsch, Louis, et al.
Pubblicazione: (2022)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
di: Saxena, Rohit, et al.
Pubblicazione: (2025)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
di: Rajani, Neel, et al.
Pubblicazione: (2025)
di: Rajani, Neel, et al.
Pubblicazione: (2025)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
Levels of AGI for Operationalizing Progress on the Path to AGI
di: Morris, Meredith Ringel, et al.
Pubblicazione: (2023)
di: Morris, Meredith Ringel, et al.
Pubblicazione: (2023)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2024)
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2024)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2025)
di: Madani, Mohammad Reza Ghasemi, et al.
Pubblicazione: (2025)
Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC
di: Du, Yilun, et al.
Pubblicazione: (2023)
di: Du, Yilun, et al.
Pubblicazione: (2023)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
di: Murphy, Alexander, et al.
Pubblicazione: (2025)
GRADA: Graph-based Reranking against Adversarial Documents Attack
di: Zheng, Jingjie, et al.
Pubblicazione: (2025)
di: Zheng, Jingjie, et al.
Pubblicazione: (2025)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2024)
MessIRve: A Large-Scale Spanish Information Retrieval Dataset
di: Valentini, Francisco, et al.
Pubblicazione: (2024)
di: Valentini, Francisco, et al.
Pubblicazione: (2024)
A Comparative Study on Patient Language across Therapeutic Domains for Effective Patient Voice Classification in Online Health Discussions
di: Lysandrou, Giorgos, et al.
Pubblicazione: (2024)
di: Lysandrou, Giorgos, et al.
Pubblicazione: (2024)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
di: Guo, Shiyuan, et al.
Pubblicazione: (2025)
di: Guo, Shiyuan, et al.
Pubblicazione: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
di: Ensign, Danielle, et al.
Pubblicazione: (2025)
di: Ensign, Danielle, et al.
Pubblicazione: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
di: Lynch, Aengus, et al.
Pubblicazione: (2025)
di: Lynch, Aengus, et al.
Pubblicazione: (2025)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
di: Mahrooghi, Ilia, et al.
Pubblicazione: (2026)
di: Mahrooghi, Ilia, et al.
Pubblicazione: (2026)
Evaluating Control Protocols for Untrusted AI Agents
di: Kutasov, Jon, et al.
Pubblicazione: (2025)
di: Kutasov, Jon, et al.
Pubblicazione: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
di: Binder, Felix J, et al.
Pubblicazione: (2024)
di: Binder, Felix J, et al.
Pubblicazione: (2024)
Stress-Testing Model Specs Reveals Character Differences among Language Models
di: Zhang, Jifan, et al.
Pubblicazione: (2025)
di: Zhang, Jifan, et al.
Pubblicazione: (2025)
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
di: Slocum, Stewart, et al.
Pubblicazione: (2025)
di: Slocum, Stewart, et al.
Pubblicazione: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
Intelligence Analysis of Language Models
di: Galanti, Liane, et al.
Pubblicazione: (2024)
di: Galanti, Liane, et al.
Pubblicazione: (2024)
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2025)
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
di: Dandamudi, Rohit, et al.
Pubblicazione: (2024)
di: Dandamudi, Rohit, et al.
Pubblicazione: (2024)
Does Self-Evaluation Enable Wireheading in Language Models?
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
di: Africa, David Demitri, et al.
Pubblicazione: (2025)
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
di: Turtayev, Rustem, et al.
Pubblicazione: (2025)
di: Turtayev, Rustem, et al.
Pubblicazione: (2025)
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
di: Abbe, Emmanuel, et al.
Pubblicazione: (2024)
di: Abbe, Emmanuel, et al.
Pubblicazione: (2024)
How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
di: Zhang, Shan, et al.
Pubblicazione: (2026)
di: Zhang, Shan, et al.
Pubblicazione: (2026)
Towards Avoiding the Data Mess: Industry Insights from Data Mesh Implementations
di: Bode, Jan, et al.
Pubblicazione: (2023)
di: Bode, Jan, et al.
Pubblicazione: (2023)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025)
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025)
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
di: Shilov, Igor, et al.
Pubblicazione: (2025)
di: Shilov, Igor, et al.
Pubblicazione: (2025)
A Sober Look at Agentic Misalignment in Automated Workflows
di: Ye, Wenqian, et al.
Pubblicazione: (2026)
di: Ye, Wenqian, et al.
Pubblicazione: (2026)
Unsupervised Elicitation of Language Models
di: Wen, Jiaxin, et al.
Pubblicazione: (2025)
di: Wen, Jiaxin, et al.
Pubblicazione: (2025)
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
di: Kim, Sejin, et al.
Pubblicazione: (2024)
di: Kim, Sejin, et al.
Pubblicazione: (2024)
How Does Naming Affect LLMs on Code Analysis Tasks?
di: Wang, Zhilong, et al.
Pubblicazione: (2023)
di: Wang, Zhilong, et al.
Pubblicazione: (2023)
AI Scientist via Synthetic Task Scaling
di: Cai, Ziyang, et al.
Pubblicazione: (2026)
di: Cai, Ziyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AI Organizations are More Effective but Less Aligned than Individual Agents
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2026) -
Inverse Scaling in Test-Time Compute
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2025) -
The boundary of neural network trainability is fractal
di: Sohl-Dickstein, Jascha
Pubblicazione: (2024) -
General-Purpose In-Context Learning by Meta-Learning Transformers
di: Kirsch, Louis, et al.
Pubblicazione: (2022) -
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
di: Saxena, Rohit, et al.
Pubblicazione: (2025)