Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Schlatter, Jeremy, Weinstein-Raun, Benjamin, Ladish, Jeffrey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
Quantifying and Mitigating Premature Closure in Frontier LLMs
by: Handler, Rebecca, et al.
Published: (2026)
by: Handler, Rebecca, et al.
Published: (2026)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
by: Sirdeshmukh, Ved, et al.
Published: (2025)
by: Sirdeshmukh, Ved, et al.
Published: (2025)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks
by: Dunivin, Zackary Okun
Published: (2024)
by: Dunivin, Zackary Okun
Published: (2024)
Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks
by: Mohanty, Dikshya, et al.
Published: (2026)
by: Mohanty, Dikshya, et al.
Published: (2026)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
by: Chi, Yizhe, et al.
Published: (2026)
by: Chi, Yizhe, et al.
Published: (2026)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
by: Zhao, Sihang, et al.
Published: (2024)
by: Zhao, Sihang, et al.
Published: (2024)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
by: Maurya, Sneha, et al.
Published: (2026)
by: Maurya, Sneha, et al.
Published: (2026)
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
Two-stage Incomplete Utterance Rewriting on Editing Operation
by: Cao, Zhiyu, et al.
Published: (2025)
by: Cao, Zhiyu, et al.
Published: (2025)
Position: Avoid Overstretching LLMs for every Enterprise Task
by: Singh, Kuldeep, et al.
Published: (2026)
by: Singh, Kuldeep, et al.
Published: (2026)
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
by: Ajjour, Yamen, et al.
Published: (2026)
by: Ajjour, Yamen, et al.
Published: (2026)
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
by: Karia, Rushang, et al.
Published: (2024)
by: Karia, Rushang, et al.
Published: (2024)
Are Long-LLMs A Necessity For Long-Context Tasks?
by: Qian, Hongjin, et al.
Published: (2024)
by: Qian, Hongjin, et al.
Published: (2024)
Conditioning LLMs to Generate Code-Switched Text
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Incomplete Utterance Rewriting with Editing Operation Guidance and Utterance Augmentation
by: Cao, Zhiyu, et al.
Published: (2025)
by: Cao, Zhiyu, et al.
Published: (2025)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
by: Arabelly, Abhinav, et al.
Published: (2025)
by: Arabelly, Abhinav, et al.
Published: (2025)
Can LLMs Generate High-Quality Task-Specific Conversations?
by: Li, Shengqi, et al.
Published: (2025)
by: Li, Shengqi, et al.
Published: (2025)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
by: Fouda, Aya E., et al.
Published: (2025)
by: Fouda, Aya E., et al.
Published: (2025)
SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
by: Wu, Mengsong, et al.
Published: (2025)
by: Wu, Mengsong, et al.
Published: (2025)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
Multi-Task Learning with LLMs for Implicit Sentiment Analysis: Data-level and Task-level Automatic Weight Learning
by: Lai, Wenna, et al.
Published: (2024)
by: Lai, Wenna, et al.
Published: (2024)
Early Signs of Steganographic Capabilities in Frontier LLMs
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
RSMLP: A light Sampled MLP Structure for Incomplete Utterance Rewrite
by: Liu, Lunjun, et al.
Published: (2025)
by: Liu, Lunjun, et al.
Published: (2025)
An Empirical Study of the Role of Incompleteness and Ambiguity in Interactions with Large Language Models
by: Naik, Riya, et al.
Published: (2025)
by: Naik, Riya, et al.
Published: (2025)
Structured Thinking Matters: Improving LLMs Generalization in Causal Inference Tasks
by: Sun, Wentao, et al.
Published: (2025)
by: Sun, Wentao, et al.
Published: (2025)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
by: Fei, Zhaoye, et al.
Published: (2025)
by: Fei, Zhaoye, et al.
Published: (2025)
Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks
by: Gong, Chang, et al.
Published: (2025)
by: Gong, Chang, et al.
Published: (2025)
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
by: Wu, Xinwei, et al.
Published: (2026)
by: Wu, Xinwei, et al.
Published: (2026)
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
by: Lin, Peiqin, et al.
Published: (2026)
by: Lin, Peiqin, et al.
Published: (2026)
Active Task Disambiguation with LLMs
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
Bottom-Up and Top-Down Analysis of Values, Agendas, and Observations in Corpora and LLMs
by: Friedman, Scott E., et al.
Published: (2024)
by: Friedman, Scott E., et al.
Published: (2024)
Similar Items
-
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025) -
Quantifying and Mitigating Premature Closure in Frontier LLMs
by: Handler, Rebecca, et al.
Published: (2026) -
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
by: Sirdeshmukh, Ved, et al.
Published: (2025) -
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025) -
Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks
by: Dunivin, Zackary Okun
Published: (2024)