Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | DiSorbo, Matthew DosSantos, Ju, Harang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Teaching AI to Handle Exceptions: Supervised Fine-Tuning with Human-Aligned Judgment
by: DiSorbo, Matthew DosSantos, et al.
Published: (2025)
by: DiSorbo, Matthew DosSantos, et al.
Published: (2025)
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
by: Bouchard, Dylan
Published: (2026)
by: Bouchard, Dylan
Published: (2026)
Early Prediction of Multi-Label Care Escalation Triggers in the Intensive Care Unit Using Electronic Health Records
by: Bukhari, Syed Ahmad Chan, et al.
Published: (2025)
by: Bukhari, Syed Ahmad Chan, et al.
Published: (2025)
Managing Escalation in Off-the-Shelf Large Language Models
by: Elbaum, Sebastian, et al.
Published: (2025)
by: Elbaum, Sebastian, et al.
Published: (2025)
SMolLM: Small Language Models Learn Small Molecular Grammar
by: Jindal, Akhil, et al.
Published: (2026)
by: Jindal, Akhil, et al.
Published: (2026)
FAMOSE: A ReAct Approach to Automated Feature Discovery
by: Burghardt, Keith, et al.
Published: (2026)
by: Burghardt, Keith, et al.
Published: (2026)
Learning When to Act: Interval-Aware Reinforcement Learning with Predictive Temporal Structure
by: Di Gioia, Davide
Published: (2026)
by: Di Gioia, Davide
Published: (2026)
Towards Automated Semantic Interpretability in Reinforcement Learning via Vision-Language Models
by: Li, Zhaoxin, et al.
Published: (2025)
by: Li, Zhaoxin, et al.
Published: (2025)
Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models
by: Wu, Zhanglin, et al.
Published: (2025)
by: Wu, Zhanglin, et al.
Published: (2025)
Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
by: Li, Kunhao, et al.
Published: (2025)
by: Li, Kunhao, et al.
Published: (2025)
Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models
by: Ciaperoni, Martino, et al.
Published: (2026)
by: Ciaperoni, Martino, et al.
Published: (2026)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
by: Biskupski, Tom, et al.
Published: (2026)
by: Biskupski, Tom, et al.
Published: (2026)
Leveraging Language Models for Automated Patient Record Linkage
by: Beheshti, Mohammad, et al.
Published: (2025)
by: Beheshti, Mohammad, et al.
Published: (2025)
Large Language Model Agent for Structural Drawing Generation Using ReAct Prompt Engineering and Retrieval Augmented Generation
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Can Language Models Explain Their Own Classification Behavior?
by: Sherburn, Dane, et al.
Published: (2024)
by: Sherburn, Dane, et al.
Published: (2024)
Behavior Injection: Preparing Language Models for Reinforcement Learning
by: Cen, Zhepeng, et al.
Published: (2025)
by: Cen, Zhepeng, et al.
Published: (2025)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
by: Rivera, Juan-Pablo, et al.
Published: (2024)
by: Rivera, Juan-Pablo, et al.
Published: (2024)
Automated Attention Pattern Discovery at Scale in Large Language Models
by: Katzy, Jonathan, et al.
Published: (2026)
by: Katzy, Jonathan, et al.
Published: (2026)
Fine-tuning Large Language Model for Automated Algorithm Design
by: Liu, Fei, et al.
Published: (2025)
by: Liu, Fei, et al.
Published: (2025)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
by: Yang, Zonghan, et al.
Published: (2024)
by: Yang, Zonghan, et al.
Published: (2024)
Learning to Act without Actions
by: Schmidt, Dominik, et al.
Published: (2023)
by: Schmidt, Dominik, et al.
Published: (2023)
LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
by: Happe, Andreas, et al.
Published: (2023)
by: Happe, Andreas, et al.
Published: (2023)
The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
by: Lu, Yao, et al.
Published: (2025)
by: Lu, Yao, et al.
Published: (2025)
The Role of Emotional Stimuli and Intensity in Shaping Large Language Model Behavior
by: Patel, Ameen, et al.
Published: (2026)
by: Patel, Ameen, et al.
Published: (2026)
Automated Federated Pipeline for Parameter-Efficient Fine-Tuning of Large Language Models
by: Fang, Zihan, et al.
Published: (2024)
by: Fang, Zihan, et al.
Published: (2024)
Embracing Large Language Models in Traffic Flow Forecasting
by: Zhao, Yusheng, et al.
Published: (2024)
by: Zhao, Yusheng, et al.
Published: (2024)
Getting out of the Big-Muddy: Escalation of Commitment in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Adaptive Sparsified Graph Learning Framework for Vessel Behavior Anomalies
by: Kim, Jeehong, et al.
Published: (2025)
by: Kim, Jeehong, et al.
Published: (2025)
Goal Recognition Design for General Behavioral Agents using Machine Learning
by: Kasumba, Robert, et al.
Published: (2024)
by: Kasumba, Robert, et al.
Published: (2024)
Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance
by: Ju, Harang, et al.
Published: (2025)
by: Ju, Harang, et al.
Published: (2025)
EvoEngineer: Mastering Automated CUDA Kernel Code Evolution with Large Language Models
by: Guo, Ping, et al.
Published: (2025)
by: Guo, Ping, et al.
Published: (2025)
An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications
by: Garrido-Merchán, Eduardo C., et al.
Published: (2023)
by: Garrido-Merchán, Eduardo C., et al.
Published: (2023)
EEGAgent: A Unified Framework for Automated EEG Analysis Using Large Language Models
by: Zhao, Sha, et al.
Published: (2025)
by: Zhao, Sha, et al.
Published: (2025)
Automating Forecasting Question Generation and Resolution for AI Evaluation
by: Bosse, Nikos I., et al.
Published: (2026)
by: Bosse, Nikos I., et al.
Published: (2026)
SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation
by: Wang, Haochun, et al.
Published: (2026)
by: Wang, Haochun, et al.
Published: (2026)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
by: Hazra, Rishi, et al.
Published: (2025)
by: Hazra, Rishi, et al.
Published: (2025)
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
CriticAL: Critic Automation with Language Models
by: Li, Michael Y., et al.
Published: (2024)
by: Li, Michael Y., et al.
Published: (2024)
Leveraging Foundation Language Models (FLMs) for Automated Cohort Extraction from Large EHR Databases
by: Mugambi, Purity, et al.
Published: (2024)
by: Mugambi, Purity, et al.
Published: (2024)
An Automated Reinforcement Learning Reward Design Framework with Large Language Model for Cooperative Platoon Coordination
by: Wei, Dixiao, et al.
Published: (2025)
by: Wei, Dixiao, et al.
Published: (2025)
Similar Items
-
Teaching AI to Handle Exceptions: Supervised Fine-Tuning with Human-Aligned Judgment
by: DiSorbo, Matthew DosSantos, et al.
Published: (2025) -
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
by: Bouchard, Dylan
Published: (2026) -
Early Prediction of Multi-Label Care Escalation Triggers in the Intensive Care Unit Using Electronic Health Records
by: Bukhari, Syed Ahmad Chan, et al.
Published: (2025) -
Managing Escalation in Off-the-Shelf Large Language Models
by: Elbaum, Sebastian, et al.
Published: (2025) -
SMolLM: Small Language Models Learn Small Molecular Grammar
by: Jindal, Akhil, et al.
Published: (2026)