ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tien, Jeremy, Anand, Abishek, Tuan, Yu-Rou, Shen, Yuchen, Kolter, J. Zico, Nayebi, Aran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
von: Nayebi, Aran
Veröffentlicht: (2026)
von: Nayebi, Aran
Veröffentlicht: (2026)
Core Safety Values for Provably Corrigible Agents
von: Nayebi, Aran
Veröffentlicht: (2025)
von: Nayebi, Aran
Veröffentlicht: (2025)
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
von: Nayebi, Aran
Veröffentlicht: (2025)
von: Nayebi, Aran
Veröffentlicht: (2025)
Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain
von: Chung, Trinity, et al.
Veröffentlicht: (2025)
von: Chung, Trinity, et al.
Veröffentlicht: (2025)
Intrinsic Goals for Autonomous Agents: Model-Based Exploration in Virtual Zebrafish Predicts Ethological Behavior and Whole-Brain Dynamics
von: Keller, Reece, et al.
Veröffentlicht: (2025)
von: Keller, Reece, et al.
Veröffentlicht: (2025)
Mimetic Initialization of MLPs
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
Weight Ensembling Improves Reasoning in Language Models
von: Dang, Xingyu, et al.
Veröffentlicht: (2025)
von: Dang, Xingyu, et al.
Veröffentlicht: (2025)
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
von: Kim, Eungyeup, et al.
Veröffentlicht: (2023)
von: Kim, Eungyeup, et al.
Veröffentlicht: (2023)
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
von: Sokota, Samuel, et al.
Veröffentlicht: (2025)
von: Sokota, Samuel, et al.
Veröffentlicht: (2025)
Compute-Optimal LLMs Provably Generalize Better With Scale
von: Finzi, Marc, et al.
Veröffentlicht: (2025)
von: Finzi, Marc, et al.
Veröffentlicht: (2025)
Understanding Optimization in Deep Learning with Central Flows
von: Cohen, Jeremy M., et al.
Veröffentlicht: (2024)
von: Cohen, Jeremy M., et al.
Veröffentlicht: (2024)
A Simple and Effective Pruning Approach for Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Provably Bounding Neural Network Preimages
von: Kotha, Suhas, et al.
Veröffentlicht: (2023)
von: Kotha, Suhas, et al.
Veröffentlicht: (2023)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
Training a Generally Curious Agent
von: Tajwar, Fahim, et al.
Veröffentlicht: (2025)
von: Tajwar, Fahim, et al.
Veröffentlicht: (2025)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
von: Bick, Aviv, et al.
Veröffentlicht: (2024)
von: Bick, Aviv, et al.
Veröffentlicht: (2024)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
von: Duan, Xintong, et al.
Veröffentlicht: (2025)
Neural Network Verification with Branch-and-Bound for General Nonlinearities
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
Contextures: Representations from Contexts
von: Zhai, Runtian, et al.
Veröffentlicht: (2025)
von: Zhai, Runtian, et al.
Veröffentlicht: (2025)
An AI Capability Threshold for Rent-Funded Universal Basic Income in an AI-Automated Economy
von: Nayebi, Aran
Veröffentlicht: (2025)
von: Nayebi, Aran
Veröffentlicht: (2025)
Base Models Look Human To AI Detectors
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
Brain-Model Evaluations Need the NeuroAI Turing Test
von: Feather, Jenelle, et al.
Veröffentlicht: (2025)
von: Feather, Jenelle, et al.
Veröffentlicht: (2025)
Ordinary Least Squares is a Special Case of Transformer
von: Tan, Xiaojun, et al.
Veröffentlicht: (2026)
von: Tan, Xiaojun, et al.
Veröffentlicht: (2026)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
von: Li, Kevin Y., et al.
Veröffentlicht: (2024)
von: Li, Kevin Y., et al.
Veröffentlicht: (2024)
The Update-Equivalence Framework for Decision-Time Planning
von: Sokota, Samuel, et al.
Veröffentlicht: (2023)
von: Sokota, Samuel, et al.
Veröffentlicht: (2023)
One-Step Diffusion Distillation through Score Implicit Matching
von: Luo, Weijian, et al.
Veröffentlicht: (2024)
von: Luo, Weijian, et al.
Veröffentlicht: (2024)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
von: Arturi, Daniel Aarao Reis, et al.
Veröffentlicht: (2025)
von: Arturi, Daniel Aarao Reis, et al.
Veröffentlicht: (2025)
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
von: Arnold, Julian, et al.
Veröffentlicht: (2025)
von: Arnold, Julian, et al.
Veröffentlicht: (2025)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
von: Wang, Bowen, et al.
Veröffentlicht: (2026)
von: Wang, Bowen, et al.
Veröffentlicht: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
DeepHGNN: Study of Graph Neural Network based Forecasting Methods for Hierarchically Related Multivariate Time Series
von: Sriramulu, Abishek, et al.
Veröffentlicht: (2024)
von: Sriramulu, Abishek, et al.
Veröffentlicht: (2024)
Context Neural Networks: A Scalable Multivariate Model for Time Series Forecasting
von: Sriramulu, Abishek, et al.
Veröffentlicht: (2024)
von: Sriramulu, Abishek, et al.
Veröffentlicht: (2024)
Grounding Computer Use Agents on Human Demonstrations
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
Blind Inverse Problem Solving Made Easy by Text-to-Image Latent Diffusion
von: Dontas, Michail, et al.
Veröffentlicht: (2024)
von: Dontas, Michail, et al.
Veröffentlicht: (2024)
Overtrained, Not Misaligned
von: Schreiber, Joel, et al.
Veröffentlicht: (2026)
von: Schreiber, Joel, et al.
Veröffentlicht: (2026)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
von: Lu, Dunjie, et al.
Veröffentlicht: (2025)
von: Lu, Dunjie, et al.
Veröffentlicht: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Antidistillation Fingerprinting
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
Model Organisms for Emergent Misalignment
von: Turner, Edward, et al.
Veröffentlicht: (2025)
von: Turner, Edward, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
von: Nayebi, Aran
Veröffentlicht: (2026) -
Core Safety Values for Provably Corrigible Agents
von: Nayebi, Aran
Veröffentlicht: (2025) -
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
von: Nayebi, Aran
Veröffentlicht: (2025) -
Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain
von: Chung, Trinity, et al.
Veröffentlicht: (2025) -
Intrinsic Goals for Autonomous Agents: Model-Based Exploration in Virtual Zebrafish Predicts Ethological Behavior and Whole-Brain Dynamics
von: Keller, Reece, et al.
Veröffentlicht: (2025)