Technical Report: Evaluating Goal Drift in Language Model Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Arike, Rauno, Donoway, Elizabeth, Bartsch, Henning, Hobbhahn, Marius |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023)
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023)
Large Language Models Often Know When They Are Being Evaluated
di: Needham, Joe, et al.
Pubblicazione: (2025)
di: Needham, Joe, et al.
Pubblicazione: (2025)
Excess Description Length of Learning Generalizable Predictors
di: Donoway, Elizabeth, et al.
Pubblicazione: (2026)
di: Donoway, Elizabeth, et al.
Pubblicazione: (2026)
Frontier Models are Capable of In-context Scheming
di: Meinke, Alexander, et al.
Pubblicazione: (2024)
di: Meinke, Alexander, et al.
Pubblicazione: (2024)
Interpreting Learned Feedback Patterns in Large Language Models
di: Marks, Luke, et al.
Pubblicazione: (2023)
di: Marks, Luke, et al.
Pubblicazione: (2023)
Evaluating the Goal-Directedness of Large Language Models
di: Everitt, Tom, et al.
Pubblicazione: (2025)
di: Everitt, Tom, et al.
Pubblicazione: (2025)
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
di: Arghal, Raghu, et al.
Pubblicazione: (2026)
di: Arghal, Raghu, et al.
Pubblicazione: (2026)
Forecasting Frontier Language Model Agent Capabilities
di: Pimpale, Govind, et al.
Pubblicazione: (2025)
di: Pimpale, Govind, et al.
Pubblicazione: (2025)
The Elicitation Game: Evaluating Capability Elicitation Techniques
di: Hofstätter, Felix, et al.
Pubblicazione: (2025)
di: Hofstätter, Felix, et al.
Pubblicazione: (2025)
AutoML for Multi-Class Anomaly Compensation of Sensor Drift
di: Schaller, Melanie, et al.
Pubblicazione: (2025)
di: Schaller, Melanie, et al.
Pubblicazione: (2025)
Cisco Time Series Model Technical Report
di: Gou, Liang, et al.
Pubblicazione: (2025)
di: Gou, Liang, et al.
Pubblicazione: (2025)
Removing Sandbagging in LLMs by Training with Weak Supervision
di: Ryd, Emil, et al.
Pubblicazione: (2026)
di: Ryd, Emil, et al.
Pubblicazione: (2026)
LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
di: Li, Guozhong, et al.
Pubblicazione: (2025)
di: Li, Guozhong, et al.
Pubblicazione: (2025)
Technical Report: Small Language Model for Japanese Clinical and Medicine
di: Watanabe, Shogo
Pubblicazione: (2024)
di: Watanabe, Shogo
Pubblicazione: (2024)
World Models with Hints of Large Language Models for Goal Achieving
di: Liu, Zeyuan, et al.
Pubblicazione: (2024)
di: Liu, Zeyuan, et al.
Pubblicazione: (2024)
EPT-2 Technical Report
di: Molinaro, Roberto, et al.
Pubblicazione: (2025)
di: Molinaro, Roberto, et al.
Pubblicazione: (2025)
INTELLECT-3: Technical Report
di: Prime Intellect Team, et al.
Pubblicazione: (2025)
di: Prime Intellect Team, et al.
Pubblicazione: (2025)
LFM2 Technical Report
di: Amini, Alexander, et al.
Pubblicazione: (2025)
di: Amini, Alexander, et al.
Pubblicazione: (2025)
How does information access affect LLM monitors' ability to detect sabotage?
di: Arike, Rauno, et al.
Pubblicazione: (2026)
di: Arike, Rauno, et al.
Pubblicazione: (2026)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
DriftXpress: Faster Drifting Models via Projected RKHS Fields
di: Falahati, Ali, et al.
Pubblicazione: (2026)
di: Falahati, Ali, et al.
Pubblicazione: (2026)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
di: Højmark, Axel, et al.
Pubblicazione: (2024)
di: Højmark, Axel, et al.
Pubblicazione: (2024)
Goal-Conditioned Agents that Learn Everything All at Once
di: Matthews, Michael, et al.
Pubblicazione: (2026)
di: Matthews, Michael, et al.
Pubblicazione: (2026)
Recourse under Model Multiplicity via Argumentative Ensembling (Technical Report)
di: Jiang, Junqi, et al.
Pubblicazione: (2023)
di: Jiang, Junqi, et al.
Pubblicazione: (2023)
Training Deliberative Monitors for Black-Box Scheming Detection
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
Overcoming Dependent Censoring in the Evaluation of Survival Models
di: Lillelund, Christian Marius, et al.
Pubblicazione: (2025)
di: Lillelund, Christian Marius, et al.
Pubblicazione: (2025)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
di: Kinniment, Megan, et al.
Pubblicazione: (2023)
di: Kinniment, Megan, et al.
Pubblicazione: (2023)
Optimal Self-Consistency for Efficient Reasoning with Large Language Models
di: Feng, Austin, et al.
Pubblicazione: (2025)
di: Feng, Austin, et al.
Pubblicazione: (2025)
Motif 2.6B Technical Report
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
RLDX-1 Technical Report
di: Kim, Dongyoung, et al.
Pubblicazione: (2026)
di: Kim, Dongyoung, et al.
Pubblicazione: (2026)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
di: Burden, John, et al.
Pubblicazione: (2025)
di: Burden, John, et al.
Pubblicazione: (2025)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
di: He, Yufei, et al.
Pubblicazione: (2025)
di: He, Yufei, et al.
Pubblicazione: (2025)
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
di: Nimonkar, Chirayu, et al.
Pubblicazione: (2025)
di: Nimonkar, Chirayu, et al.
Pubblicazione: (2025)
Reaching Consensus in Cooperative Multi-Agent Reinforcement Learning with Goal Imagination
di: Wang, Liangzhou, et al.
Pubblicazione: (2024)
di: Wang, Liangzhou, et al.
Pubblicazione: (2024)
Goal Recognition Design for General Behavioral Agents using Machine Learning
di: Kasumba, Robert, et al.
Pubblicazione: (2024)
di: Kasumba, Robert, et al.
Pubblicazione: (2024)
Argumentative Debates for Transparent Bias Detection [Technical Report]
di: Ayoobi, Hamed, et al.
Pubblicazione: (2025)
di: Ayoobi, Hamed, et al.
Pubblicazione: (2025)
On the Wasserstein Gradient Flow Interpretation of Drifting Models
di: Gretton, Arthur, et al.
Pubblicazione: (2026)
di: Gretton, Arthur, et al.
Pubblicazione: (2026)
Incoherence in Goal-Conditioned Autoregressive Models
di: Karwowski, Jacek, et al.
Pubblicazione: (2025)
di: Karwowski, Jacek, et al.
Pubblicazione: (2025)
Documenti analoghi
-
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023) -
Large Language Models Often Know When They Are Being Evaluated
di: Needham, Joe, et al.
Pubblicazione: (2025) -
Excess Description Length of Learning Generalizable Predictors
di: Donoway, Elizabeth, et al.
Pubblicazione: (2026) -
Frontier Models are Capable of In-context Scheming
di: Meinke, Alexander, et al.
Pubblicazione: (2024)