Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | Bercovich, Ivan, Segal, Ivgeni, Zhang, Kexun, Saxena, Shashwat, Raghunathan, Aditi, Zhong, Ziqian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
by: Bercovich, Ivan
Published: (2026)
by: Bercovich, Ivan
Published: (2026)
Convolutional Model Trees
by: Armstrong, William Ward, et al.
Published: (2025)
by: Armstrong, William Ward, et al.
Published: (2025)
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
by: Huang, Xuanxiang, et al.
Published: (2025)
by: Huang, Xuanxiang, et al.
Published: (2025)
$γ(3,4)$ `Attention' in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics
by: Burgess, Mark
Published: (2025)
by: Burgess, Mark
Published: (2025)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
by: Li, Shenghao
Published: (2025)
by: Li, Shenghao
Published: (2025)
The Company You Keep: How LLMs Respond to Dark Triad Traits
by: Lu, Zeyi, et al.
Published: (2026)
by: Lu, Zeyi, et al.
Published: (2026)
Tactile-Proprioceptive Sensor Fusion for Contact Wrench Estimation in Whole-Body Physical Human-Robot Interaction
by: Min, Junha, et al.
Published: (2026)
by: Min, Junha, et al.
Published: (2026)
A Sequent Calculus for General Inductive Definitions
by: Eede, Robbe Van den, et al.
Published: (2026)
by: Eede, Robbe Van den, et al.
Published: (2026)
Membership Inference Attacks against Large Audio Language Models
by: Dong, Jia-Kai, et al.
Published: (2026)
by: Dong, Jia-Kai, et al.
Published: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
by: Lucas, Tom, et al.
Published: (2026)
by: Lucas, Tom, et al.
Published: (2026)
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
by: von Cossel, Oskar
Published: (2026)
by: von Cossel, Oskar
Published: (2026)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
by: Xu, Qiyuan, et al.
Published: (2026)
by: Xu, Qiyuan, et al.
Published: (2026)
MultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging Modalities
by: Lincetto, Federico, et al.
Published: (2025)
by: Lincetto, Federico, et al.
Published: (2025)
Fusing Structure from Motion and Simulation-Augmented Pose Regression from Optical Flow for Challenging Indoor Environments
by: Ott, Felix, et al.
Published: (2023)
by: Ott, Felix, et al.
Published: (2023)
On The Role of Intentionality in Knowledge Representation: Analyzing Scene Context for Cognitive Agents with a Tiny Language Model
by: Burgess, Mark
Published: (2025)
by: Burgess, Mark
Published: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024)
by: Kashyap, Pankhi, et al.
Published: (2024)
The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete
by: Barmettler, Joel
Published: (2026)
by: Barmettler, Joel
Published: (2026)
Integration of Contextual Descriptors in Ontology Alignment for Enrichment of Semantic Correspondence
by: Manziuk, Eduard, et al.
Published: (2024)
by: Manziuk, Eduard, et al.
Published: (2024)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
by: Beltoft, Stine, et al.
Published: (2025)
by: Beltoft, Stine, et al.
Published: (2025)
Consensus and Synchronization of Multi-agent Systems over Finite Fields -- Graph Topologies
by: Hengster-Movrić, Kristian, et al.
Published: (2026)
by: Hengster-Movrić, Kristian, et al.
Published: (2026)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
KernelOracle: Predicting the Linux Scheduler's Next Move with Deep Learning
by: Kahu, Sampanna Yashwant
Published: (2025)
by: Kahu, Sampanna Yashwant
Published: (2025)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
by: Cohen, Liran, et al.
Published: (2025)
by: Cohen, Liran, et al.
Published: (2025)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
Privacy as Commodity: MFG-RegretNet for Large-Scale Privacy Trading in Federated Learning
by: Sun, Kangkang, et al.
Published: (2026)
by: Sun, Kangkang, et al.
Published: (2026)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
by: Poddar, Aheli, et al.
Published: (2025)
by: Poddar, Aheli, et al.
Published: (2025)
HybridVFL: Disentangled Feature Learning for Edge-Enabled Vertical Federated Multimodal Classification
by: Anoosha, Mostafa, et al.
Published: (2025)
by: Anoosha, Mostafa, et al.
Published: (2025)
Automated Theorem Provers Help Improve Large Language Model Reasoning
by: McGinness, Lachlan, et al.
Published: (2024)
by: McGinness, Lachlan, et al.
Published: (2024)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
by: Yagoubi, Faouzi El, et al.
Published: (2026)
by: Yagoubi, Faouzi El, et al.
Published: (2026)
Why we need an AI-resilient society
by: Bartz-Beielstein, Thomas
Published: (2019)
by: Bartz-Beielstein, Thomas
Published: (2019)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026)
by: Tang, Wenjie, et al.
Published: (2026)
Ontology Neural Network and ORTSF: A Framework for Topological Reasoning and Delay-Robust Control
by: Oh, Jaehong
Published: (2025)
by: Oh, Jaehong
Published: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
by: Saxena, Udit
Published: (2025)
by: Saxena, Udit
Published: (2025)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
by: Abzianidze, Lasha
Published: (2023)
by: Abzianidze, Lasha
Published: (2023)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
by: Yun, Bhada, et al.
Published: (2026)
by: Yun, Bhada, et al.
Published: (2026)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
by: Naik, Akshat, et al.
Published: (2025)
by: Naik, Akshat, et al.
Published: (2025)
Selecting for Less Discriminatory Algorithms: A Relational Search Framework for Navigating Fairness-Accuracy Trade-offs in Practice
by: Samad, Hana, et al.
Published: (2025)
by: Samad, Hana, et al.
Published: (2025)
SPARQL in N3: SPARQL CONSTRUCT as a rule language for the Semantic Web (Extended Version)
by: Arndt, Dörthe, et al.
Published: (2025)
by: Arndt, Dörthe, et al.
Published: (2025)
Similar Items
-
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
by: Bercovich, Ivan
Published: (2026) -
Convolutional Model Trees
by: Armstrong, William Ward, et al.
Published: (2025) -
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
by: Huang, Xuanxiang, et al.
Published: (2025) -
$γ(3,4)$ `Attention' in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics
by: Burgess, Mark
Published: (2025) -
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
by: Li, Shenghao
Published: (2025)