Safety Must Precede the Deployment of Open-Ended AI
Fuente:
arXiv
Saved in:
| Main Authors: | Sheth, Ivaxi, Wehner, Jan, Abdelnabi, Sahar, Binkyte, Ruta, Fritz, Mario |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution
by: Binkyte, Ruta, et al.
Published: (2026)
by: Binkyte, Ruta, et al.
Published: (2026)
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
by: Binkyte, Ruta, et al.
Published: (2025)
by: Binkyte, Ruta, et al.
Published: (2025)
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
by: Vasu, Sai Suresh Macharla, et al.
Published: (2025)
by: Vasu, Sai Suresh Macharla, et al.
Published: (2025)
LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation
by: Afonja, Tejumade, et al.
Published: (2024)
by: Afonja, Tejumade, et al.
Published: (2024)
Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables
by: Sheth, Ivaxi, et al.
Published: (2024)
by: Sheth, Ivaxi, et al.
Published: (2024)
Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework
by: Binkyte, Ruta
Published: (2025)
by: Binkyte, Ruta
Published: (2025)
ProtocolLLM: RTL Benchmark for SystemVerilog Generation of Communication Protocols
by: Sheth, Arnav, et al.
Published: (2025)
by: Sheth, Arnav, et al.
Published: (2025)
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
Inspectable AI for Science: A Research Object Approach to Generative AI Governance
by: Binkyte, Ruta, et al.
Published: (2026)
by: Binkyte, Ruta, et al.
Published: (2026)
On the Need and Applicability of Causality for Fairness: A Unified Framework for AI Auditing and Legal Analysis
by: Binkyte, Ruta, et al.
Published: (2022)
by: Binkyte, Ruta, et al.
Published: (2022)
IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery
by: Sheth, Ivaxi, et al.
Published: (2026)
by: Sheth, Ivaxi, et al.
Published: (2026)
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs
by: Labroo, Arya, et al.
Published: (2026)
by: Labroo, Arya, et al.
Published: (2026)
A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive
by: Sivaprasad, Sarath, et al.
Published: (2024)
by: Sivaprasad, Sarath, et al.
Published: (2024)
BaBE: Enhancing Fairness via Estimation of Latent Explaining Variables
by: Binkyte, Ruta, et al.
Published: (2023)
by: Binkyte, Ruta, et al.
Published: (2023)
Survey on AI Ethics: A Socio-technical Perspective
by: Mbiazi, Dave, et al.
Published: (2023)
by: Mbiazi, Dave, et al.
Published: (2023)
Position: AI Safety Must Embrace an Antifragile Perspective
by: Jin, Ming, et al.
Published: (2025)
by: Jin, Ming, et al.
Published: (2025)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
Models That Know How Evaluations Are Designed Score Safer
by: Deckenbach, Katharina, et al.
Published: (2026)
by: Deckenbach, Katharina, et al.
Published: (2026)
Mental Health AI Safety Claims Must Preserve Temporal Evidence
by: Dutta, Srimonti, et al.
Published: (2026)
by: Dutta, Srimonti, et al.
Published: (2026)
On Creativity and Open-Endedness
by: Soros, L. B., et al.
Published: (2024)
by: Soros, L. B., et al.
Published: (2024)
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
by: Wehner, Jan, et al.
Published: (2025)
by: Wehner, Jan, et al.
Published: (2025)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
AI Must not be Fully Autonomous
by: Adewumi, Tosin, et al.
Published: (2025)
by: Adewumi, Tosin, et al.
Published: (2025)
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
by: Liu, Zewen, et al.
Published: (2026)
by: Liu, Zewen, et al.
Published: (2026)
The Deployment Gap in AI Media Detection: Platform-Aware and Visually Constrained Adversarial Evaluation
by: Budhkar, Aishwarya, et al.
Published: (2026)
by: Budhkar, Aishwarya, et al.
Published: (2026)
Fundamental Risks in the Current Deployment of General-Purpose AI Models: What Have We (Not) Learnt From Cybersecurity?
by: Fritz, Mario
Published: (2024)
by: Fritz, Mario
Published: (2024)
Causal Discovery Under Local Privacy
by: Binkytė, Rūta, et al.
Published: (2023)
by: Binkytė, Rūta, et al.
Published: (2023)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
by: Ying, Lance, et al.
Published: (2026)
by: Ying, Lance, et al.
Published: (2026)
Pessimistic Verification for Open Ended Math Questions
by: Huang, Yanxing, et al.
Published: (2025)
by: Huang, Yanxing, et al.
Published: (2025)
CausalGraph2LLM: Evaluating LLMs for Causal Queries
by: Sheth, Ivaxi, et al.
Published: (2024)
by: Sheth, Ivaxi, et al.
Published: (2024)
Amorphous Fortress Online: Collaboratively Designing Open-Ended Multi-Agent AI and Game Environments
by: Charity, M, et al.
Published: (2025)
by: Charity, M, et al.
Published: (2025)
When Precedents Clash
by: Di Florio, Cecilia, et al.
Published: (2024)
by: Di Florio, Cecilia, et al.
Published: (2024)
On Improvisation and Open-Endedness: Insights for Experiential AI
by: Hu, Botao 'Amber'
Published: (2025)
by: Hu, Botao 'Amber'
Published: (2025)
LLM Agents Beyond Utility: An Open-Ended Perspective
by: Nachkov, Asen, et al.
Published: (2025)
by: Nachkov, Asen, et al.
Published: (2025)
Trustworthy AI Must Account for Interactions
by: Cresswell, Jesse C.
Published: (2025)
by: Cresswell, Jesse C.
Published: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
by: Lu, Chris, et al.
Published: (2024)
by: Lu, Chris, et al.
Published: (2024)
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
Similar Items
-
Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution
by: Binkyte, Ruta, et al.
Published: (2026) -
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
by: Binkyte, Ruta, et al.
Published: (2025) -
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
by: Vasu, Sai Suresh Macharla, et al.
Published: (2025) -
LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation
by: Afonja, Tejumade, et al.
Published: (2024) -
Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables
by: Sheth, Ivaxi, et al.
Published: (2024)