Saved in:
| Main Authors: | Horvitz, Eric, Conitzer, Vincent, McIlraith, Sheila, Stone, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.04750 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
by: Alamdari, Parand A., et al.
Published: (2026)
by: Alamdari, Parand A., et al.
Published: (2026)
Pluralistic Alignment Over Time
by: Klassen, Toryn Q., et al.
Published: (2024)
by: Klassen, Toryn Q., et al.
Published: (2024)
Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI
by: Alamdari, Parand A., et al.
Published: (2024)
by: Alamdari, Parand A., et al.
Published: (2024)
Remembering to Be Fair: Non-Markovian Fairness in Sequential Decision Making
by: Alamdari, Parand A., et al.
Published: (2023)
by: Alamdari, Parand A., et al.
Published: (2023)
Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
by: Lifshitz, Shalev, et al.
Published: (2025)
by: Lifshitz, Shalev, et al.
Published: (2025)
Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
by: Chen, Dillon Z., et al.
Published: (2025)
by: Chen, Dillon Z., et al.
Published: (2025)
Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning
by: Chen, Dillon Z., et al.
Published: (2026)
by: Chen, Dillon Z., et al.
Published: (2026)
Accurate Measures of Vaccination and Concerns of Vaccine Holdouts from Web Search Logs
by: Chang, Serina, et al.
Published: (2023)
by: Chang, Serina, et al.
Published: (2023)
Narrative Frames: A New Approach to Analysing Metaphors in AI Ethics and Policy Discourse
by: Stone, Daniel
Published: (2026)
by: Stone, Daniel
Published: (2026)
Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest
by: Shi, Jerick, et al.
Published: (2026)
by: Shi, Jerick, et al.
Published: (2026)
From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception
by: Shi, Jerick, et al.
Published: (2026)
by: Shi, Jerick, et al.
Published: (2026)
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft
by: Lifshitz, Shalev, et al.
Published: (2023)
by: Lifshitz, Shalev, et al.
Published: (2023)
Satisficing and Optimal Generalised Planning via Goal Regression (Extended Version)
by: Chen, Dillon Z., et al.
Published: (2025)
by: Chen, Dillon Z., et al.
Published: (2025)
Expert Survey: AI Reliability & Security Research Priorities
by: O'Brien, Joe, et al.
Published: (2025)
by: O'Brien, Joe, et al.
Published: (2025)
The Singapore Consensus on Global AI Safety Research Priorities
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
by: Keswani, Vijay, et al.
Published: (2025)
by: Keswani, Vijay, et al.
Published: (2025)
Now More Than Ever, Foundational AI Research and Infrastructure Depends on the Federal Government
by: Taufer, Michela, et al.
Published: (2025)
by: Taufer, Michela, et al.
Published: (2025)
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
by: Keswani, Vijay, et al.
Published: (2025)
by: Keswani, Vijay, et al.
Published: (2025)
Managing extreme AI risks amid rapid progress
by: Bengio, Yoshua, et al.
Published: (2023)
by: Bengio, Yoshua, et al.
Published: (2023)
It's Not the AI - It's Each of Us! Ten Commandments for the Wise & Responsible Use of AI
by: Steffen, Barbara, et al.
Published: (2025)
by: Steffen, Barbara, et al.
Published: (2025)
On the Pros and Cons of Active Learning for Moral Preference Elicitation
by: Keswani, Vijay, et al.
Published: (2024)
by: Keswani, Vijay, et al.
Published: (2024)
Better Training Data Attribution via Better Inverse Hessian-Vector Products
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
An LLM's Apology: Outsourcing Awkwardness in the Age of AI
by: Stone, Twm, et al.
Published: (2025)
by: Stone, Twm, et al.
Published: (2025)
Defense Priorities in the Open-Source AI Debate: A Preliminary Assessment
by: Dahlgren, Masao
Published: (2024)
by: Dahlgren, Masao
Published: (2024)
The Essentials of AI for Life and Society: A Full-Scale AI Literacy Course Accessible to All
by: Xu, Zifan, et al.
Published: (2025)
by: Xu, Zifan, et al.
Published: (2025)
Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
by: Li, Andrew C., et al.
Published: (2025)
by: Li, Andrew C., et al.
Published: (2025)
Pushdown Reward Machines for Reinforcement Learning
by: Varricchione, Giovanni, et al.
Published: (2025)
by: Varricchione, Giovanni, et al.
Published: (2025)
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
by: Boerstler, Kyle, et al.
Published: (2024)
by: Boerstler, Kyle, et al.
Published: (2024)
The Complexity of Computing Robust Mediated Equilibria in Ordinal Games
by: Conitzer, Vincent
Published: (2024)
by: Conitzer, Vincent
Published: (2024)
AI-Generated Figures in Academic Publishing: Policies, Tools, and Practical Guidelines
by: Chen, Davie
Published: (2026)
by: Chen, Davie
Published: (2026)
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
by: Dias, Gabriela Aránguiz, et al.
Published: (2026)
by: Dias, Gabriela Aránguiz, et al.
Published: (2026)
Academics and Generative AI: Empirical and Epistemic Indicators of Policy-Practice Voids
by: Ravenor, R. Yamamoto
Published: (2025)
by: Ravenor, R. Yamamoto
Published: (2025)
Assessing Privacy Policies with AI: Ethical, Legal, and Technical Challenges
by: Aydin, Irem, et al.
Published: (2024)
by: Aydin, Irem, et al.
Published: (2024)
The Essentials of AI for Life and Society: An AI Literacy Course for the University Community
by: Biswas, Joydeep, et al.
Published: (2025)
by: Biswas, Joydeep, et al.
Published: (2025)
Analysis of Generative AI Policies in Computing Course Syllabi
by: Ali, Areej, et al.
Published: (2024)
by: Ali, Areej, et al.
Published: (2024)
Bridging the Gap: Integrating Ethics and Environmental Sustainability in AI Research and Practice
by: Luccioni, Alexandra Sasha, et al.
Published: (2025)
by: Luccioni, Alexandra Sasha, et al.
Published: (2025)
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025)
by: Kim, HyunJin, et al.
Published: (2025)
Gauss-Newton Unlearning for the LLM Era
by: McKinney, Lev, et al.
Published: (2026)
by: McKinney, Lev, et al.
Published: (2026)
Centering Policy and Practice: Research Gaps around Usable Differential Privacy
by: Cummings, Rachel, et al.
Published: (2024)
by: Cummings, Rachel, et al.
Published: (2024)
Now You See Me: Designing Responsible AI Dashboards for Early-Stage Health Innovation
by: Surodina, Svitlana, et al.
Published: (2026)
by: Surodina, Svitlana, et al.
Published: (2026)
Similar Items
-
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
by: Alamdari, Parand A., et al.
Published: (2026) -
Pluralistic Alignment Over Time
by: Klassen, Toryn Q., et al.
Published: (2024) -
Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI
by: Alamdari, Parand A., et al.
Published: (2024) -
Remembering to Be Fair: Non-Markovian Fairness in Sequential Decision Making
by: Alamdari, Parand A., et al.
Published: (2023) -
Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
by: Lifshitz, Shalev, et al.
Published: (2025)