Saved in:
| Main Author: | Han, Shanshan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.18114 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
by: Brophy, Matthew
Published: (2025)
by: Brophy, Matthew
Published: (2025)
AI Safety Should Prioritize the Future of Work
by: Hazra, Sanchaita, et al.
Published: (2025)
by: Hazra, Sanchaita, et al.
Published: (2025)
Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning
by: Wang, Olivia Peiyu, et al.
Published: (2026)
by: Wang, Olivia Peiyu, et al.
Published: (2026)
Charting the Future of AI-supported Science Education: A Human-Centered Vision
by: Zhai, Xiaoming, et al.
Published: (2026)
by: Zhai, Xiaoming, et al.
Published: (2026)
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction
by: Ehsan, Upol, et al.
Published: (2026)
by: Ehsan, Upol, et al.
Published: (2026)
Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics
by: Baum, Kevin
Published: (2025)
by: Baum, Kevin
Published: (2025)
NeuroAI and Beyond: Bridging Between Advances in Neuroscience and ArtificialIntelligence
by: Zador, Anthony, et al.
Published: (2026)
by: Zador, Anthony, et al.
Published: (2026)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025)
by: Dobbe, Roel
Published: (2025)
Human-AI Safety: A Descendant of Generative AI and Control Systems Safety
by: Bajcsy, Andrea, et al.
Published: (2024)
by: Bajcsy, Andrea, et al.
Published: (2024)
AI Fairness Beyond Complete Demographics: Current Achievements and Future Directions
by: Wang, Zichong, et al.
Published: (2025)
by: Wang, Zichong, et al.
Published: (2025)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
by: Scholefield, Rebecca, et al.
Published: (2025)
by: Scholefield, Rebecca, et al.
Published: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Bridging the Gap in the Responsible AI Divides
by: Gyevnár, Bálint, et al.
Published: (2026)
by: Gyevnár, Bálint, et al.
Published: (2026)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
Toward an African Agenda for AI Safety
by: Segun, Samuel T., et al.
Published: (2025)
by: Segun, Samuel T., et al.
Published: (2025)
Concrete Problems in AI Safety, Revisited
by: Raji, Inioluwa Deborah, et al.
Published: (2023)
by: Raji, Inioluwa Deborah, et al.
Published: (2023)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)
by: Vaccaro, Michelle, et al.
Published: (2026)
AI and the Future of Digital Public Squares
by: Goldberg, Beth, et al.
Published: (2024)
by: Goldberg, Beth, et al.
Published: (2024)
AI Safety: Necessary, but insufficient and possibly problematic
by: P, Deepak
Published: (2024)
by: P, Deepak
Published: (2024)
Emerging Practices in Frontier AI Safety Frameworks
by: Buhl, Marie Davidsen, et al.
Published: (2025)
by: Buhl, Marie Davidsen, et al.
Published: (2025)
Beyond Explainability: The Case for AI Validation
by: Feldman, Dalit Ken-Dror, et al.
Published: (2025)
by: Feldman, Dalit Ken-Dror, et al.
Published: (2025)
Governing AI Beyond the Pretraining Frontier
by: Caputo, Nicholas A.
Published: (2025)
by: Caputo, Nicholas A.
Published: (2025)
AI Consciousness and Public Perceptions: Four Futures
by: Fernandez, Ines, et al.
Published: (2024)
by: Fernandez, Ines, et al.
Published: (2024)
Trust in AI: Progress, Challenges, and Future Directions
by: Afroogh, Saleh, et al.
Published: (2024)
by: Afroogh, Saleh, et al.
Published: (2024)
AI-Enhanced Deliberative Democracy and the Future of the Collective Will
by: Revel, Manon, et al.
Published: (2025)
by: Revel, Manon, et al.
Published: (2025)
Upstream and Downstream AI Safety: Both on the Same River?
by: McDermid, John, et al.
Published: (2024)
by: McDermid, John, et al.
Published: (2024)
Probabilistic Analysis of Copyright Disputes and Generative AI Safety
by: Chiba-Okabe, Hiroaki
Published: (2024)
by: Chiba-Okabe, Hiroaki
Published: (2024)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
by: Costa, Mariana Lins
Published: (2026)
by: Costa, Mariana Lins
Published: (2026)
Combining Cost-Constrained Runtime Monitors for AI Safety
by: Hua, Tim Tian, et al.
Published: (2025)
by: Hua, Tim Tian, et al.
Published: (2025)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
by: Li, Miles Q., et al.
Published: (2026)
by: Li, Miles Q., et al.
Published: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Building Effective Safety Guardrails in AI Education Tools
by: Clark, Hannah-Beth, et al.
Published: (2025)
by: Clark, Hannah-Beth, et al.
Published: (2025)
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
by: Chin, Yik Chan, et al.
Published: (2026)
by: Chin, Yik Chan, et al.
Published: (2026)
What Is AI Safety? What Do We Want It to Be?
by: Harding, Jacqueline, et al.
Published: (2025)
by: Harding, Jacqueline, et al.
Published: (2025)
The Singapore Consensus on Global AI Safety Research Priorities
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
AI-Driven Human-Autonomy Teaming in Tactical Operations: Proposed Framework, Challenges, and Future Directions
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
Exploring AI Writers: Technology, Impact, and Future Prospects
by: Huang, Zhiqian
Published: (2025)
by: Huang, Zhiqian
Published: (2025)
AI-Educational Development Loop (AI-EDL): A Conceptual Framework to Bridge AI Capabilities with Classical Educational Theories
by: Yu, Ning, et al.
Published: (2025)
by: Yu, Ning, et al.
Published: (2025)
Similar Items
-
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
by: Brophy, Matthew
Published: (2025) -
AI Safety Should Prioritize the Future of Work
by: Hazra, Sanchaita, et al.
Published: (2025) -
Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning
by: Wang, Olivia Peiyu, et al.
Published: (2026) -
Charting the Future of AI-supported Science Education: A Human-Centered Vision
by: Zhai, Xiaoming, et al.
Published: (2026) -
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction
by: Ehsan, Upol, et al.
Published: (2026)