Enregistré dans:
| Auteur principal: | Bird, Steven |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2512.24863 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AI Safety Training Can be Clinically Harmful
par: BN, Suhas, et autres
Publié: (2026)
par: BN, Suhas, et autres
Publié: (2026)
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
par: Wong, Qi Han
Publié: (2026)
par: Wong, Qi Han
Publié: (2026)
Revealing Hidden Bias in AI: Lessons from Large Language Models
par: Beatty, Django, et autres
Publié: (2024)
par: Beatty, Django, et autres
Publié: (2024)
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
par: von Cossel, Oskar
Publié: (2026)
par: von Cossel, Oskar
Publié: (2026)
Growing a Tail: Increasing Output Diversity in Large Language Models
par: Shur-Ofry, Michal, et autres
Publié: (2024)
par: Shur-Ofry, Michal, et autres
Publié: (2024)
Pro-AI Bias in Large Language Models
par: Trabelsi, Benaya, et autres
Publié: (2026)
par: Trabelsi, Benaya, et autres
Publié: (2026)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
par: Zmanovskii, Nikita
Publié: (2025)
par: Zmanovskii, Nikita
Publié: (2025)
The AI Fiction Paradox
par: Elkins, Katherine
Publié: (2026)
par: Elkins, Katherine
Publié: (2026)
AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains
par: Shintani, Seine A.
Publié: (2026)
par: Shintani, Seine A.
Publié: (2026)
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
par: Wu, Shuai, et autres
Publié: (2026)
par: Wu, Shuai, et autres
Publié: (2026)
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
par: Arora, Sunil, et autres
Publié: (2025)
par: Arora, Sunil, et autres
Publié: (2025)
Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild
par: Loth, Alexander, et autres
Publié: (2026)
par: Loth, Alexander, et autres
Publié: (2026)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
par: Han, Pengrui, et autres
Publié: (2024)
par: Han, Pengrui, et autres
Publié: (2024)
Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems
par: Loth, Alexander, et autres
Publié: (2026)
par: Loth, Alexander, et autres
Publié: (2026)
Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels
par: Sharma, Anika, et autres
Publié: (2025)
par: Sharma, Anika, et autres
Publié: (2025)
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
par: Zhu, Pengyun, et autres
Publié: (2026)
par: Zhu, Pengyun, et autres
Publié: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
par: Beltoft, Stine, et autres
Publié: (2025)
par: Beltoft, Stine, et autres
Publié: (2025)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
par: Kelly, Matthew
Publié: (2025)
par: Kelly, Matthew
Publié: (2025)
Superhuman Game AI Disclosure: Expertise and Context Moderate Effects on Trust and Fairness
par: Chua, Jaymari, et autres
Publié: (2025)
par: Chua, Jaymari, et autres
Publié: (2025)
Large language models can replicate cross-cultural differences in personality
par: Niszczota, Paweł, et autres
Publié: (2023)
par: Niszczota, Paweł, et autres
Publié: (2023)
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts
par: Zahraei, Pardis Sadat, et autres
Publié: (2024)
par: Zahraei, Pardis Sadat, et autres
Publié: (2024)
Leveraging Multi-Source Textural UGC for Neighbourhood Housing Quality Assessment: A GPT-Enhanced Framework
par: Hong, Qiyuan, et autres
Publié: (2025)
par: Hong, Qiyuan, et autres
Publié: (2025)
Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistants
par: Vekaria, Yash, et autres
Publié: (2025)
par: Vekaria, Yash, et autres
Publié: (2025)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
par: Basu, Abhinaba, et autres
Publié: (2026)
par: Basu, Abhinaba, et autres
Publié: (2026)
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
par: Lehr, Steven A., et autres
Publié: (2024)
par: Lehr, Steven A., et autres
Publié: (2024)
The Fragility Of Moral Judgment In Large Language Models
par: van Nuenen, Tom, et autres
Publié: (2026)
par: van Nuenen, Tom, et autres
Publié: (2026)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
par: Yeste, Víctor, et autres
Publié: (2026)
par: Yeste, Víctor, et autres
Publié: (2026)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
par: Wang, Xinyue, et autres
Publié: (2026)
par: Wang, Xinyue, et autres
Publié: (2026)
The Company You Keep: How LLMs Respond to Dark Triad Traits
par: Lu, Zeyi, et autres
Publié: (2026)
par: Lu, Zeyi, et autres
Publié: (2026)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
par: Chen, Renmiao, et autres
Publié: (2025)
par: Chen, Renmiao, et autres
Publié: (2025)
The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete
par: Barmettler, Joel
Publié: (2026)
par: Barmettler, Joel
Publié: (2026)
Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle
par: Thorne, William, et autres
Publié: (2026)
par: Thorne, William, et autres
Publié: (2026)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
par: Yun, Bhada, et autres
Publié: (2026)
par: Yun, Bhada, et autres
Publié: (2026)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
par: Naik, Akshat, et autres
Publié: (2025)
par: Naik, Akshat, et autres
Publié: (2025)
Computable Gap Assessment of Artificial Intelligence Governance in Children's Centres: Evidence-Mechanism-Governance-Indicator Modelling of UNICEF's Guidance on AI and Children 3.0 Based on the Graph-GAP Framework
par: Meng, Wei
Publié: (2025)
par: Meng, Wei
Publié: (2025)
Powerful Training-Free Membership Inference Against Autoregressive Language Models
par: Ilić, David, et autres
Publié: (2026)
par: Ilić, David, et autres
Publié: (2026)
PromptAug: Fine-grained Conflict Classification Using Data Augmentation
par: Warke, Oliver, et autres
Publié: (2025)
par: Warke, Oliver, et autres
Publié: (2025)
Generative midtended cognition and Artificial Intelligence. Thinging with thinging things
par: Barandiaran, Xabier E., et autres
Publié: (2024)
par: Barandiaran, Xabier E., et autres
Publié: (2024)
Building the Web for Agents: A Declarative Framework for Agent-Web Interaction
par: Schultze, Sven, et autres
Publié: (2025)
par: Schultze, Sven, et autres
Publié: (2025)
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
par: Madan, Kapil
Publié: (2025)
par: Madan, Kapil
Publié: (2025)
Documents similaires
-
AI Safety Training Can be Clinically Harmful
par: BN, Suhas, et autres
Publié: (2026) -
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
par: Wong, Qi Han
Publié: (2026) -
Revealing Hidden Bias in AI: Lessons from Large Language Models
par: Beatty, Django, et autres
Publié: (2024) -
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
par: von Cossel, Oskar
Publié: (2026) -
Growing a Tail: Increasing Output Diversity in Large Language Models
par: Shur-Ofry, Michal, et autres
Publié: (2024)