Know Thyself? On the Incapability and Implications of AI Self-Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Xiaoyan, Shrivastava, Aryan, Holtzman, Ari, Tan, Chenhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
von: Liu, Haokun, et al.
Veröffentlicht: (2025)
von: Liu, Haokun, et al.
Veröffentlicht: (2025)
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
von: Liu, Haokun, et al.
Veröffentlicht: (2024)
von: Liu, Haokun, et al.
Veröffentlicht: (2024)
Hypothesis Generation with Large Language Models
von: Zhou, Yangqiaoyu, et al.
Veröffentlicht: (2024)
von: Zhou, Yangqiaoyu, et al.
Veröffentlicht: (2024)
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
von: Heddaya, Mourad, et al.
Veröffentlicht: (2024)
von: Heddaya, Mourad, et al.
Veröffentlicht: (2024)
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2026)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024)
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024)
AI as Entertainment
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
Forking Paths in Neural Text Generation
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
von: Kıcıman, Emre, et al.
Veröffentlicht: (2023)
von: Kıcıman, Emre, et al.
Veröffentlicht: (2023)
Interpretable Recognition of Cognitive Distortions in Natural Language Texts
von: Kolonin, Anton, et al.
Veröffentlicht: (2025)
von: Kolonin, Anton, et al.
Veröffentlicht: (2025)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2024)
AI-AI Bias: large language models favor communications generated by large language models
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
von: Schroeder, Daniel Thilo, et al.
Veröffentlicht: (2025)
von: Schroeder, Daniel Thilo, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Risks of Generative AI in Financial Services
von: Gehrmann, Sebastian, et al.
Veröffentlicht: (2025)
von: Gehrmann, Sebastian, et al.
Veröffentlicht: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
Managing extreme AI risks amid rapid progress
von: Bengio, Yoshua, et al.
Veröffentlicht: (2023)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2023)
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2025)
von: Han, Pengrui, et al.
Veröffentlicht: (2025)
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
von: Imperial, Joseph Marvin, et al.
Veröffentlicht: (2025)
von: Imperial, Joseph Marvin, et al.
Veröffentlicht: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
Impacts of Racial Bias in Historical Training Data for News AI
von: Bhargava, Rahul, et al.
Veröffentlicht: (2025)
von: Bhargava, Rahul, et al.
Veröffentlicht: (2025)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2024)
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2024)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
von: Shojaei, Mostafa Faghih, et al.
Veröffentlicht: (2025)
von: Shojaei, Mostafa Faghih, et al.
Veröffentlicht: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024)
von: Ren, Richard, et al.
Veröffentlicht: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
von: Yang, Chao, et al.
Veröffentlicht: (2024)
von: Yang, Chao, et al.
Veröffentlicht: (2024)
Frontier AI systems have surpassed the self-replicating red line
von: Pan, Xudong, et al.
Veröffentlicht: (2024)
von: Pan, Xudong, et al.
Veröffentlicht: (2024)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
von: Shin, Kwan Soo
Veröffentlicht: (2026)
von: Shin, Kwan Soo
Veröffentlicht: (2026)
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
von: Lerner, Emilia Agis, et al.
Veröffentlicht: (2024)
von: Lerner, Emilia Agis, et al.
Veröffentlicht: (2024)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
von: Gringras, David
Veröffentlicht: (2026)
von: Gringras, David
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026) -
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025) -
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026) -
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026) -
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)