Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Kaiqu, Hu, Haimin, Zhao, Xuandong, Song, Dawn, Griffiths, Thomas L., Fisac, Jaime Fernández |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
von: Liang, Kaiqu, et al.
Veröffentlicht: (2024)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
Who Plays First? Optimizing the Order of Play in Stackelberg Games with Many Robots
von: Hu, Haimin, et al.
Veröffentlicht: (2024)
von: Hu, Haimin, et al.
Veröffentlicht: (2024)
In-Context Watermarks for Large Language Models
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
A Practical Examination of AI-Generated Text Detectors for Large Language Models
von: Tufts, Brian, et al.
Veröffentlicht: (2024)
von: Tufts, Brian, et al.
Veröffentlicht: (2024)
Emergent Semantic Role Understanding in Language Models
von: Griffiths, Carla, et al.
Veröffentlicht: (2026)
von: Griffiths, Carla, et al.
Veröffentlicht: (2026)
Incoherent Probability Judgments in Large Language Models
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
GRATH: Gradual Self-Truthifying for Large Language Models
von: Chen, Weixin, et al.
Veröffentlicht: (2024)
von: Chen, Weixin, et al.
Veröffentlicht: (2024)
Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models
von: Hu, Yueqing, et al.
Veröffentlicht: (2026)
von: Hu, Yueqing, et al.
Veröffentlicht: (2026)
Learning Personalized Agents from Human Feedback
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2025)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2025)
Multimodal Situational Safety
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2024)
MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety
von: Wang, Justin, et al.
Veröffentlicht: (2024)
von: Wang, Justin, et al.
Veröffentlicht: (2024)
Recovering Event Probabilities from Large Language Model Embeddings via Axiomatic Constraints
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2025)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2025)
Recovering Mental Representations from Large Language Models with Markov Chain Monte Carlo
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
What is a Number, That a Large Language Model May Know It?
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
An Undetectable Watermark for Generative Image Models
von: Gunn, Sam, et al.
Veröffentlicht: (2024)
von: Gunn, Sam, et al.
Veröffentlicht: (2024)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
Emergent Introspective Awareness in Large Language Models
von: Lindsey, Jack
Veröffentlicht: (2026)
von: Lindsey, Jack
Veröffentlicht: (2026)
Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate
von: Pecher, Tom
Veröffentlicht: (2026)
von: Pecher, Tom
Veröffentlicht: (2026)
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
von: Geng, Jiayi, et al.
Veröffentlicht: (2025)
von: Geng, Jiayi, et al.
Veröffentlicht: (2025)
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2025)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2025)
Ranking Large Language Models without Ground Truth
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
von: Song, Yeongwoo, et al.
Veröffentlicht: (2025)
von: Song, Yeongwoo, et al.
Veröffentlicht: (2025)
Rational Metareasoning for Large Language Models
von: De Sabbata, C. Nicolò, et al.
Veröffentlicht: (2024)
von: De Sabbata, C. Nicolò, et al.
Veröffentlicht: (2024)
Are Large Language Models Sensitive to the Motives Behind Communication?
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
von: Wu, Addison J., et al.
Veröffentlicht: (2025)
InfoSynth: Information-Guided Benchmark Synthesis for LLMs
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
von: Xie, Jingxu, et al.
Veröffentlicht: (2025)
von: Xie, Jingxu, et al.
Veröffentlicht: (2025)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
von: Zhang, Liyi, et al.
Veröffentlicht: (2025)
von: Zhang, Liyi, et al.
Veröffentlicht: (2025)
To Tell The Truth: Language of Deception and Language Models
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
von: Veselovsky, Veniamin, et al.
Veröffentlicht: (2025)
von: Veselovsky, Veniamin, et al.
Veröffentlicht: (2025)
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
von: Crispino, Nicholas, et al.
Veröffentlicht: (2023)
von: Crispino, Nicholas, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025) -
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
von: Liang, Kaiqu, et al.
Veröffentlicht: (2024) -
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025) -
Who Plays First? Optimizing the Order of Play in Stackelberg Games with Many Robots
von: Hu, Haimin, et al.
Veröffentlicht: (2024) -
In-Context Watermarks for Large Language Models
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)