To Tell The Truth: Language of Deception and Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hazra, Sanchaita, Majumder, Bodhisattwa Prasad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AI Safety Should Prioritize the Future of Work
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
von: Bhattacharya, Haimanti, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Haimanti, et al.
Veröffentlicht: (2024)
Data-driven Discovery with Large Generative Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
von: Kalluri, Tarun, et al.
Veröffentlicht: (2024)
von: Kalluri, Tarun, et al.
Veröffentlicht: (2024)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
von: Chen, Jiangjie, et al.
Veröffentlicht: (2023)
von: Chen, Jiangjie, et al.
Veröffentlicht: (2023)
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
von: Shah, Arya, et al.
Veröffentlicht: (2026)
von: Shah, Arya, et al.
Veröffentlicht: (2026)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
von: Jansen, Peter, et al.
Veröffentlicht: (2024)
von: Jansen, Peter, et al.
Veröffentlicht: (2024)
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
On the Relationship between Truth and Political Bias in Language Models
von: Fulay, Suyash, et al.
Veröffentlicht: (2024)
von: Fulay, Suyash, et al.
Veröffentlicht: (2024)
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
von: Jansen, Peter, et al.
Veröffentlicht: (2025)
von: Jansen, Peter, et al.
Veröffentlicht: (2025)
Truth Knows No Language: Evaluating Truthfulness Beyond English
von: Figueras, Blanca Calvo, et al.
Veröffentlicht: (2025)
von: Figueras, Blanca Calvo, et al.
Veröffentlicht: (2025)
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023)
von: Hagendorff, Thilo
Veröffentlicht: (2023)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
von: Lermen, Simon, et al.
Veröffentlicht: (2025)
von: Lermen, Simon, et al.
Veröffentlicht: (2025)
Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models
von: Guo, Linge
Veröffentlicht: (2024)
von: Guo, Linge
Veröffentlicht: (2024)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
von: Merrill, Scott, et al.
Veröffentlicht: (2026)
von: Merrill, Scott, et al.
Veröffentlicht: (2026)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
From Deception to Detection: The Dual Roles of Large Language Models in Fake News
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Models
von: Li, Mingda, et al.
Veröffentlicht: (2026)
von: Li, Mingda, et al.
Veröffentlicht: (2026)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2026)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2026)
Personas as a Way to Model Truthfulness in Language Models
von: Joshi, Nitish, et al.
Veröffentlicht: (2023)
von: Joshi, Nitish, et al.
Veröffentlicht: (2023)
A Study on Effect of Reference Knowledge Choice in Generating Technical Content Relevant to SAPPhIRE Model Using Large Language Model
von: Bhattacharya, Kausik, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Kausik, et al.
Veröffentlicht: (2024)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
Tell me the truth: A system to measure the trustworthiness of Large Language Models
von: Lipizzi, Carlo
Veröffentlicht: (2024)
von: Lipizzi, Carlo
Veröffentlicht: (2024)
Ranking Large Language Models without Ground Truth
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
An Assessment of Model-On-Model Deception
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models
von: Ismailov, Altynbek, et al.
Veröffentlicht: (2025)
von: Ismailov, Altynbek, et al.
Veröffentlicht: (2025)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
von: de Carvalho, Gonçalo Hora, et al.
Veröffentlicht: (2024)
von: de Carvalho, Gonçalo Hora, et al.
Veröffentlicht: (2024)
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models
von: Berg, Cameron, et al.
Veröffentlicht: (2026)
von: Berg, Cameron, et al.
Veröffentlicht: (2026)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
von: Järviniemi, Olli, et al.
Veröffentlicht: (2024)
von: Järviniemi, Olli, et al.
Veröffentlicht: (2024)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AI Safety Should Prioritize the Future of Work
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025) -
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025) -
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
von: Bhattacharya, Haimanti, et al.
Veröffentlicht: (2024) -
Data-driven Discovery with Large Generative Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024) -
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
von: Kalluri, Tarun, et al.
Veröffentlicht: (2024)