Effective faking of verbal deception detection with target-aligned adversarial attacks
Fuente:
arXiv
Salvato in:
| Autori principali: | Kleinberg, Bennett, Loconte, Riccardo, Verschuere, Bruno |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Effective faking of verbal deception detection with target‐aligned adversarial attacks
di: Bennett Kleinberg, et al.
Pubblicazione: (2025)
di: Bennett Kleinberg, et al.
Pubblicazione: (2025)
When lies are mostly truthful: automated verbal deception detection for embedded lies
di: Loconte, Riccardo, et al.
Pubblicazione: (2025)
di: Loconte, Riccardo, et al.
Pubblicazione: (2025)
Linguistic traces of stochastic empathy in language models
di: Kleinberg, Bennett, et al.
Pubblicazione: (2024)
di: Kleinberg, Bennett, et al.
Pubblicazione: (2024)
Humans incorrectly reject confident accusatory AI judgments
di: Loconte, Riccardo, et al.
Pubblicazione: (2025)
di: Loconte, Riccardo, et al.
Pubblicazione: (2025)
Large Language Models in Cryptocurrency Securities Cases: Can a GPT Model Meaningfully Assist Lawyers?
di: Trozze, Arianna, et al.
Pubblicazione: (2023)
di: Trozze, Arianna, et al.
Pubblicazione: (2023)
Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?
di: Rennard, Virgile, et al.
Pubblicazione: (2024)
di: Rennard, Virgile, et al.
Pubblicazione: (2024)
Are aligned neural networks adversarially aligned?
di: Carlini, Nicholas, et al.
Pubblicazione: (2023)
di: Carlini, Nicholas, et al.
Pubblicazione: (2023)
Can adversarial attacks by large language models be attributed?
di: Cebrian, Manuel, et al.
Pubblicazione: (2024)
di: Cebrian, Manuel, et al.
Pubblicazione: (2024)
An exploration of features to improve the generalisability of fake news detection models
di: Hoy, Nathaniel, et al.
Pubblicazione: (2025)
di: Hoy, Nathaniel, et al.
Pubblicazione: (2025)
GETAE: Graph information Enhanced deep neural NeTwork ensemble ArchitecturE for fake news detection
di: Truică, Ciprian-Octavian, et al.
Pubblicazione: (2024)
di: Truică, Ciprian-Octavian, et al.
Pubblicazione: (2024)
Alignment faking in large language models
di: Greenblatt, Ryan, et al.
Pubblicazione: (2024)
di: Greenblatt, Ryan, et al.
Pubblicazione: (2024)
VelLMes: A high-interaction AI-based deception framework
di: Sladić, Muris, et al.
Pubblicazione: (2025)
di: Sladić, Muris, et al.
Pubblicazione: (2025)
Anti-adversarial Learning: Desensitizing Prompts for Large Language Models
di: Li, Xuan, et al.
Pubblicazione: (2025)
di: Li, Xuan, et al.
Pubblicazione: (2025)
Evaluating the World Model Implicit in a Generative Model
di: Vafa, Keyon, et al.
Pubblicazione: (2024)
di: Vafa, Keyon, et al.
Pubblicazione: (2024)
Multi-round jailbreak attack on large language models
di: Zhou, Yihua, et al.
Pubblicazione: (2024)
di: Zhou, Yihua, et al.
Pubblicazione: (2024)
OrderBkd: Textual backdoor attack through repositioning
di: Alekseevskaia, Irina, et al.
Pubblicazione: (2024)
di: Alekseevskaia, Irina, et al.
Pubblicazione: (2024)
Language Generation in the Limit
di: Kleinberg, Jon, et al.
Pubblicazione: (2024)
di: Kleinberg, Jon, et al.
Pubblicazione: (2024)
Sparse Autoencoders for Hypothesis Generation
di: Movva, Rajiv, et al.
Pubblicazione: (2025)
di: Movva, Rajiv, et al.
Pubblicazione: (2025)
From text to multimodal: a survey of adversarial example generation in question answering systems
di: Yigit, Gulsum, et al.
Pubblicazione: (2023)
di: Yigit, Gulsum, et al.
Pubblicazione: (2023)
COVID-19 Infodemic. Understanding content features in detecting fake news using a machine learning approach
di: Balakrishnan, Vimala, et al.
Pubblicazione: (2026)
di: Balakrishnan, Vimala, et al.
Pubblicazione: (2026)
Cognitive phantoms in LLMs through the lens of latent variables
di: Peereboom, Sanne, et al.
Pubblicazione: (2024)
di: Peereboom, Sanne, et al.
Pubblicazione: (2024)
Grammar and Gameplay-aligned RL for Game Description Generation with LLMs
di: Tanaka, Tsunehiko, et al.
Pubblicazione: (2025)
di: Tanaka, Tsunehiko, et al.
Pubblicazione: (2025)
Directed Graph-alignment Approach for Identification of Gaps in Short Answers
di: Sahu, Archana, et al.
Pubblicazione: (2025)
di: Sahu, Archana, et al.
Pubblicazione: (2025)
EDUMATH: Generating Standards-aligned Educational Math Word Problems
di: Christ, Bryan R., et al.
Pubblicazione: (2025)
di: Christ, Bryan R., et al.
Pubblicazione: (2025)
Strong and weak alignment of large language models with human values
di: Khamassi, Mehdi, et al.
Pubblicazione: (2024)
di: Khamassi, Mehdi, et al.
Pubblicazione: (2024)
Language models align with human judgments on key grammatical constructions
di: Hu, Jennifer, et al.
Pubblicazione: (2024)
di: Hu, Jennifer, et al.
Pubblicazione: (2024)
PBCAT: Patch-based composite adversarial training against physically realizable attacks on object detection
di: Li, Xiao, et al.
Pubblicazione: (2025)
di: Li, Xiao, et al.
Pubblicazione: (2025)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
di: Garg, Nikhil, et al.
Pubblicazione: (2026)
di: Garg, Nikhil, et al.
Pubblicazione: (2026)
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
di: Peng, Kenny, et al.
Pubblicazione: (2025)
di: Peng, Kenny, et al.
Pubblicazione: (2025)
Position: Stop Acting Like Language Model Agents Are Normal Agents
di: Perrier, Elija, et al.
Pubblicazione: (2025)
di: Perrier, Elija, et al.
Pubblicazione: (2025)
ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
di: Zhang, Xiaobo, et al.
Pubblicazione: (2025)
di: Zhang, Xiaobo, et al.
Pubblicazione: (2025)
RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
di: Wang, Bowen, et al.
Pubblicazione: (2025)
di: Wang, Bowen, et al.
Pubblicazione: (2025)
Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
di: Hu, Xiang, et al.
Pubblicazione: (2025)
di: Hu, Xiang, et al.
Pubblicazione: (2025)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
di: Hu, Mengxuan, et al.
Pubblicazione: (2026)
di: Hu, Mengxuan, et al.
Pubblicazione: (2026)
Analysis and prevention of AI-based phishing email attacks
di: Eze, Chibuike Samuel, et al.
Pubblicazione: (2024)
di: Eze, Chibuike Samuel, et al.
Pubblicazione: (2024)
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
di: Rozner, Joshua, et al.
Pubblicazione: (2026)
di: Rozner, Joshua, et al.
Pubblicazione: (2026)
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
di: Jing, Dong, et al.
Pubblicazione: (2025)
di: Jing, Dong, et al.
Pubblicazione: (2025)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
di: Sclar, Melanie, et al.
Pubblicazione: (2024)
di: Sclar, Melanie, et al.
Pubblicazione: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Effective faking of verbal deception detection with target‐aligned adversarial attacks
di: Bennett Kleinberg, et al.
Pubblicazione: (2025) -
When lies are mostly truthful: automated verbal deception detection for embedded lies
di: Loconte, Riccardo, et al.
Pubblicazione: (2025) -
Linguistic traces of stochastic empathy in language models
di: Kleinberg, Bennett, et al.
Pubblicazione: (2024) -
Humans incorrectly reject confident accusatory AI judgments
di: Loconte, Riccardo, et al.
Pubblicazione: (2025) -
Large Language Models in Cryptocurrency Securities Cases: Can a GPT Model Meaningfully Assist Lawyers?
di: Trozze, Arianna, et al.
Pubblicazione: (2023)