Too Big to Fool: Resisting Deception in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Samsami, Mohammad Reza, Richter, Mats Leon, Rodriguez, Juan, Thakkar, Megh, Chandar, Sarath, Gasse, Maxime |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
by: Guiroy, Simon, et al.
Published: (2025)
by: Guiroy, Simon, et al.
Published: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
Do Large Language Models Know How Much They Know?
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
Steering Large Language Model Activations in Sparse Spaces
by: Bayat, Reza, et al.
Published: (2025)
by: Bayat, Reza, et al.
Published: (2025)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
by: Lermen, Simon, et al.
Published: (2025)
by: Lermen, Simon, et al.
Published: (2025)
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
by: Prato, Gabriele, et al.
Published: (2023)
by: Prato, Gabriele, et al.
Published: (2023)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
by: Thakkar, Megh, et al.
Published: (2024)
by: Thakkar, Megh, et al.
Published: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
Deception Abilities Emerged in Large Language Models
by: Hagendorff, Thilo
Published: (2023)
by: Hagendorff, Thilo
Published: (2023)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
by: Aghajohari, Milad, et al.
Published: (2025)
by: Aghajohari, Milad, et al.
Published: (2025)
Small Encoders Can Rival Large Decoders in Detecting Groundedness
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
Language Models Optimized to Fool Detectors Still Have a Distinct Style (And How to Change It)
by: Soto, Rafael Rivera, et al.
Published: (2025)
by: Soto, Rafael Rivera, et al.
Published: (2025)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
by: Balestriero, Randall, et al.
Published: (2023)
by: Balestriero, Randall, et al.
Published: (2023)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
by: Drouin, Alexandre, et al.
Published: (2024)
by: Drouin, Alexandre, et al.
Published: (2024)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
by: Järviniemi, Olli, et al.
Published: (2024)
by: Järviniemi, Olli, et al.
Published: (2024)
Language Model Re-rankers are Fooled by Lexical Similarities
by: Hagström, Lovisa, et al.
Published: (2025)
by: Hagström, Lovisa, et al.
Published: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
by: Ibrahim, Adam, et al.
Published: (2024)
by: Ibrahim, Adam, et al.
Published: (2024)
Intelligent Switching for Reset-Free RL
by: Patil, Darshan, et al.
Published: (2024)
by: Patil, Darshan, et al.
Published: (2024)
Lookbehind-SAM: k steps back, 1 step forward
by: Mordido, Gonçalo, et al.
Published: (2023)
by: Mordido, Gonçalo, et al.
Published: (2023)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Post-edits Are Preferences Too
by: Berger, Nathaniel, et al.
Published: (2024)
by: Berger, Nathaniel, et al.
Published: (2024)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
by: Yao, Yang, et al.
Published: (2025)
by: Yao, Yang, et al.
Published: (2025)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
Reasoning in Large Language Models: A Geometric Perspective
by: Cosentino, Romain, et al.
Published: (2024)
by: Cosentino, Romain, et al.
Published: (2024)
To Tell The Truth: Language of Deception and Language Models
by: Hazra, Sanchaita, et al.
Published: (2023)
by: Hazra, Sanchaita, et al.
Published: (2023)
Leviathan: Decoupling Input and Output Representations in Language Models
by: Batley, Reza T., et al.
Published: (2026)
by: Batley, Reza T., et al.
Published: (2026)
Endogenous Resistance to Activation Steering in Language Models
by: McKenzie, Alex, et al.
Published: (2026)
by: McKenzie, Alex, et al.
Published: (2026)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Similar Items
-
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024) -
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
by: Guiroy, Simon, et al.
Published: (2025) -
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025) -
Do Large Language Models Know How Much They Know?
by: Prato, Gabriele, et al.
Published: (2025) -
Steering Large Language Model Activations in Sparse Spaces
by: Bayat, Reza, et al.
Published: (2025)