LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Olson, Matthew Lyle, Ratzlaff, Neale, Hinck, Musashi, Nguyen, Tri, Lal, Vasudev, Campbell, Joseph, Stepputtis, Simon, Tseng, Shao-Yen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering Large Language Models to Evaluate and Amplify Creativity
by: Olson, Matthew Lyle, et al.
Published: (2024)
by: Olson, Matthew Lyle, et al.
Published: (2024)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Probing Semantic Routing in Large Mixture-of-Expert Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
by: Hinck, Musashi, et al.
Published: (2024)
by: Hinck, Musashi, et al.
Published: (2024)
Why do LLaVA Vision-Language Models Reply to Images in English?
by: Hinck, Musashi, et al.
Published: (2024)
by: Hinck, Musashi, et al.
Published: (2024)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
by: Yu, Sungduk, et al.
Published: (2024)
by: Yu, Sungduk, et al.
Published: (2024)
Bayesian Social Deduction with Graph-Informed Language Models
by: Rahimirad, Shahab, et al.
Published: (2025)
by: Rahimirad, Shahab, et al.
Published: (2025)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
by: Luo, Man, et al.
Published: (2025)
by: Luo, Man, et al.
Published: (2025)
Multi-Agent Transfer Learning via Temporal Contrastive Learning
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments
by: Saenger, Till Raphael, et al.
Published: (2024)
by: Saenger, Till Raphael, et al.
Published: (2024)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
by: Egami, Naoki, et al.
Published: (2023)
by: Egami, Naoki, et al.
Published: (2023)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
by: Aflalo, Estelle, et al.
Published: (2024)
by: Aflalo, Estelle, et al.
Published: (2024)
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
by: Rosenman, Shachar, et al.
Published: (2023)
by: Rosenman, Shachar, et al.
Published: (2023)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
Lies, Lies, and Lies. On Truth, Dishonesty, Deception, and SelfDeception
by: Guido Löhrer
Published: (2022)
by: Guido Löhrer
Published: (2022)
HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation
by: Zhang, Ce, et al.
Published: (2024)
by: Zhang, Ce, et al.
Published: (2024)
FastRM: An efficient and automatic explainability framework for multimodal generative models
by: Stan, Gabriela Ben-Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben-Melech, et al.
Published: (2024)
Theory of Mind for Multi-Agent Collaboration via Large Language Models
by: Li, Huao, et al.
Published: (2023)
by: Li, Huao, et al.
Published: (2023)
Intentional Deception as Controllable Capability in LLM Agents
by: Starace, Jason, et al.
Published: (2026)
by: Starace, Jason, et al.
Published: (2026)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Model-Agnostic Policy Explanations with Large Language Models
by: Xi-Jia, Zhang, et al.
Published: (2025)
by: Xi-Jia, Zhang, et al.
Published: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
L-MAGIC: Language Model Assisted Generation of Images with Coherence
by: Cai, Zhipeng, et al.
Published: (2024)
by: Cai, Zhipeng, et al.
Published: (2024)
Utilizando Recursos Computacionais (Planilha) na Compreensão dos Números Racionais
by: Rosane Ratzlaff da Rosa
Published: (2008)
by: Rosane Ratzlaff da Rosa
Published: (2008)
ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition
by: Li, Samuel, et al.
Published: (2024)
by: Li, Samuel, et al.
Published: (2024)
DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents
by: Mukhopadhyay, Snehasis
Published: (2026)
by: Mukhopadhyay, Snehasis
Published: (2026)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
Symbolic Graph Inference for Compound Scene Understanding
by: Aryan, FNU, et al.
Published: (2024)
by: Aryan, FNU, et al.
Published: (2024)
Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning
by: Lin, Muhan, et al.
Published: (2025)
by: Lin, Muhan, et al.
Published: (2025)
Fuzzing with Agents? Generators Are All You Need
by: Vikram, Vasudev, et al.
Published: (2026)
by: Vikram, Vasudev, et al.
Published: (2026)
A Time Bomb Lies Buried
by: V. Lal, Brij
Published: (2013)
by: V. Lal, Brij
Published: (2013)
Interaction-Enabled Two- and Three-Fold Exceptional Points
by: Kato, Musashi, et al.
Published: (2026)
by: Kato, Musashi, et al.
Published: (2026)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
Similar Items
-
Steering Large Language Models to Evaluate and Amplify Creativity
by: Olson, Matthew Lyle, et al.
Published: (2024) -
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024) -
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025) -
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
by: Ratzlaff, Neale, et al.
Published: (2024) -
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
by: Olson, Matthew Lyle, et al.
Published: (2025)