Caught in the Act: a mechanistic approach to detecting deception
Fuente:
arXiv
Saved in:
| Main Authors: | Boxo, Gerard, Socha, Ryan, Yoo, Daniel, Raval, Shivam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
by: Boxo, Gerard, et al.
Published: (2025)
by: Boxo, Gerard, et al.
Published: (2025)
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
by: Gulati, Idhant, et al.
Published: (2026)
by: Gulati, Idhant, et al.
Published: (2026)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
by: Maltbie, Benjamin, et al.
Published: (2026)
by: Maltbie, Benjamin, et al.
Published: (2026)
Effective faking of verbal deception detection with target-aligned adversarial attacks
by: Kleinberg, Bennett, et al.
Published: (2025)
by: Kleinberg, Bennett, et al.
Published: (2025)
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026)
by: Dawes, Cutter, et al.
Published: (2026)
SMART-OC: A Real-time Time-risk Optimal Replanning Algorithm for Dynamic Obstacles and Spatio-temporally Varying Currents
by: Raval, Reema, et al.
Published: (2025)
by: Raval, Reema, et al.
Published: (2025)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
LLM is Not All You Need: A Systematic Evaluation of ML vs. Foundation Models for text and image based Medical Classification
by: Raval, Meet, et al.
Published: (2026)
by: Raval, Meet, et al.
Published: (2026)
Optimal sensor deception in stochastic environments with partial observability to mislead a robot to a decoy goal
by: Rahmani, Hazhar, et al.
Published: (2025)
by: Rahmani, Hazhar, et al.
Published: (2025)
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
by: Yun, Hye Sun, et al.
Published: (2025)
by: Yun, Hye Sun, et al.
Published: (2025)
ActDroid: An active learning framework for Android malware detection
by: Muzaffar, Ali, et al.
Published: (2024)
by: Muzaffar, Ali, et al.
Published: (2024)
Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs
by: Dubey, Shivam
Published: (2025)
by: Dubey, Shivam
Published: (2025)
Multi-Domain ABSA Conversation Dataset Generation via LLMs for Real-World Evaluation and Model Comparison
by: Pandit, Tejul, et al.
Published: (2025)
by: Pandit, Tejul, et al.
Published: (2025)
VelLMes: A high-interaction AI-based deception framework
by: Sladić, Muris, et al.
Published: (2025)
by: Sladić, Muris, et al.
Published: (2025)
An introduction to graphical tensor notation for mechanistic interpretability
by: Taylor, Jordan K.
Published: (2024)
by: Taylor, Jordan K.
Published: (2024)
Plant detection from ultra high resolution remote sensing images: A Semantic Segmentation approach based on fuzzy loss
by: Pande, Shivam, et al.
Published: (2024)
by: Pande, Shivam, et al.
Published: (2024)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
by: Hryszko, Jaroslaw
Published: (2026)
by: Hryszko, Jaroslaw
Published: (2026)
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
by: Kang, Jiazheng, et al.
Published: (2026)
by: Kang, Jiazheng, et al.
Published: (2026)
Focused ReAct: Improving ReAct through Reiterate and Early Stop
by: Li, Shuoqiu, et al.
Published: (2024)
by: Li, Shuoqiu, et al.
Published: (2024)
The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
by: Pandit, Tejul, et al.
Published: (2025)
by: Pandit, Tejul, et al.
Published: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
A technical note for the 91-clauses SAT resolution with Indirect QAOA based approach
by: Fleury, Gerard, et al.
Published: (2024)
by: Fleury, Gerard, et al.
Published: (2024)
Attention-gated U-Net model for semantic segmentation of brain tumors and feature extraction for survival prognosis
by: Pate, Rut, et al.
Published: (2026)
by: Pate, Rut, et al.
Published: (2026)
Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music
by: Chauhan, Shivam, et al.
Published: (2026)
by: Chauhan, Shivam, et al.
Published: (2026)
Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment
by: Kitkana, Chayanon, et al.
Published: (2026)
by: Kitkana, Chayanon, et al.
Published: (2026)
When to Act: Calibrated Confidence for Reliable Human Intention Prediction in Assistive Robotics
by: Gaus, Johannes A., et al.
Published: (2026)
by: Gaus, Johannes A., et al.
Published: (2026)
Joint data imputation and mechanistic modelling for simulating heart-brain interactions in incomplete datasets
by: Banus, Jaume, et al.
Published: (2020)
by: Banus, Jaume, et al.
Published: (2020)
ProAct: Agentic Lookahead in Interactive Environments
by: Yu, Yangbin, et al.
Published: (2026)
by: Yu, Yangbin, et al.
Published: (2026)
The Reasons that Agents Act: Intention and Instrumental Goals
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
by: Ratnakar, Shivam, et al.
Published: (2025)
by: Ratnakar, Shivam, et al.
Published: (2025)
Mesh-based Super-resolution of Detonation Flows with Multiscale Graph Transformers
by: Barwey, Shivam, et al.
Published: (2025)
by: Barwey, Shivam, et al.
Published: (2025)
Beyond Winning: Margin of Victory Relative to Expectation Unlocks Accurate Skill Ratings
by: Shorewala, Shivam, et al.
Published: (2025)
by: Shorewala, Shivam, et al.
Published: (2025)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
by: Zhao, Enyu, et al.
Published: (2025)
by: Zhao, Enyu, et al.
Published: (2025)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
A BERT-Based Summarization approach for depression detection
by: Gavalan, Hossein Salahshoor, et al.
Published: (2024)
by: Gavalan, Hossein Salahshoor, et al.
Published: (2024)
cotomi Act: Learning to Automate Work by Watching You
by: Oyamada, Masafumi, et al.
Published: (2026)
by: Oyamada, Masafumi, et al.
Published: (2026)
Learning to Act without Actions
by: Schmidt, Dominik, et al.
Published: (2023)
by: Schmidt, Dominik, et al.
Published: (2023)
Similar Items
-
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
by: Boxo, Gerard, et al.
Published: (2025) -
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
by: Gulati, Idhant, et al.
Published: (2026) -
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
by: Maltbie, Benjamin, et al.
Published: (2026) -
Effective faking of verbal deception detection with target-aligned adversarial attacks
by: Kleinberg, Bennett, et al.
Published: (2025) -
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026)