Semantic Deception: When Reasoning Models Can't Compute an Addition
Fuente:
arXiv
Saved in:
| Main Authors: | de Leeuw, Nathaniël, Nahon, Marceau, Reymond, Mathis, Chatila, Raja, Khamassi, Mehdi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024)
by: Khamassi, Mehdi, et al.
Published: (2024)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
by: Morlat, Geoffroy, et al.
Published: (2025)
by: Morlat, Geoffroy, et al.
Published: (2025)
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
by: Chen, Kang, et al.
Published: (2024)
by: Chen, Kang, et al.
Published: (2024)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
by: Li, Bangzheng, et al.
Published: (2023)
by: Li, Bangzheng, et al.
Published: (2023)
Can't say cant? Measuring and Reasoning of Dark Jargons in Large Language Models
by: Ji, Xu, et al.
Published: (2024)
by: Ji, Xu, et al.
Published: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
You Can't Fight in Here! This is BBS!
by: Futrell, Richard, et al.
Published: (2026)
by: Futrell, Richard, et al.
Published: (2026)
AbsenceBench: Language Models Can't Tell What's Missing
by: Fu, Harvey Yiyun, et al.
Published: (2025)
by: Fu, Harvey Yiyun, et al.
Published: (2025)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
by: Kamath, Amita, et al.
Published: (2026)
by: Kamath, Amita, et al.
Published: (2026)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
by: Young, Richard J., et al.
Published: (2025)
by: Young, Richard J., et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
Can Factual Statements be Deceptive? The DeFaBel Corpus of Belief-based Deception
by: Velutharambath, Aswathy, et al.
Published: (2024)
by: Velutharambath, Aswathy, et al.
Published: (2024)
Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models
by: Min, Dehai, et al.
Published: (2026)
by: Min, Dehai, et al.
Published: (2026)
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
by: Dash, Saloni, et al.
Published: (2025)
by: Dash, Saloni, et al.
Published: (2025)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
by: Liu, Yongkang, et al.
Published: (2023)
by: Liu, Yongkang, et al.
Published: (2023)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
We Can't Understand AI Using our Existing Vocabulary
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
by: Prakash, Nirmalendu, et al.
Published: (2025)
by: Prakash, Nirmalendu, et al.
Published: (2025)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
by: Baldelli, Davide, et al.
Published: (2026)
by: Baldelli, Davide, et al.
Published: (2026)
When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning
by: Zhu, Rongzhi, et al.
Published: (2025)
by: Zhu, Rongzhi, et al.
Published: (2025)
AdaptThink: Reasoning Models Can Learn When to Think
by: Zhang, Jiajie, et al.
Published: (2025)
by: Zhang, Jiajie, et al.
Published: (2025)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
by: Merrill, Scott, et al.
Published: (2026)
by: Merrill, Scott, et al.
Published: (2026)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
by: Rani, Anku, et al.
Published: (2023)
by: Rani, Anku, et al.
Published: (2023)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
by: Song, Maojia, et al.
Published: (2025)
by: Song, Maojia, et al.
Published: (2025)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
by: Liang, Tuo, et al.
Published: (2025)
by: Liang, Tuo, et al.
Published: (2025)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
by: Riachi, Roland, et al.
Published: (2025)
by: Riachi, Roland, et al.
Published: (2025)
An Assessment of Model-On-Model Deception
by: Heitkoetter, Julius, et al.
Published: (2024)
by: Heitkoetter, Julius, et al.
Published: (2024)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
Can LLMs Compute with Reasons?
by: Sandilya, Harshit, et al.
Published: (2024)
by: Sandilya, Harshit, et al.
Published: (2024)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
by: Khalifa, Muhammad, et al.
Published: (2024)
by: Khalifa, Muhammad, et al.
Published: (2024)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
by: Liu, Yanjiang, et al.
Published: (2026)
by: Liu, Yanjiang, et al.
Published: (2026)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2024)
When to Reason: Semantic Router for vLLM
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Make an Offer They Can't Refuse: Grounding Bayesian Persuasion in Real-World Dialogues without Pre-Commitment
by: He, Buwei, et al.
Published: (2025)
by: He, Buwei, et al.
Published: (2025)
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Similar Items
-
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024) -
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
by: Morlat, Geoffroy, et al.
Published: (2025) -
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
by: Chen, Kang, et al.
Published: (2024) -
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
by: Li, Bangzheng, et al.
Published: (2023) -
Can't say cant? Measuring and Reasoning of Dark Jargons in Large Language Models
by: Ji, Xu, et al.
Published: (2024)