The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sahoo, Subramanyam, Chadha, Aman, Jain, Vinija, Chaudhary, Divya |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
par: Sinha, Neelabh, et autres
Publié: (2024)
par: Sinha, Neelabh, et autres
Publié: (2024)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
par: Singh, Smriti, et autres
Publié: (2024)
par: Singh, Smriti, et autres
Publié: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
par: Sinha, Neelabh, et autres
Publié: (2024)
par: Sinha, Neelabh, et autres
Publié: (2024)
How Culturally Aware are Vision-Language Models?
par: Burda-Lassen, Olena, et autres
Publié: (2024)
par: Burda-Lassen, Olena, et autres
Publié: (2024)
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
par: Sahoo, Subramanyam, et autres
Publié: (2025)
par: Sahoo, Subramanyam, et autres
Publié: (2025)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
par: Sahoo, Subramanyam, et autres
Publié: (2025)
par: Sahoo, Subramanyam, et autres
Publié: (2025)
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training
par: Kumar, Harsh, et autres
Publié: (2026)
par: Kumar, Harsh, et autres
Publié: (2026)
Decoding the Diversity: A Review of the Indic AI Research Landscape
par: KJ, Sankalp, et autres
Publié: (2024)
par: KJ, Sankalp, et autres
Publié: (2024)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
par: Balne, Charith Chandra Sai, et autres
Publié: (2024)
par: Balne, Charith Chandra Sai, et autres
Publié: (2024)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
par: Das, Arion, et autres
Publié: (2026)
par: Das, Arion, et autres
Publié: (2026)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
par: Sahoo, Subramanyam
Publié: (2026)
par: Sahoo, Subramanyam
Publié: (2026)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
par: Sahoo, Pranab, et autres
Publié: (2024)
par: Sahoo, Pranab, et autres
Publié: (2024)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
par: Saha, Partha Pratim, et autres
Publié: (2026)
par: Saha, Partha Pratim, et autres
Publié: (2026)
The Controllability Trap: A Governance Framework for Military AI Agents
par: Sahoo, Subramanyam
Publié: (2026)
par: Sahoo, Subramanyam
Publié: (2026)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
par: Aggarwal, Yash, et autres
Publié: (2026)
par: Aggarwal, Yash, et autres
Publié: (2026)
Deliberative Alignment: Reasoning Enables Safer Language Models
par: Guan, Melody Y., et autres
Publié: (2024)
par: Guan, Melody Y., et autres
Publié: (2024)
Overview of Factify5WQA: Fact Verification through 5W Question-Answering
par: Suresh, Suryavardan, et autres
Publié: (2024)
par: Suresh, Suryavardan, et autres
Publié: (2024)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
par: Islam, Tunazzina
Publié: (2026)
par: Islam, Tunazzina
Publié: (2026)
Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
par: Zhao, Yibo, et autres
Publié: (2025)
par: Zhao, Yibo, et autres
Publié: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
par: AlDahoul, Nouar, et autres
Publié: (2025)
par: AlDahoul, Nouar, et autres
Publié: (2025)
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
par: Patil, Avinash, et autres
Publié: (2025)
par: Patil, Avinash, et autres
Publié: (2025)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
par: Das, Amitava, et autres
Publié: (2025)
par: Das, Amitava, et autres
Publié: (2025)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
par: Bai, Xiaoyan, et autres
Publié: (2026)
par: Bai, Xiaoyan, et autres
Publié: (2026)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
par: Yagnik, Suryash, et autres
Publié: (2026)
par: Yagnik, Suryash, et autres
Publié: (2026)
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
par: Yu, Jeffy, et autres
Publié: (2024)
par: Yu, Jeffy, et autres
Publié: (2024)
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
par: Sahoo, Pranab, et autres
Publié: (2024)
par: Sahoo, Pranab, et autres
Publié: (2024)
Situational Awareness Matters in 3D Vision Language Reasoning
par: Man, Yunze, et autres
Publié: (2024)
par: Man, Yunze, et autres
Publié: (2024)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
par: Krishnappa, Pushwitha, et autres
Publié: (2026)
par: Krishnappa, Pushwitha, et autres
Publié: (2026)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
par: Saha, Anusa, et autres
Publié: (2026)
par: Saha, Anusa, et autres
Publié: (2026)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
par: Khoshnoodi, Mahsa, et autres
Publié: (2024)
par: Khoshnoodi, Mahsa, et autres
Publié: (2024)
Multilingual State Space Models for Structured Question Answering in Indic Languages
par: Vats, Arpita, et autres
Publié: (2025)
par: Vats, Arpita, et autres
Publié: (2025)
Mars: Situated Inductive Reasoning in an Open-World Environment
par: Tang, Xiaojuan, et autres
Publié: (2024)
par: Tang, Xiaojuan, et autres
Publié: (2024)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
par: Joshi, Abhinav, et autres
Publié: (2025)
par: Joshi, Abhinav, et autres
Publié: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
par: Lin, Bill Yuchen, et autres
Publié: (2025)
par: Lin, Bill Yuchen, et autres
Publié: (2025)
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
par: Shin, Kwan Soo
Publié: (2026)
par: Shin, Kwan Soo
Publié: (2026)
Learning to Reason with Mixture of Tokens
par: Jain, Adit, et autres
Publié: (2025)
par: Jain, Adit, et autres
Publié: (2025)
Documents similaires
-
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
par: Sahoo, Subramanyam, et autres
Publié: (2026) -
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
par: Sahoo, Subramanyam, et autres
Publié: (2026) -
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
par: Sahoo, Subramanyam, et autres
Publié: (2026) -
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
par: Sahoo, Subramanyam, et autres
Publié: (2026) -
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
par: Sinha, Neelabh, et autres
Publié: (2024)