Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Basu, Abhinaba, Chakraborty, Pavan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
von: Sela, Omer
Veröffentlicht: (2026)
von: Sela, Omer
Veröffentlicht: (2026)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
von: Zhang, Xue
Veröffentlicht: (2025)
von: Zhang, Xue
Veröffentlicht: (2025)
ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026)
von: Jia, Xiao
Veröffentlicht: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
von: Wang, Zhixiang
Veröffentlicht: (2025)
von: Wang, Zhixiang
Veröffentlicht: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
von: Xia, Bowei, et al.
Veröffentlicht: (2026)
von: Xia, Bowei, et al.
Veröffentlicht: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
von: Pizzo, David Alejandro Trejo
Veröffentlicht: (2026)
von: Pizzo, David Alejandro Trejo
Veröffentlicht: (2026)
Measuring Intent Comprehension in LLMs
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2025)
von: Kunievsky, Nadav, et al.
Veröffentlicht: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2026)
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
Improving Commonsense Bias Classification by Mitigating the Influence of Demographic Terms
von: Lee, JinKyu, et al.
Veröffentlicht: (2024)
von: Lee, JinKyu, et al.
Veröffentlicht: (2024)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
Towards Ontology-Enhanced Representation Learning for Large Language Models
von: Ronzano, Francesco, et al.
Veröffentlicht: (2024)
von: Ronzano, Francesco, et al.
Veröffentlicht: (2024)
Mitigating Position-Shift Failures in Text-Based Modular Arithmetic via Position Curriculum and Template Diversity
von: Yudin, Nikolay
Veröffentlicht: (2026)
von: Yudin, Nikolay
Veröffentlicht: (2026)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
von: Juzek, Tom S., et al.
Veröffentlicht: (2025)
von: Juzek, Tom S., et al.
Veröffentlicht: (2025)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
von: Gupta, Aayush
Veröffentlicht: (2025)
von: Gupta, Aayush
Veröffentlicht: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
Classifier-Augmented Generation for Structured Workflow Prediction
von: Gschwind, Thomas, et al.
Veröffentlicht: (2025)
von: Gschwind, Thomas, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
von: Kaiser, Daniel, et al.
Veröffentlicht: (2026)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2026)
Monotonicity as an Architectural Bias for Robust Language Models
von: Cooper, Patrick, et al.
Veröffentlicht: (2026)
von: Cooper, Patrick, et al.
Veröffentlicht: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
von: Basu, Abhinaba
Veröffentlicht: (2026) -
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026) -
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026) -
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)