Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bozoukov, Matthew, Nguyen, Matthew, Singh, Shubkarman, Bussmann, Bart, Leask, Patrick |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
par: Roytburg, Dani, et autres
Publié: (2025)
par: Roytburg, Dani, et autres
Publié: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
par: Roytburg, Dani, et autres
Publié: (2026)
par: Roytburg, Dani, et autres
Publié: (2026)
BatchTopK Sparse Autoencoders
par: Bussmann, Bart, et autres
Publié: (2024)
par: Bussmann, Bart, et autres
Publié: (2024)
Guided Self-Evolving LLMs with Minimal Human Supervision
par: Yu, Wenhao, et autres
Publié: (2025)
par: Yu, Wenhao, et autres
Publié: (2025)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
par: Lee, Yu-Ting, et autres
Publié: (2025)
par: Lee, Yu-Ting, et autres
Publié: (2025)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
par: Reddy, Avinash, et autres
Publié: (2026)
par: Reddy, Avinash, et autres
Publié: (2026)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
par: Beniwal, Himanshu, et autres
Publié: (2026)
par: Beniwal, Himanshu, et autres
Publié: (2026)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
par: Liu, Liangxin, et autres
Publié: (2024)
par: Liu, Liangxin, et autres
Publié: (2024)
Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
par: Guan, Zihan, et autres
Publié: (2026)
par: Guan, Zihan, et autres
Publié: (2026)
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise
par: Khoriaty, Matthew, et autres
Publié: (2026)
par: Khoriaty, Matthew, et autres
Publié: (2026)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
par: Han, Pengrui, et autres
Publié: (2025)
par: Han, Pengrui, et autres
Publié: (2025)
Mechanistic?
par: Saphra, Naomi, et autres
Publié: (2024)
par: Saphra, Naomi, et autres
Publié: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
par: Du, Hongzhe, et autres
Publié: (2025)
par: Du, Hongzhe, et autres
Publié: (2025)
On the Performance of LLMs for Real Estate Appraisal
par: Geerts, Margot, et autres
Publié: (2025)
par: Geerts, Margot, et autres
Publié: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
par: Tran, Nguyen Phuc, et autres
Publié: (2025)
par: Tran, Nguyen Phuc, et autres
Publié: (2025)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
par: Kutakh, Matthew
Publié: (2026)
par: Kutakh, Matthew
Publié: (2026)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
par: Leask, Patrick, et autres
Publié: (2025)
par: Leask, Patrick, et autres
Publié: (2025)
Interpreting the Effects of Quantization on LLMs
par: Singh, Manpreet, et autres
Publié: (2025)
par: Singh, Manpreet, et autres
Publié: (2025)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
par: Xiong, Boya, et autres
Publié: (2025)
par: Xiong, Boya, et autres
Publié: (2025)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
par: Yang, Hongming, et autres
Publié: (2025)
par: Yang, Hongming, et autres
Publié: (2025)
Low-Rank Quantization-Aware Training for LLMs
par: Bondarenko, Yelysei, et autres
Publié: (2024)
par: Bondarenko, Yelysei, et autres
Publié: (2024)
Transformer-Squared: Self-adaptive LLMs
par: Sun, Qi, et autres
Publié: (2025)
par: Sun, Qi, et autres
Publié: (2025)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
par: Nainani, Jatin, et autres
Publié: (2024)
par: Nainani, Jatin, et autres
Publié: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
par: Smith, Matthew L., et autres
Publié: (2026)
par: Smith, Matthew L., et autres
Publié: (2026)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
par: Yang, Zhuonan, et autres
Publié: (2026)
par: Yang, Zhuonan, et autres
Publié: (2026)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
par: Xu, Tianyang, et autres
Publié: (2024)
par: Xu, Tianyang, et autres
Publié: (2024)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
par: Finlayson, Matthew, et autres
Publié: (2025)
par: Finlayson, Matthew, et autres
Publié: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
par: Li, He, et autres
Publié: (2024)
par: Li, He, et autres
Publié: (2024)
Certifying Knowledge Comprehension in LLMs
par: Chaudhary, Isha, et autres
Publié: (2024)
par: Chaudhary, Isha, et autres
Publié: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
par: Sheshadri, Abhay, et autres
Publié: (2024)
par: Sheshadri, Abhay, et autres
Publié: (2024)
Mechanistic Fine-tuning for In-context Learning
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
MIB: A Mechanistic Interpretability Benchmark
par: Mueller, Aaron, et autres
Publié: (2025)
par: Mueller, Aaron, et autres
Publié: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
par: Singh, Janvijay, et autres
Publié: (2026)
par: Singh, Janvijay, et autres
Publié: (2026)
Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective
par: Mao, Liyuan, et autres
Publié: (2026)
par: Mao, Liyuan, et autres
Publié: (2026)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
par: Yang, Matthew Y. R., et autres
Publié: (2026)
par: Yang, Matthew Y. R., et autres
Publié: (2026)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
par: Laine, Rudolf, et autres
Publié: (2024)
par: Laine, Rudolf, et autres
Publié: (2024)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
par: Wang, Qibin, et autres
Publié: (2025)
par: Wang, Qibin, et autres
Publié: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
par: Chaudhary, Maheep, et autres
Publié: (2025)
par: Chaudhary, Maheep, et autres
Publié: (2025)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
par: Schiekiera, Louis, et autres
Publié: (2026)
par: Schiekiera, Louis, et autres
Publié: (2026)
Documents similaires
-
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
par: Roytburg, Dani, et autres
Publié: (2025) -
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
par: Roytburg, Dani, et autres
Publié: (2026) -
BatchTopK Sparse Autoencoders
par: Bussmann, Bart, et autres
Publié: (2024) -
Guided Self-Evolving LLMs with Minimal Human Supervision
par: Yu, Wenhao, et autres
Publié: (2025) -
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
par: Lee, Yu-Ting, et autres
Publié: (2025)