Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
Fuente:
arXiv
Guardado en:
| Autor principal: | Keeman, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
por: Keeman, Michael
Publicado: (2026)
por: Keeman, Michael
Publicado: (2026)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
por: Mahale, Ajay Pravin
Publicado: (2026)
por: Mahale, Ajay Pravin
Publicado: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
por: Kubica, Dominick, et al.
Publicado: (2025)
por: Kubica, Dominick, et al.
Publicado: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Exemplar Retrieval Without Overhypothesis Induction: Limits of Distributional Sequence Learning in Early Word Learning
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
por: Hilasaca, Kenji, et al.
Publicado: (2026)
por: Hilasaca, Kenji, et al.
Publicado: (2026)
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law
por: Shportun, Nazarii
Publicado: (2026)
por: Shportun, Nazarii
Publicado: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
por: Vera, Patricio
Publicado: (2026)
por: Vera, Patricio
Publicado: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
por: Young, Richard J.
Publicado: (2026)
por: Young, Richard J.
Publicado: (2026)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
por: Yang, Jisoo, et al.
Publicado: (2026)
por: Yang, Jisoo, et al.
Publicado: (2026)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
por: Zhou, Yangyang, et al.
Publicado: (2026)
por: Zhou, Yangyang, et al.
Publicado: (2026)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
por: Shafique, Muhammad Ali, et al.
Publicado: (2026)
por: Shafique, Muhammad Ali, et al.
Publicado: (2026)
Eyla: Toward an Identity-Anchored LLM Architecture with Integrated Biological Priors -- Vision, Implementation Attempt, and Lessons from AI-Assisted Development
por: Aditto, Arif
Publicado: (2026)
por: Aditto, Arif
Publicado: (2026)
Truth as a Compression Artifact in Language Model Training
por: Krestnikov, Konstantin
Publicado: (2026)
por: Krestnikov, Konstantin
Publicado: (2026)
Machine Unlearning for Masked Diffusion Language Models
por: Lee, Georu, et al.
Publicado: (2026)
por: Lee, Georu, et al.
Publicado: (2026)
KAConvText: Novel Approach to Burmese Sentence Classification using Kolmogorov-Arnold Convolution
por: Thu, Ye Kyaw, et al.
Publicado: (2025)
por: Thu, Ye Kyaw, et al.
Publicado: (2025)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
por: Agarwal, Mahak, et al.
Publicado: (2025)
por: Agarwal, Mahak, et al.
Publicado: (2025)
A Hierarchical Error Framework for Reliable Automated Coding in Communication Research: Applications to Health and Political Communication
por: Zhao, Zhilong, et al.
Publicado: (2025)
por: Zhao, Zhilong, et al.
Publicado: (2025)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
por: Bouchekif, Abdessalam, et al.
Publicado: (2025)
por: Bouchekif, Abdessalam, et al.
Publicado: (2025)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
por: Zhang, Yuxuan
Publicado: (2025)
por: Zhang, Yuxuan
Publicado: (2025)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
por: Pratyush, Spandan
Publicado: (2026)
por: Pratyush, Spandan
Publicado: (2026)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
por: Yordanov, Yordan, et al.
Publicado: (2026)
por: Yordanov, Yordan, et al.
Publicado: (2026)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
por: Lan, Guangchen, et al.
Publicado: (2025)
por: Lan, Guangchen, et al.
Publicado: (2025)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
por: Zhou, Tianyang, et al.
Publicado: (2026)
por: Zhou, Tianyang, et al.
Publicado: (2026)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
por: Manczak, Blazej, et al.
Publicado: (2025)
por: Manczak, Blazej, et al.
Publicado: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
por: Gwak, Jiho, et al.
Publicado: (2025)
por: Gwak, Jiho, et al.
Publicado: (2025)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
por: Kumar, Sachin
Publicado: (2026)
por: Kumar, Sachin
Publicado: (2026)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
por: Kajare, Prajwal Vijay, et al.
Publicado: (2026)
por: Kajare, Prajwal Vijay, et al.
Publicado: (2026)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
por: Zhang, Luyan
Publicado: (2025)
por: Zhang, Luyan
Publicado: (2025)
Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation
por: Lyu, Bohan, et al.
Publicado: (2024)
por: Lyu, Bohan, et al.
Publicado: (2024)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
por: Li, Zilin, et al.
Publicado: (2026)
por: Li, Zilin, et al.
Publicado: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
por: Resck, Lucas, et al.
Publicado: (2026)
por: Resck, Lucas, et al.
Publicado: (2026)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
por: Cui, Jian, et al.
Publicado: (2026)
por: Cui, Jian, et al.
Publicado: (2026)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
por: Sauter, Andreas, et al.
Publicado: (2026)
por: Sauter, Andreas, et al.
Publicado: (2026)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
por: Heydari, Sina, et al.
Publicado: (2026)
por: Heydari, Sina, et al.
Publicado: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
por: Jehu-Appiah, Rodney
Publicado: (2026)
por: Jehu-Appiah, Rodney
Publicado: (2026)
Ejemplares similares
-
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
por: Keeman, Michael
Publicado: (2026) -
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
por: Mahale, Ajay Pravin
Publicado: (2026) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025) -
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
por: Kubica, Dominick, et al.
Publicado: (2025) -
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)