Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bianco, Francesca, Shiller, Derek |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
par: Chen, Jianhui, et autres
Publié: (2026)
par: Chen, Jianhui, et autres
Publié: (2026)
Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
par: Camassa, Carolina, et autres
Publié: (2026)
par: Camassa, Carolina, et autres
Publié: (2026)
On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse
par: Kalavasis, Alkis, et autres
Publié: (2024)
par: Kalavasis, Alkis, et autres
Publié: (2024)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
par: Tsujimura, Hikaru, et autres
Publié: (2025)
par: Tsujimura, Hikaru, et autres
Publié: (2025)
Mechanistic?
par: Saphra, Naomi, et autres
Publié: (2024)
par: Saphra, Naomi, et autres
Publié: (2024)
ACAR: Adaptive Complexity Routing for Multi-Model Ensembles with Auditable Decision Traces
par: Kumaresan, Ramchand
Publié: (2026)
par: Kumaresan, Ramchand
Publié: (2026)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
par: Habibi, Reza, et autres
Publié: (2026)
par: Habibi, Reza, et autres
Publié: (2026)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
par: Hammoud, Hasan Abed Al Kader, et autres
Publié: (2025)
par: Hammoud, Hasan Abed Al Kader, et autres
Publié: (2025)
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries
par: Hicke, Rebecca M. M., et autres
Publié: (2026)
par: Hicke, Rebecca M. M., et autres
Publié: (2026)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
par: Das, Rocktim Jyoti, et autres
Publié: (2023)
par: Das, Rocktim Jyoti, et autres
Publié: (2023)
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
par: Bouchard, Dylan
Publié: (2026)
par: Bouchard, Dylan
Publié: (2026)
Evaluating LLM Understanding via Structured Tabular Decision Simulations
par: Li, Sichao, et autres
Publié: (2025)
par: Li, Sichao, et autres
Publié: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
par: Salinas, David, et autres
Publié: (2025)
par: Salinas, David, et autres
Publié: (2025)
LLM-Powered Ensemble Learning for Paper Source Tracing: A GPU-Free Approach
par: Chen, Kunlong, et autres
Publié: (2024)
par: Chen, Kunlong, et autres
Publié: (2024)
Mechanistic Fine-tuning for In-context Learning
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
MIB: A Mechanistic Interpretability Benchmark
par: Mueller, Aaron, et autres
Publié: (2025)
par: Mueller, Aaron, et autres
Publié: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
par: Fan, Dongyang, et autres
Publié: (2025)
par: Fan, Dongyang, et autres
Publié: (2025)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
par: Xia, Fanzeng, et autres
Publié: (2024)
par: Xia, Fanzeng, et autres
Publié: (2024)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
par: Li, Zhe, et autres
Publié: (2025)
par: Li, Zhe, et autres
Publié: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
par: Landesberg, Eddie
Publié: (2026)
par: Landesberg, Eddie
Publié: (2026)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
par: Joshi, Abhinav, et autres
Publié: (2025)
par: Joshi, Abhinav, et autres
Publié: (2025)
LOLAMEME: Logic, Language, Memory, Mechanistic Framework
par: Desai, Jay, et autres
Publié: (2024)
par: Desai, Jay, et autres
Publié: (2024)
CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making
par: Ma, Mingyu Derek, et autres
Publié: (2024)
par: Ma, Mingyu Derek, et autres
Publié: (2024)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
par: Wang, Yiming, et autres
Publié: (2024)
par: Wang, Yiming, et autres
Publié: (2024)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
par: Sun, Jiuding, et autres
Publié: (2025)
par: Sun, Jiuding, et autres
Publié: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
par: Mishra, Anurag
Publié: (2025)
par: Mishra, Anurag
Publié: (2025)
Iteration Head: A Mechanistic Study of Chain-of-Thought
par: Cabannes, Vivien, et autres
Publié: (2024)
par: Cabannes, Vivien, et autres
Publié: (2024)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
par: Bozoukov, Matthew, et autres
Publié: (2025)
par: Bozoukov, Matthew, et autres
Publié: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
par: Méloux, Maxime, et autres
Publié: (2025)
par: Méloux, Maxime, et autres
Publié: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
par: Méloux, Maxime, et autres
Publié: (2025)
par: Méloux, Maxime, et autres
Publié: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
par: Alam, Firoj, et autres
Publié: (2026)
par: Alam, Firoj, et autres
Publié: (2026)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
par: Zhang, Shenao, et autres
Publié: (2025)
par: Zhang, Shenao, et autres
Publié: (2025)
Augmenting Legal Decision Support Systems with LLM-based NLI for Analyzing Social Media Evidence
par: Kadiyala, Ram Mohan Rao, et autres
Publié: (2024)
par: Kadiyala, Ram Mohan Rao, et autres
Publié: (2024)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
par: Kim, Jeonghye, et autres
Publié: (2025)
par: Kim, Jeonghye, et autres
Publié: (2025)
"Oh LLM, I'm Asking Thee, Please Give Me a Decision Tree": Zero-Shot Decision Tree Induction and Embedding with Large Language Models
par: Knauer, Ricardo, et autres
Publié: (2024)
par: Knauer, Ricardo, et autres
Publié: (2024)
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
par: Deng, Xinle, et autres
Publié: (2026)
par: Deng, Xinle, et autres
Publié: (2026)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
par: Song, Xiangchen, et autres
Publié: (2025)
par: Song, Xiangchen, et autres
Publié: (2025)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
par: Jain, Ojas, et autres
Publié: (2026)
par: Jain, Ojas, et autres
Publié: (2026)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
par: Wang, Zehong, et autres
Publié: (2026)
par: Wang, Zehong, et autres
Publié: (2026)
Documents similaires
-
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
par: Chen, Jianhui, et autres
Publié: (2026) -
Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
par: Camassa, Carolina, et autres
Publié: (2026) -
On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse
par: Kalavasis, Alkis, et autres
Publié: (2024) -
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
par: Tsujimura, Hikaru, et autres
Publié: (2025) -
Mechanistic?
par: Saphra, Naomi, et autres
Publié: (2024)