Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
Fuente:
arXiv
Saved in:
| Main Authors: | Ortu, Francesco, Jin, Zhijing, Doimo, Diego, Sachan, Mrinmaya, Cazzaniga, Alberto, Schölkopf, Bernhard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
by: Ortu, Francesco, et al.
Published: (2025)
by: Ortu, Francesco, et al.
Published: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
by: Dotsinski, Asen, et al.
Published: (2025)
by: Dotsinski, Asen, et al.
Published: (2025)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
by: Jenny, David F., et al.
Published: (2023)
by: Jenny, David F., et al.
Published: (2023)
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
by: Lyu, Zhiheng, et al.
Published: (2024)
by: Lyu, Zhiheng, et al.
Published: (2024)
Language Model Alignment in Multilingual Trolley Problems
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
by: Piatti, Giorgio, et al.
Published: (2024)
by: Piatti, Giorgio, et al.
Published: (2024)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)
by: Kumar, Ishan, et al.
Published: (2023)
The representation landscape of few-shot learning and fine-tuning in large language models
by: Doimo, Diego, et al.
Published: (2024)
by: Doimo, Diego, et al.
Published: (2024)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
by: Ortu, Francesco, et al.
Published: (2026)
by: Ortu, Francesco, et al.
Published: (2026)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)
by: Basile, Lorenzo, et al.
Published: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
by: Kassem, Aly M., et al.
Published: (2025)
by: Kassem, Aly M., et al.
Published: (2025)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
by: Cui, Peng, et al.
Published: (2025)
by: Cui, Peng, et al.
Published: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Learning to Reason Efficiently with A* Post-Training
by: Opedal, Andreas, et al.
Published: (2026)
by: Opedal, Andreas, et al.
Published: (2026)
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
On Affine Homotopy between Language Encoders
by: Chan, Robin SM, et al.
Published: (2024)
by: Chan, Robin SM, et al.
Published: (2024)
Probing for Arithmetic Errors in Language Models
by: Sun, Yucheng, et al.
Published: (2025)
by: Sun, Yucheng, et al.
Published: (2025)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
by: Lalwani, Abhinav, et al.
Published: (2024)
by: Lalwani, Abhinav, et al.
Published: (2024)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
Are Language Models Efficient Reasoners? A Perspective from Logic Programming
by: Opedal, Andreas, et al.
Published: (2025)
by: Opedal, Andreas, et al.
Published: (2025)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
by: Cui, Shaobo, et al.
Published: (2024)
by: Cui, Shaobo, et al.
Published: (2024)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Efficiently Computing Susceptibility to Context in Language Models
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
Do Vision-Language Models Really Understand Visual Language?
by: Hou, Yifan, et al.
Published: (2024)
by: Hou, Yifan, et al.
Published: (2024)
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025)
by: Samway, Keenan, et al.
Published: (2025)
What Do Language Models Learn in Context? The Structured Task Hypothesis
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026)
by: Samway, Keenan, et al.
Published: (2026)
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
by: Serra, Alessandro Pietro, et al.
Published: (2024)
by: Serra, Alessandro Pietro, et al.
Published: (2024)
Grammar Control in Dialogue Response Generation for Language Learning Chatbots
by: Glandorf, Dominik, et al.
Published: (2025)
by: Glandorf, Dominik, et al.
Published: (2025)
Similar Items
-
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
by: Ortu, Francesco, et al.
Published: (2025) -
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025) -
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
by: Dotsinski, Asen, et al.
Published: (2025) -
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
by: Jenny, David F., et al.
Published: (2023) -
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025)