Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
Fuente:
arXiv
Salvato in:
| Autori principali: | Ortu, Francesco, Jin, Zhijing, Doimo, Diego, Sachan, Mrinmaya, Cazzaniga, Alberto, Schölkopf, Bernhard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
di: Ortu, Francesco, et al.
Pubblicazione: (2025)
di: Ortu, Francesco, et al.
Pubblicazione: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025)
di: Simko, Samuel, et al.
Pubblicazione: (2025)
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
di: Dotsinski, Asen, et al.
Pubblicazione: (2025)
di: Dotsinski, Asen, et al.
Pubblicazione: (2025)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
di: Jenny, David F., et al.
Pubblicazione: (2023)
di: Jenny, David F., et al.
Pubblicazione: (2023)
Are LLMs Good Safety Agents or a Propaganda Engine?
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
di: Lyu, Zhiheng, et al.
Pubblicazione: (2024)
di: Lyu, Zhiheng, et al.
Pubblicazione: (2024)
Language Model Alignment in Multilingual Trolley Problems
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
di: Piatti, Giorgio, et al.
Pubblicazione: (2024)
di: Piatti, Giorgio, et al.
Pubblicazione: (2024)
Implicit Personalization in Language Models: A Systematic Study
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
CausalCite: A Causal Formulation of Paper Citations
di: Kumar, Ishan, et al.
Pubblicazione: (2023)
di: Kumar, Ishan, et al.
Pubblicazione: (2023)
The representation landscape of few-shot learning and fine-tuning in large language models
di: Doimo, Diego, et al.
Pubblicazione: (2024)
di: Doimo, Diego, et al.
Pubblicazione: (2024)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
di: Ortu, Francesco, et al.
Pubblicazione: (2026)
di: Ortu, Francesco, et al.
Pubblicazione: (2026)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
di: Basile, Lorenzo, et al.
Pubblicazione: (2025)
di: Basile, Lorenzo, et al.
Pubblicazione: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
di: Kassem, Aly M., et al.
Pubblicazione: (2025)
di: Kassem, Aly M., et al.
Pubblicazione: (2025)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
di: Cui, Peng, et al.
Pubblicazione: (2025)
di: Cui, Peng, et al.
Pubblicazione: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
di: Opedal, Andreas, et al.
Pubblicazione: (2024)
di: Opedal, Andreas, et al.
Pubblicazione: (2024)
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
di: Zouhar, Vilém, et al.
Pubblicazione: (2025)
di: Zouhar, Vilém, et al.
Pubblicazione: (2025)
Learning to Reason Efficiently with A* Post-Training
di: Opedal, Andreas, et al.
Pubblicazione: (2026)
di: Opedal, Andreas, et al.
Pubblicazione: (2026)
CLadder: Assessing Causal Reasoning in Language Models
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
di: Ozyurt, Yilmazcan, et al.
Pubblicazione: (2024)
di: Ozyurt, Yilmazcan, et al.
Pubblicazione: (2024)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
On Affine Homotopy between Language Encoders
di: Chan, Robin SM, et al.
Pubblicazione: (2024)
di: Chan, Robin SM, et al.
Pubblicazione: (2024)
Probing for Arithmetic Errors in Language Models
di: Sun, Yucheng, et al.
Pubblicazione: (2025)
di: Sun, Yucheng, et al.
Pubblicazione: (2025)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
di: Lalwani, Abhinav, et al.
Pubblicazione: (2024)
di: Lalwani, Abhinav, et al.
Pubblicazione: (2024)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
Are Language Models Efficient Reasoners? A Perspective from Logic Programming
di: Opedal, Andreas, et al.
Pubblicazione: (2025)
di: Opedal, Andreas, et al.
Pubblicazione: (2025)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
di: Cui, Shaobo, et al.
Pubblicazione: (2024)
di: Cui, Shaobo, et al.
Pubblicazione: (2024)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
di: Chowdhury, Sankalan Pal, et al.
Pubblicazione: (2024)
di: Chowdhury, Sankalan Pal, et al.
Pubblicazione: (2024)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
di: Opedal, Andreas, et al.
Pubblicazione: (2024)
di: Opedal, Andreas, et al.
Pubblicazione: (2024)
Efficiently Computing Susceptibility to Context in Language Models
di: Liu, Tianyu, et al.
Pubblicazione: (2024)
di: Liu, Tianyu, et al.
Pubblicazione: (2024)
Do Vision-Language Models Really Understand Visual Language?
di: Hou, Yifan, et al.
Pubblicazione: (2024)
di: Hou, Yifan, et al.
Pubblicazione: (2024)
Are Language Models Consequentialist or Deontological Moral Reasoners?
di: Samway, Keenan, et al.
Pubblicazione: (2025)
di: Samway, Keenan, et al.
Pubblicazione: (2025)
What Do Language Models Learn in Context? The Structured Task Hypothesis
di: Li, Jiaoda, et al.
Pubblicazione: (2024)
di: Li, Jiaoda, et al.
Pubblicazione: (2024)
When Do Language Models Endorse Limitations on Human Rights Principles?
di: Samway, Keenan, et al.
Pubblicazione: (2026)
di: Samway, Keenan, et al.
Pubblicazione: (2026)
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
di: Serra, Alessandro Pietro, et al.
Pubblicazione: (2024)
di: Serra, Alessandro Pietro, et al.
Pubblicazione: (2024)
Grammar Control in Dialogue Response Generation for Language Learning Chatbots
di: Glandorf, Dominik, et al.
Pubblicazione: (2025)
di: Glandorf, Dominik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
di: Ortu, Francesco, et al.
Pubblicazione: (2025) -
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025) -
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
di: Dotsinski, Asen, et al.
Pubblicazione: (2025) -
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
di: Jenny, David F., et al.
Pubblicazione: (2023) -
Are LLMs Good Safety Agents or a Propaganda Engine?
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)