On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
Fuente:
arXiv
Salvato in:
| Autori principali: | Dotsinski, Asen, Thakur, Udit, Ivanov, Marko, Khan, Mohammad Hafeez, Heuss, Maria |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
di: Campregher, Dante, et al.
Pubblicazione: (2025)
di: Campregher, Dante, et al.
Pubblicazione: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
di: Dotsinski, Asen, et al.
Pubblicazione: (2026)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
di: Ortu, Francesco, et al.
Pubblicazione: (2024)
di: Ortu, Francesco, et al.
Pubblicazione: (2024)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
di: Christoph, Daniel, et al.
Pubblicazione: (2025)
di: Christoph, Daniel, et al.
Pubblicazione: (2025)
From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
di: Khan, Imran
Pubblicazione: (2025)
di: Khan, Imran
Pubblicazione: (2025)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2026)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
di: Kim, Hazel H.
Pubblicazione: (2024)
di: Kim, Hazel H.
Pubblicazione: (2024)
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
di: Kutakh, Matthew
Pubblicazione: (2026)
di: Kutakh, Matthew
Pubblicazione: (2026)
On Large Language Models' Hallucination with Regard to Known Facts
di: Jiang, Che, et al.
Pubblicazione: (2024)
di: Jiang, Che, et al.
Pubblicazione: (2024)
Unsupervised Pretraining for Fact Verification by Language Model Distillation
di: Bazaga, Adrián, et al.
Pubblicazione: (2023)
di: Bazaga, Adrián, et al.
Pubblicazione: (2023)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
di: Behnia, Tina, et al.
Pubblicazione: (2025)
di: Behnia, Tina, et al.
Pubblicazione: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
di: Yoon, Youngsik, et al.
Pubblicazione: (2026)
di: Yoon, Youngsik, et al.
Pubblicazione: (2026)
Gumbel Counterfactual Generation From Language Models
di: Ravfogel, Shauli, et al.
Pubblicazione: (2024)
di: Ravfogel, Shauli, et al.
Pubblicazione: (2024)
Counterfactual Token Generation in Large Language Models
di: Chatzi, Ivi, et al.
Pubblicazione: (2024)
di: Chatzi, Ivi, et al.
Pubblicazione: (2024)
Stacking Small Language Models for Generalizability
di: Liang, Laurence
Pubblicazione: (2024)
di: Liang, Laurence
Pubblicazione: (2024)
TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
di: Chang, Shenxu, et al.
Pubblicazione: (2025)
di: Chang, Shenxu, et al.
Pubblicazione: (2025)
Are Large Language Models Table-based Fact-Checkers?
di: Zhang, Hanwen, et al.
Pubblicazione: (2024)
di: Zhang, Hanwen, et al.
Pubblicazione: (2024)
Fact or Guesswork? Evaluating Large Language Models' Medical Knowledge with Structured One-Hop Judgments
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
Reasoning Elicitation in Language Models via Counterfactual Feedback
di: Hüyük, Alihan, et al.
Pubblicazione: (2024)
di: Hüyük, Alihan, et al.
Pubblicazione: (2024)
CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval
di: Putta, Akshith Reddy, et al.
Pubblicazione: (2026)
di: Putta, Akshith Reddy, et al.
Pubblicazione: (2026)
Fine-Tuned Large Language Models for Symptom Recognition from Spanish Clinical Text
di: Shaaban, Mai A., et al.
Pubblicazione: (2024)
di: Shaaban, Mai A., et al.
Pubblicazione: (2024)
BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques
di: Kabir, Muhammad Rafsan, et al.
Pubblicazione: (2024)
di: Kabir, Muhammad Rafsan, et al.
Pubblicazione: (2024)
FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models
di: Wang, Yanling, et al.
Pubblicazione: (2024)
di: Wang, Yanling, et al.
Pubblicazione: (2024)
Large Language Models as Generalizable Policies for Embodied Tasks
di: Szot, Andrew, et al.
Pubblicazione: (2023)
di: Szot, Andrew, et al.
Pubblicazione: (2023)
Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2025)
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2025)
Building a Strong Instruction Language Model for a Less-Resourced Language
di: Vreš, Domen, et al.
Pubblicazione: (2026)
di: Vreš, Domen, et al.
Pubblicazione: (2026)
Tracing Uncertainty in Language Model "Reasoning"
di: Grünefeld, Nils, et al.
Pubblicazione: (2026)
di: Grünefeld, Nils, et al.
Pubblicazione: (2026)
Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning
di: Liu, Shuo, et al.
Pubblicazione: (2026)
di: Liu, Shuo, et al.
Pubblicazione: (2026)
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective
di: Ghosh, Bishwamittra, et al.
Pubblicazione: (2026)
di: Ghosh, Bishwamittra, et al.
Pubblicazione: (2026)
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
di: Deng, Xinle, et al.
Pubblicazione: (2026)
di: Deng, Xinle, et al.
Pubblicazione: (2026)
Counterfactual Generation with Identifiability Guarantees
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
Fact Checking Beyond Training Set
di: Karisani, Payam, et al.
Pubblicazione: (2024)
di: Karisani, Payam, et al.
Pubblicazione: (2024)
How Reliable is Language Model Micro-Benchmarking?
di: Yauney, Gregory, et al.
Pubblicazione: (2025)
di: Yauney, Gregory, et al.
Pubblicazione: (2025)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
di: Somayajula, Sai Ashish, et al.
Pubblicazione: (2024)
di: Somayajula, Sai Ashish, et al.
Pubblicazione: (2024)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2024)
di: Fadeeva, Ekaterina, et al.
Pubblicazione: (2024)
Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective for Molecular Property Prediction
di: He, Yinhan, et al.
Pubblicazione: (2024)
di: He, Yinhan, et al.
Pubblicazione: (2024)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
di: Brahman, Faeze, et al.
Pubblicazione: (2023)
di: Brahman, Faeze, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
di: Campregher, Dante, et al.
Pubblicazione: (2025) -
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
di: Dotsinski, Asen, et al.
Pubblicazione: (2026) -
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
di: Ortu, Francesco, et al.
Pubblicazione: (2024) -
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
di: Christoph, Daniel, et al.
Pubblicazione: (2025) -
From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
di: Khan, Imran
Pubblicazione: (2025)