Tracing and Reversing Edits in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Youssef, Paul, Zhao, Zhixue, Seifert, Christin, Schlötterer, Jörg |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025)
by: Cheng, Yinjie, et al.
Published: (2025)
Persuasion Tokens for Editing Factual Knowledge in LLMs
by: Youssef, Paul, et al.
Published: (2026)
by: Youssef, Paul, et al.
Published: (2026)
Position: Editing Large Language Models Poses Serious Safety Risks
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Enhancing Fact Retrieval in PLMs through Truthfulness
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
by: Nguyen, Van Bach, et al.
Published: (2025)
by: Nguyen, Van Bach, et al.
Published: (2025)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
by: Nguyen, Van Bach, et al.
Published: (2022)
by: Nguyen, Van Bach, et al.
Published: (2022)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
by: Koraş, Osman Alperen, et al.
Published: (2024)
by: Koraş, Osman Alperen, et al.
Published: (2024)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
Behavioral Analysis of Information Salience in Large Language Models
by: Trienes, Jan, et al.
Published: (2025)
by: Trienes, Jan, et al.
Published: (2025)
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
by: Trienes, Jan, et al.
Published: (2025)
by: Trienes, Jan, et al.
Published: (2025)
An XAI-based Analysis of Shortcut Learning in Neural Networks
by: Le, Phuong Quynh, et al.
Published: (2025)
by: Le, Phuong Quynh, et al.
Published: (2025)
Is Last Layer Re-Training Truly Sufficient for Robustness to Spurious Correlations?
by: Le, Phuong Quynh, et al.
Published: (2023)
by: Le, Phuong Quynh, et al.
Published: (2023)
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
by: Trienes, Jan, et al.
Published: (2024)
by: Trienes, Jan, et al.
Published: (2024)
Towards Interpretable Deep Neural Networks for Tabular Data
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
Investigating the Impact of Randomness on Reproducibility in Computer Vision: A Study on Applications in Civil Engineering and Medicine
by: Eryılmaz, Bahadır, et al.
Published: (2024)
by: Eryılmaz, Bahadır, et al.
Published: (2024)
Explanation Generation for Contradiction Reconciliation with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
by: Chan, Jason, et al.
Published: (2025)
by: Chan, Jason, et al.
Published: (2025)
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
by: Chan, Jason, et al.
Published: (2024)
by: Chan, Jason, et al.
Published: (2024)
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
Invariant Learning with Annotation-free Environments
by: Le, Phuong Quynh, et al.
Published: (2025)
by: Le, Phuong Quynh, et al.
Published: (2025)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
by: Le, Phuong Quynh, et al.
Published: (2024)
by: Le, Phuong Quynh, et al.
Published: (2024)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
Explanation format does not matter; but explanations do -- An Eggsbert study on explaining Bayesian Optimisation tasks
by: Chakraborty, Tanmay, et al.
Published: (2025)
by: Chakraborty, Tanmay, et al.
Published: (2025)
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
by: Sarumi, Olufunke O., et al.
Published: (2025)
by: Sarumi, Olufunke O., et al.
Published: (2025)
The Impact of Annotator Personas on LLM Behavior Across the Perspectivism Spectrum
by: Sarumi, Olufunke O., et al.
Published: (2025)
by: Sarumi, Olufunke O., et al.
Published: (2025)
Patch-based Intuitive Multimodal Prototypes Network (PIMPNet) for Alzheimer's Disease classification
by: De Santi, Lisa Anita, et al.
Published: (2024)
by: De Santi, Lisa Anita, et al.
Published: (2024)
From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks
by: Datta, Ayan, et al.
Published: (2026)
by: Datta, Ayan, et al.
Published: (2026)
Trace and Edit Relation Associations in GPT
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
Efficient Unsupervised Shortcut Learning Detection and Mitigation in Transformers
by: Kuhn, Lukas, et al.
Published: (2025)
by: Kuhn, Lukas, et al.
Published: (2025)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Prototype-based Interpretable Breast Cancer Prediction Models: Analysis and Challenges
by: Pathak, Shreyasi, et al.
Published: (2024)
by: Pathak, Shreyasi, et al.
Published: (2024)
Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
by: Nie, Shangrui, et al.
Published: (2025)
by: Nie, Shangrui, et al.
Published: (2025)
Label Set Optimization via Activation Distribution Kurtosis for Zero-shot Classification with Generative Models
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Learning to Edit: Aligning LLMs with Knowledge Editing
by: Jiang, Yuxin, et al.
Published: (2024)
by: Jiang, Yuxin, et al.
Published: (2024)
Similar Items
-
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
by: Youssef, Paul, et al.
Published: (2024) -
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024) -
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025) -
Persuasion Tokens for Editing Factual Knowledge in LLMs
by: Youssef, Paul, et al.
Published: (2026) -
Position: Editing Large Language Models Poses Serious Safety Risks
by: Youssef, Paul, et al.
Published: (2025)