Persuasion Tokens for Editing Factual Knowledge in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Youssef, Paul, Seifert, Christin, Schlötterer, Jörg |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
por: Koraş, Osman Alperen, et al.
Publicado: (2024)
por: Koraş, Osman Alperen, et al.
Publicado: (2024)
Tracing and Reversing Edits in LLMs
por: Youssef, Paul, et al.
Publicado: (2025)
por: Youssef, Paul, et al.
Publicado: (2025)
Enhancing Fact Retrieval in PLMs through Truthfulness
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
por: Cheng, Yinjie, et al.
Publicado: (2025)
por: Cheng, Yinjie, et al.
Publicado: (2025)
Position: Editing Large Language Models Poses Serious Safety Risks
por: Youssef, Paul, et al.
Publicado: (2025)
por: Youssef, Paul, et al.
Publicado: (2025)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
por: Nguyen, Van Bach, et al.
Publicado: (2024)
por: Nguyen, Van Bach, et al.
Publicado: (2024)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
por: Nguyen, Van Bach, et al.
Publicado: (2025)
por: Nguyen, Van Bach, et al.
Publicado: (2025)
An XAI-based Analysis of Shortcut Learning in Neural Networks
por: Le, Phuong Quynh, et al.
Publicado: (2025)
por: Le, Phuong Quynh, et al.
Publicado: (2025)
Is Last Layer Re-Training Truly Sufficient for Robustness to Spurious Correlations?
por: Le, Phuong Quynh, et al.
Publicado: (2023)
por: Le, Phuong Quynh, et al.
Publicado: (2023)
Towards Interpretable Deep Neural Networks for Tabular Data
por: Elhadri, Khawla, et al.
Publicado: (2025)
por: Elhadri, Khawla, et al.
Publicado: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
por: Elhadri, Khawla, et al.
Publicado: (2025)
por: Elhadri, Khawla, et al.
Publicado: (2025)
Invariant Learning with Annotation-free Environments
por: Le, Phuong Quynh, et al.
Publicado: (2025)
por: Le, Phuong Quynh, et al.
Publicado: (2025)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
por: Nguyen, Van Bach, et al.
Publicado: (2022)
por: Nguyen, Van Bach, et al.
Publicado: (2022)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
por: Nguyen, Van Bach, et al.
Publicado: (2024)
por: Nguyen, Van Bach, et al.
Publicado: (2024)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
por: Le, Phuong Quynh, et al.
Publicado: (2024)
por: Le, Phuong Quynh, et al.
Publicado: (2024)
Behavioral Analysis of Information Salience in Large Language Models
por: Trienes, Jan, et al.
Publicado: (2025)
por: Trienes, Jan, et al.
Publicado: (2025)
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
por: Zhang, Zhuoran, et al.
Publicado: (2024)
por: Zhang, Zhuoran, et al.
Publicado: (2024)
Mitigating Heterogeneous Token Overfitting in LLM Knowledge Editing
por: Liu, Tianci, et al.
Publicado: (2025)
por: Liu, Tianci, et al.
Publicado: (2025)
One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them
por: Holmov, Ali, et al.
Publicado: (2026)
por: Holmov, Ali, et al.
Publicado: (2026)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
por: Yuan, Jiaqing, et al.
Publicado: (2024)
por: Yuan, Jiaqing, et al.
Publicado: (2024)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
por: Wu, Qinyuan, et al.
Publicado: (2024)
por: Wu, Qinyuan, et al.
Publicado: (2024)
Understanding Finetuning for Factual Knowledge Extraction
por: Ghosal, Gaurav, et al.
Publicado: (2024)
por: Ghosal, Gaurav, et al.
Publicado: (2024)
Efficient Unsupervised Shortcut Learning Detection and Mitigation in Transformers
por: Kuhn, Lukas, et al.
Publicado: (2025)
por: Kuhn, Lukas, et al.
Publicado: (2025)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
por: Mahaut, Matéo, et al.
Publicado: (2024)
por: Mahaut, Matéo, et al.
Publicado: (2024)
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
por: Thede, Lukas, et al.
Publicado: (2025)
por: Thede, Lukas, et al.
Publicado: (2025)
Shortcut Mitigation via Spurious-Positive Samples
por: Le, Phuong Quynh, et al.
Publicado: (2026)
por: Le, Phuong Quynh, et al.
Publicado: (2026)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
por: Chughtai, Bilal, et al.
Publicado: (2024)
por: Chughtai, Bilal, et al.
Publicado: (2024)
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
por: Trienes, Jan, et al.
Publicado: (2025)
por: Trienes, Jan, et al.
Publicado: (2025)
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models
por: Mekala, Anmol, et al.
Publicado: (2024)
por: Mekala, Anmol, et al.
Publicado: (2024)
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
por: Liu, Jinzhe, et al.
Publicado: (2025)
por: Liu, Jinzhe, et al.
Publicado: (2025)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
This looks like what? Challenges and Future Research Directions for Part-Prototype Models
por: Elhadri, Khawla, et al.
Publicado: (2025)
por: Elhadri, Khawla, et al.
Publicado: (2025)
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
por: Khodja, Hichem Ammar, et al.
Publicado: (2025)
por: Khodja, Hichem Ammar, et al.
Publicado: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
por: Sreenivas, Sharath Turuvekere, et al.
Publicado: (2026)
por: Sreenivas, Sharath Turuvekere, et al.
Publicado: (2026)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
por: Rohera, Pritika, et al.
Publicado: (2025)
por: Rohera, Pritika, et al.
Publicado: (2025)
Toward a Theory of Tokenization in LLMs
por: Rajaraman, Nived, et al.
Publicado: (2024)
por: Rajaraman, Nived, et al.
Publicado: (2024)
Silent Tokens, Loud Effects: Padding in LLMs
por: Himelstein, Rom, et al.
Publicado: (2025)
por: Himelstein, Rom, et al.
Publicado: (2025)
Ejemplares similares
-
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
por: Youssef, Paul, et al.
Publicado: (2024) -
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
por: Youssef, Paul, et al.
Publicado: (2024) -
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
por: Youssef, Paul, et al.
Publicado: (2024) -
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
por: Koraş, Osman Alperen, et al.
Publicado: (2024) -
Tracing and Reversing Edits in LLMs
por: Youssef, Paul, et al.
Publicado: (2025)