Understanding Knowledge Drift in LLMs through Misinformation
Fuente:
arXiv
Salvato in:
| Autori principali: | Fastowski, Alina, Kasneci, Gjergji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
From Confidence to Collapse in LLM Factual Robustness
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)
Analysing the Safety Pitfalls of Steering Vectors
di: Li, Yuxiao, et al.
Pubblicazione: (2026)
di: Li, Yuxiao, et al.
Pubblicazione: (2026)
Consolidating Rewarded Perturbations for LLM Post-Training
di: Zhang, Zheyu, et al.
Pubblicazione: (2026)
di: Zhang, Zheyu, et al.
Pubblicazione: (2026)
Emergent Abilities in Large Language Models: A Survey
di: Berti, Leonardo, et al.
Pubblicazione: (2025)
di: Berti, Leonardo, et al.
Pubblicazione: (2025)
Not All Features Deserve Attention: Graph-Guided Dependency Learning for Tabular Data Generation with Language Models
di: Zhang, Zheyu, et al.
Pubblicazione: (2025)
di: Zhang, Zheyu, et al.
Pubblicazione: (2025)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
di: Kocak, Aysenur, et al.
Pubblicazione: (2025)
di: Kocak, Aysenur, et al.
Pubblicazione: (2025)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
Entry Dependent Expert Selection in Distributed Gaussian Processes Using Multilabel Classification
di: Jalali, Hamed, et al.
Pubblicazione: (2022)
di: Jalali, Hamed, et al.
Pubblicazione: (2022)
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers
di: Kasneci, Gjergji, et al.
Pubblicazione: (2024)
di: Kasneci, Gjergji, et al.
Pubblicazione: (2024)
Enhancing Fairness through Reweighting: A Path to Attain the Sufficiency Rule
di: Zhao, Xuan, et al.
Pubblicazione: (2024)
di: Zhao, Xuan, et al.
Pubblicazione: (2024)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
di: Lu, Taiming, et al.
Pubblicazione: (2024)
di: Lu, Taiming, et al.
Pubblicazione: (2024)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
di: Abramov, Roman, et al.
Pubblicazione: (2025)
di: Abramov, Roman, et al.
Pubblicazione: (2025)
Mitigating Semantic Drift: Evaluating LLMs' Efficacy in Psychotherapy through MI Dialogue Summarization
di: Kumar, Vivek, et al.
Pubblicazione: (2025)
di: Kumar, Vivek, et al.
Pubblicazione: (2025)
SCISSOR: Mitigating Semantic Bias through Cluster-Aware Siamese Networks for Robust Classification
di: Yang, Shuo, et al.
Pubblicazione: (2025)
di: Yang, Shuo, et al.
Pubblicazione: (2025)
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
di: Xiao, Yuzhuo, et al.
Pubblicazione: (2025)
di: Xiao, Yuzhuo, et al.
Pubblicazione: (2025)
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
di: Thede, Lukas, et al.
Pubblicazione: (2025)
di: Thede, Lukas, et al.
Pubblicazione: (2025)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
Understanding Chain-of-Thought in LLMs through Information Theory
di: Ton, Jean-Francois, et al.
Pubblicazione: (2024)
di: Ton, Jean-Francois, et al.
Pubblicazione: (2024)
Benchmarking Large Language Models for Math Reasoning Tasks
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
Reinforcement Unlearning via Group Relative Policy Optimization
di: Zaradoukas, Efstratios, et al.
Pubblicazione: (2026)
di: Zaradoukas, Efstratios, et al.
Pubblicazione: (2026)
Graph Inverse Style Transfer for Counterfactual Explainability
di: Prenkaj, Bardh, et al.
Pubblicazione: (2025)
di: Prenkaj, Bardh, et al.
Pubblicazione: (2025)
DART-ing Through the Drift: Dynamic Tracing of Knowledge Neurons for Adaptive Inference-Time Pruning
di: Tyagi, Abhishek, et al.
Pubblicazione: (2026)
di: Tyagi, Abhishek, et al.
Pubblicazione: (2026)
TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
di: Berti, Leonardo, et al.
Pubblicazione: (2025)
di: Berti, Leonardo, et al.
Pubblicazione: (2025)
Understanding Finetuning for Factual Knowledge Extraction
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
SoK: Machine Learning for Misinformation Detection
di: Xiao, Madelyne, et al.
Pubblicazione: (2023)
di: Xiao, Madelyne, et al.
Pubblicazione: (2023)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
di: Qi, Yanlin, et al.
Pubblicazione: (2026)
di: Qi, Yanlin, et al.
Pubblicazione: (2026)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
di: Ni, Ruikang, et al.
Pubblicazione: (2024)
di: Ni, Ruikang, et al.
Pubblicazione: (2024)
I Prefer not to Say: Protecting User Consent in Models with Optional Personal Data
di: Leemann, Tobias, et al.
Pubblicazione: (2022)
di: Leemann, Tobias, et al.
Pubblicazione: (2022)
Persuasion Tokens for Editing Factual Knowledge in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2026)
di: Youssef, Paul, et al.
Pubblicazione: (2026)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
di: Nyandwi, Jean de Dieu, et al.
Pubblicazione: (2025)
di: Nyandwi, Jean de Dieu, et al.
Pubblicazione: (2025)
Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
Understanding Memorisation in LLMs: Dynamics, Influencing Factors, and Implications
di: Speicher, Till, et al.
Pubblicazione: (2024)
di: Speicher, Till, et al.
Pubblicazione: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
Efficient Knowledge Injection in LLMs via Self-Distillation
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
di: Leemann, Tobias, et al.
Pubblicazione: (2024) -
From Confidence to Collapse in LLM Factual Robustness
di: Fastowski, Alina, et al.
Pubblicazione: (2025) -
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
di: Fastowski, Alina, et al.
Pubblicazione: (2025) -
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
di: Yang, Shuo, et al.
Pubblicazione: (2024) -
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)