Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Dentan, Jérémie, Buscaldi, Davide, Vanier, Sonia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
di: Bail, Mathis Le, et al.
Pubblicazione: (2025)
di: Bail, Mathis Le, et al.
Pubblicazione: (2025)
MUCH: A Multilingual Claim Hallucination Benchmark
di: Dentan, Jérémie, et al.
Pubblicazione: (2025)
di: Dentan, Jérémie, et al.
Pubblicazione: (2025)
Predicting memorization within Large Language Models fine-tuned for classification
di: Dentan, Jérémie, et al.
Pubblicazione: (2024)
di: Dentan, Jérémie, et al.
Pubblicazione: (2024)
Activation Surgery: Jailbreaking White-box LLMs without Touching the Prompt
di: Jenny, Maël, et al.
Pubblicazione: (2026)
di: Jenny, Maël, et al.
Pubblicazione: (2026)
PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
di: Dhouib, Mohamed, et al.
Pubblicazione: (2025)
di: Dhouib, Mohamed, et al.
Pubblicazione: (2025)
Leveraging Contrastive Learning for a Similarity-Guided Tampered Document Data Generation Pipeline
di: Dhouib, Mohamed, et al.
Pubblicazione: (2026)
di: Dhouib, Mohamed, et al.
Pubblicazione: (2026)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
di: Zhang, Ying, et al.
Pubblicazione: (2025)
di: Zhang, Ying, et al.
Pubblicazione: (2025)
Word Sense Induction with Hierarchical Clustering and Mutual Information Maximization
di: Abdine, Hadi, et al.
Pubblicazione: (2022)
di: Abdine, Hadi, et al.
Pubblicazione: (2022)
What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models
di: Zhao, Xin, et al.
Pubblicazione: (2024)
di: Zhao, Xin, et al.
Pubblicazione: (2024)
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
di: Srivastava, Alisha, et al.
Pubblicazione: (2025)
di: Srivastava, Alisha, et al.
Pubblicazione: (2025)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
di: Chang, Ting-Yun, et al.
Pubblicazione: (2023)
di: Chang, Ting-Yun, et al.
Pubblicazione: (2023)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
di: Hans, Abhimanyu, et al.
Pubblicazione: (2024)
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
di: Yu, Qingchen, et al.
Pubblicazione: (2025)
di: Yu, Qingchen, et al.
Pubblicazione: (2025)
Sensivity of LLMs' Explanations to the Training Randomness:Context, Class & Task Dependencies
di: Loncour, Romain, et al.
Pubblicazione: (2026)
di: Loncour, Romain, et al.
Pubblicazione: (2026)
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
di: Kassem, Aly M., et al.
Pubblicazione: (2024)
di: Kassem, Aly M., et al.
Pubblicazione: (2024)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
di: Kim, Jisu, et al.
Pubblicazione: (2025)
di: Kim, Jisu, et al.
Pubblicazione: (2025)
Mitigating Memorization in LLMs using Activation Steering
di: Suri, Manan, et al.
Pubblicazione: (2025)
di: Suri, Manan, et al.
Pubblicazione: (2025)
Arithmetic with Language Models: from Memorization to Computation
di: Maltoni, Davide, et al.
Pubblicazione: (2023)
di: Maltoni, Davide, et al.
Pubblicazione: (2023)
Memorization and Knowledge Injection in Gated LLMs
di: Pan, Xu, et al.
Pubblicazione: (2025)
di: Pan, Xu, et al.
Pubblicazione: (2025)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
di: Ruan, Zhiwen, et al.
Pubblicazione: (2025)
Understanding Verbatim Memorization in LLMs Through Circuit Discovery
di: Lasy, Ilya, et al.
Pubblicazione: (2025)
di: Lasy, Ilya, et al.
Pubblicazione: (2025)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2026)
di: Gekhman, Zorik, et al.
Pubblicazione: (2026)
Exploring Precision and Recall to assess the quality and diversity of LLMs
di: Bronnec, Florian Le, et al.
Pubblicazione: (2024)
di: Bronnec, Florian Le, et al.
Pubblicazione: (2024)
SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
di: Gashkov, Aleksandr, et al.
Pubblicazione: (2025)
di: Gashkov, Aleksandr, et al.
Pubblicazione: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
di: Lei, Ge, et al.
Pubblicazione: (2025)
di: Lei, Ge, et al.
Pubblicazione: (2025)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
di: Kale, Sahil
Pubblicazione: (2025)
di: Kale, Sahil
Pubblicazione: (2025)
Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
di: Wu, Qinyuan, et al.
Pubblicazione: (2025)
di: Wu, Qinyuan, et al.
Pubblicazione: (2025)
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
di: Jiang, Yuxuan, et al.
Pubblicazione: (2024)
di: Jiang, Yuxuan, et al.
Pubblicazione: (2024)
From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals
di: Yu, Yeyong, et al.
Pubblicazione: (2026)
di: Yu, Yeyong, et al.
Pubblicazione: (2026)
Localizing Paragraph Memorization in Language Models
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
di: Stoehr, Niklas, et al.
Pubblicazione: (2024)
GuessBench: Sensemaking Multimodal Creativity in the Wild
di: Zhu, Zifeng, et al.
Pubblicazione: (2025)
di: Zhu, Zifeng, et al.
Pubblicazione: (2025)
Memory Dial: A Training Framework for Controllable Memorization in Language Models
di: Zhang, Xiangbo, et al.
Pubblicazione: (2026)
di: Zhang, Xiangbo, et al.
Pubblicazione: (2026)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
di: Xiong, Alexander, et al.
Pubblicazione: (2025)
Identifying Legal Holdings with LLMs: A Systematic Study of Performance, Scale, and Memorization
di: Arvin, Chuck
Pubblicazione: (2025)
di: Arvin, Chuck
Pubblicazione: (2025)
Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
di: Ruzzetti, Elena Sofia, et al.
Pubblicazione: (2025)
di: Ruzzetti, Elena Sofia, et al.
Pubblicazione: (2025)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
di: Luo, Xiaoyu, et al.
Pubblicazione: (2026)
di: Luo, Xiaoyu, et al.
Pubblicazione: (2026)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
di: Tang, Ethan
Pubblicazione: (2026)
di: Tang, Ethan
Pubblicazione: (2026)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
di: Xu, Yixuan, et al.
Pubblicazione: (2025)
di: Xu, Yixuan, et al.
Pubblicazione: (2025)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
di: Uddin, Md Nayem, et al.
Pubblicazione: (2024)
di: Uddin, Md Nayem, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
di: Bail, Mathis Le, et al.
Pubblicazione: (2025) -
MUCH: A Multilingual Claim Hallucination Benchmark
di: Dentan, Jérémie, et al.
Pubblicazione: (2025) -
Predicting memorization within Large Language Models fine-tuned for classification
di: Dentan, Jérémie, et al.
Pubblicazione: (2024) -
Activation Surgery: Jailbreaking White-box LLMs without Touching the Prompt
di: Jenny, Maël, et al.
Pubblicazione: (2026) -
PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
di: Dhouib, Mohamed, et al.
Pubblicazione: (2025)