Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Qinyuan, Das, Soumi, Amani, Mahsa, Ghosh, Bishwamittra, Khan, Mohammad Aflah, Gummadi, Krishna P., Zafar, Muhammad Bilal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914360148885504
author Wu, Qinyuan
Das, Soumi
Amani, Mahsa
Ghosh, Bishwamittra
Khan, Mohammad Aflah
Gummadi, Krishna P.
Zafar, Muhammad Bilal
author_facet Wu, Qinyuan
Das, Soumi
Amani, Mahsa
Ghosh, Bishwamittra
Khan, Mohammad Aflah
Gummadi, Krishna P.
Zafar, Muhammad Bilal
contents Rote learning is a memorization technique based on repetition. Many researchers argue that rote learning hinders generalization because it encourages verbatim memorization rather than deeper understanding. This concern extends even to factual knowledge, which inevitably requires a certain degree of memorization. In this work, we challenge this view and demonstrate that large language models (LLMs) can, in fact, generalize over rote memorized data. We introduce a two-phase "memorize-then-generalize" framework, where the model first rote memorizes factual subject-object associations using a synthetic semantically meaningless key token and then learns to generalize by fine-tuning on a small set of semantically meaningful prompts. Extensive experiments over 8 LLMs show that the models can reinterpret rote memorized data through the semantically meaningful prompts, as evidenced by the emergence of structured, semantically aligned latent representations between the key token and the semantically meaningful prompts. This surprising finding opens the door to both effective and efficient knowledge injection as well as possible risks of repurposing the memorized data for malicious usage.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
Wu, Qinyuan
Das, Soumi
Amani, Mahsa
Ghosh, Bishwamittra
Khan, Mohammad Aflah
Gummadi, Krishna P.
Zafar, Muhammad Bilal
Computation and Language
Rote learning is a memorization technique based on repetition. Many researchers argue that rote learning hinders generalization because it encourages verbatim memorization rather than deeper understanding. This concern extends even to factual knowledge, which inevitably requires a certain degree of memorization. In this work, we challenge this view and demonstrate that large language models (LLMs) can, in fact, generalize over rote memorized data. We introduce a two-phase "memorize-then-generalize" framework, where the model first rote memorizes factual subject-object associations using a synthetic semantically meaningless key token and then learns to generalize by fine-tuning on a small set of semantically meaningful prompts. Extensive experiments over 8 LLMs show that the models can reinterpret rote memorized data through the semantically meaningful prompts, as evidenced by the emergence of structured, semantically aligned latent representations between the key token and the semantically meaningful prompts. This surprising finding opens the door to both effective and efficient knowledge injection as well as possible risks of repurposing the memorized data for malicious usage.
title Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
topic Computation and Language
url https://arxiv.org/abs/2507.21914