Large-Language Memorization During the Classification of United States Supreme Court Cases

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ortega, John E., Joshi, Dhruv D., Borkowski, Matt P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909963359617024
author Ortega, John E.
Joshi, Dhruv D.
Borkowski, Matt P.
author_facet Ortega, John E.
Joshi, Dhruv D.
Borkowski, Matt P.
contents Large-language models (LLMs) have been shown to respond in a variety of ways for classification tasks outside of question-answering. LLM responses are sometimes called "hallucinations" since the output is not what is ex pected. Memorization strategies in LLMs are being studied in detail, with the goal of understanding how LLMs respond. We perform a deep dive into a classification task based on United States Supreme Court (SCOTUS) decisions. The SCOTUS corpus is an ideal classification task to study for LLM memory accuracy because it presents significant challenges due to extensive sentence length, complex legal terminology, non-standard structure, and domain-specific vocabulary. Experimentation is performed with the latest LLM fine tuning and retrieval-based approaches, such as parameter-efficient fine-tuning, auto-modeling, and others, on two traditional category-based SCOTUS classification tasks: one with 15 labeled topics and another with 279. We show that prompt-based models with memories, such as DeepSeek, can be more robust than previous BERT-based models on both tasks scoring about 2 points better than previous models not based on prompting.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large-Language Memorization During the Classification of United States Supreme Court Cases
Ortega, John E.
Joshi, Dhruv D.
Borkowski, Matt P.
Computation and Language
Artificial Intelligence
Emerging Technologies
Information Retrieval
Large-language models (LLMs) have been shown to respond in a variety of ways for classification tasks outside of question-answering. LLM responses are sometimes called "hallucinations" since the output is not what is ex pected. Memorization strategies in LLMs are being studied in detail, with the goal of understanding how LLMs respond. We perform a deep dive into a classification task based on United States Supreme Court (SCOTUS) decisions. The SCOTUS corpus is an ideal classification task to study for LLM memory accuracy because it presents significant challenges due to extensive sentence length, complex legal terminology, non-standard structure, and domain-specific vocabulary. Experimentation is performed with the latest LLM fine tuning and retrieval-based approaches, such as parameter-efficient fine-tuning, auto-modeling, and others, on two traditional category-based SCOTUS classification tasks: one with 15 labeled topics and another with 279. We show that prompt-based models with memories, such as DeepSeek, can be more robust than previous BERT-based models on both tasks scoring about 2 points better than previous models not based on prompting.
title Large-Language Memorization During the Classification of United States Supreme Court Cases
topic Computation and Language
Artificial Intelligence
Emerging Technologies
Information Retrieval
url https://arxiv.org/abs/2512.13654