Classifying long legal documents using short random chunks
Fuente:
arXiv
Saved in:
| Main Author: | Cabrera-Diego, Luis Adrián |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
by: Vargas, Francisco, et al.
Published: (2025)
by: Vargas, Francisco, et al.
Published: (2025)
Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap
by: Panda, Akash Kumar, et al.
Published: (2026)
by: Panda, Akash Kumar, et al.
Published: (2026)
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
by: Yuan, Mingruo, et al.
Published: (2025)
by: Yuan, Mingruo, et al.
Published: (2025)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025)
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025)
Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
by: Lenz, Stefan, et al.
Published: (2025)
by: Lenz, Stefan, et al.
Published: (2025)
Coverage-based Fairness in Multi-document Summarization
by: Li, Haoyuan, et al.
Published: (2024)
by: Li, Haoyuan, et al.
Published: (2024)
Making Language Model a Hierarchical Classifier
by: Wang, Yihong, et al.
Published: (2025)
by: Wang, Yihong, et al.
Published: (2025)
Building Efficient Universal Classifiers with Natural Language Inference
by: Laurer, Moritz, et al.
Published: (2023)
by: Laurer, Moritz, et al.
Published: (2023)
LBC: Language-Based-Classifier for Out-Of-Variable Generalization
by: Noh, Kangjun, et al.
Published: (2024)
by: Noh, Kangjun, et al.
Published: (2024)
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
Chip-Tuning: Classify Before Language Models Say
by: Zhu, Fangwei, et al.
Published: (2024)
by: Zhu, Fangwei, et al.
Published: (2024)
Rethinking Transformer-based Multi-document Summarization: An Empirical Investigation
by: Ma, Congbo, et al.
Published: (2024)
by: Ma, Congbo, et al.
Published: (2024)
Linear Cross-document Event Coreference Resolution with X-AMR
by: Ahmed, Shafiuddin Rehan, et al.
Published: (2024)
by: Ahmed, Shafiuddin Rehan, et al.
Published: (2024)
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
by: McAleese, Stephen, et al.
Published: (2024)
by: McAleese, Stephen, et al.
Published: (2024)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025)
by: Bonomo, Tommaso, et al.
Published: (2025)
Classifying German Language Proficiency Levels Using Large Language Models
by: Ahlers, Elias-Leander, et al.
Published: (2025)
by: Ahlers, Elias-Leander, et al.
Published: (2025)
Classifying Cancer Stage with Open-Source Clinical Large Language Models
by: Chang, Chia-Hsuan, et al.
Published: (2024)
by: Chang, Chia-Hsuan, et al.
Published: (2024)
Classifying Human-Generated and AI-Generated Election Claims in Social Media
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
Contextual Evaluation of Large Language Models for Classifying Tropical and Infectious Diseases
by: Asiedu, Mercy, et al.
Published: (2024)
by: Asiedu, Mercy, et al.
Published: (2024)
GROUNDEDKG-RAG: Grounded Knowledge Graph Index for Long-document Question Answering
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
GerPS-Compare: Comparing NER methods for legal norm analysis
by: Bachinger, Sarah T., et al.
Published: (2024)
by: Bachinger, Sarah T., et al.
Published: (2024)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
by: Huo, Feiye, et al.
Published: (2025)
by: Huo, Feiye, et al.
Published: (2025)
Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
by: Medvecki, Darija, et al.
Published: (2024)
by: Medvecki, Darija, et al.
Published: (2024)
LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability
by: Xu, Xuan, et al.
Published: (2025)
by: Xu, Xuan, et al.
Published: (2025)
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
by: Xu, Jillian, et al.
Published: (2025)
by: Xu, Jillian, et al.
Published: (2025)
Learning the Topic, Not the Language: How LLMs Classify Online Immigration Discourse Across Languages
by: Nasuto, Andrea, et al.
Published: (2025)
by: Nasuto, Andrea, et al.
Published: (2025)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
by: Achara, Akshit, et al.
Published: (2025)
by: Achara, Akshit, et al.
Published: (2025)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
by: Jung-Mok, Lee, et al.
Published: (2026)
by: Jung-Mok, Lee, et al.
Published: (2026)
multiMentalRoBERTa: A Fine-tuned Multiclass Classifier for Mental Health Disorder
by: Islam, K M Sajjadul, et al.
Published: (2025)
by: Islam, K M Sajjadul, et al.
Published: (2025)
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
by: Ye, Liqin, et al.
Published: (2025)
by: Ye, Liqin, et al.
Published: (2025)
Berta: an open-source, modular tool for AI-enabled clinical documentation
by: Vaid, Samridhi, et al.
Published: (2026)
by: Vaid, Samridhi, et al.
Published: (2026)
An In-Vitro Study on Cross-Lingual Generalization in Language Models
by: Cosma, Adrian
Published: (2026)
by: Cosma, Adrian
Published: (2026)
It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers
by: Clavié, Benjamin, et al.
Published: (2025)
by: Clavié, Benjamin, et al.
Published: (2025)
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
by: AlDahoul, Nouar, et al.
Published: (2024)
by: AlDahoul, Nouar, et al.
Published: (2024)
Dopamin: Transformer-based Comment Classifiers through Domain Post-Training and Multi-level Layer Aggregation
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT
by: Xylogiannopoulos, Konstantinos, et al.
Published: (2025)
by: Xylogiannopoulos, Konstantinos, et al.
Published: (2025)
AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents
by: Briva-Iglesias, Vicent, et al.
Published: (2026)
by: Briva-Iglesias, Vicent, et al.
Published: (2026)
Similar Items
-
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
by: Vargas, Francisco, et al.
Published: (2025) -
Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap
by: Panda, Akash Kumar, et al.
Published: (2026) -
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
by: Yuan, Mingruo, et al.
Published: (2025) -
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025) -
Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
by: Lenz, Stefan, et al.
Published: (2025)