WisdomInterrogatory (LuWen): An Open-Source Legal Large Language Model Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yiquan, Liu, Yuhang, Liu, Yifei, Li, Ang, Zhou, Siying, Kuang, Kun, Wu, Fei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918437697093632
author Wu, Yiquan
Liu, Yuhang
Liu, Yifei
Li, Ang
Zhou, Siying
Kuang, Kun
Wu, Fei
author_facet Wu, Yiquan
Liu, Yuhang
Liu, Yifei
Li, Ang
Zhou, Siying
Kuang, Kun
Wu, Fei
contents Large language models have demonstrated remarkable capabilities across a wide range of natural language processing tasks, yet their application in the legal domain remains challenging due to the specialized terminology, complex reasoning requirements, and rapidly evolving legal knowledge involved. In this paper, we present WisdomInterrogatory (LuWen), an open-source Chinese legal language model built upon the Baichuan foundation model through three key techniques: continual pre-training on a large-scale legal corpus, supervised fine-tuning with carefully curated legal instruction data, and retrieval-augmented generation integrated with a comprehensive legal knowledge base. We evaluate LuWen on five representative legal tasks spanning both prediction and generation settings, including legal judgment prediction, judicial examination, legal text summarization, law article question answering, and judicial decision reasoning. Experimental results show that LuWen outperforms several strong baselines, demonstrating the effectiveness of our approach in adapting general-purpose language models to the legal domain.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06737
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WisdomInterrogatory (LuWen): An Open-Source Legal Large Language Model Technical Report
Wu, Yiquan
Liu, Yuhang
Liu, Yifei
Li, Ang
Zhou, Siying
Kuang, Kun
Wu, Fei
Computation and Language
Artificial Intelligence
Large language models have demonstrated remarkable capabilities across a wide range of natural language processing tasks, yet their application in the legal domain remains challenging due to the specialized terminology, complex reasoning requirements, and rapidly evolving legal knowledge involved. In this paper, we present WisdomInterrogatory (LuWen), an open-source Chinese legal language model built upon the Baichuan foundation model through three key techniques: continual pre-training on a large-scale legal corpus, supervised fine-tuning with carefully curated legal instruction data, and retrieval-augmented generation integrated with a comprehensive legal knowledge base. We evaluate LuWen on five representative legal tasks spanning both prediction and generation settings, including legal judgment prediction, judicial examination, legal text summarization, law article question answering, and judicial decision reasoning. Experimental results show that LuWen outperforms several strong baselines, demonstrating the effectiveness of our approach in adapting general-purpose language models to the legal domain.
title WisdomInterrogatory (LuWen): An Open-Source Legal Large Language Model Technical Report
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.06737