The advancement of natural language processing (NLP) has expanded AI-based text classification in the legal domain.
However, accurately classifying legal documents remains challenging due to the complexity of legal texts and subtle differences between legal categories.
This study evaluates legal text classification models
using ten categories of Korean sexual offense precedents.
The results show that fine-tuning small-scale models such as KLUE-BERT on legal data outperforms general-purpose models such as GPT-3.5 and GPT-4.0, as well as traditional machine learning models.
KLUE-BERT achieved the highest accuracy of 99.3%, indicating that domain adaptation and fine-tuning can be more important than model size for legal document classification.
Explainable AI (XAI) techniques
We further employ explainable AI (XAI) techniques to analyze model predictions and misclassification cases.
XAI analysis identifies linguistic features influencing model decisions and limitations in capturing subtle textual cues.
Using KICS data
Using KICS data, which closely resembles real-world legal case records, we further evaluate the model's generalization capabilities.
and find that it struggles to interpret implicit contextual cues.
These findings highlight the importance of both performance and interpretability in legal AI and demonstrate how XAI can improve transparency in legal text classification.
AI-assisted tools
AI-assisted tools can support legal professionals in tasks including document classification, legal information retrieval, and case assessment.