首页 > AI前沿 > Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

arXiv自然语言 2026-09-17 17:38 5 阅读 查看原文

Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely unexplored.

To fill this gap, we introduce the Profiling, Investigation, and Judgment (PIJ), comprising 2,500 real homicide cases from five countries.

PIJ evaluates LLMs across three tasks that span the entire criminal investigation pipeline: criminal profiling, which requires abductive reasoning to infer suspect attributes from fragmentary scene evidence, crime process reconstruction, which tests structured information extraction, and sentence prediction, which demands legal deductive reasoning.

We evaluate 9 powerful LLMs and find that performance degrades systematically as tasks shift from explicit fact extraction to implicit reasoning over unknown suspect profiles.

Categories requiring inferential reasoning, such as motivation and victim-offender relationships, remain the primary bottlenecks.

Further analysis reveals substantial gaps between LLMs and human experts, along with pervasive biases in gender, age, and motive attribution.

Our findings indicate that pre-arrest inference from incomplete evidence remains an open challenge.