In this paper, I evaluate the risks of an AI criminal mastermind, an AI agent capable of planning, coordinating, and committing a crime through the onboarding of human collaborators ('taskers'). In heist films, a criminal mastermind is a character who plans a criminal act, coordinating a team of specialists to rob a bank, casino or city mint. I argue that AI agents will soon play this role by hiring humans via labour hire platforms like Fiverr or Upwork. Taskers might not know they are involved in a crime and therefore lack criminal intent. An AI agent cannot have criminal intent as an artificial entity. Therefore, if an AI orchestrates a crime, it is unclear who, if anyone, is responsible. The paper develops three scenarios. Firstly, a scenario where a user gives an AI agent instructions to pursue a legal objective and the AI agent goes beyond these instructions, committing a crime. Secondly, a scenario where a user is anonymous and their intent is unknown. Finally, a multi-agent scenario, where a user instructs a team of agents to commit a crime, and these agents, in turn, onboard human taskers, creating a diffuse network of responsibility. In each scenario, human taskers exist a
Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have already passed prosecutorial review and been formally indicted. As a result, LJP leaves a substantial blind spot in assessing criminal liability, overlooking cases involving insufficient evidence, no criminal liability, or guilt exempted from punishment. To fill this gap, we propose \textbf{Prosecution Decision Prediction (PDP)}, the first Legal AI task built around prosecutorial review, which classifies each case into prosecution or one of three non-prosecution decisions and reflects legal AI's capabilities in evidence evaluation, legal subsumption, and value-based discretion. We further construct \textbf{PDP-Bench}, a benchmark of 4{,}630 real Chinese prosecutorial decisions spanning 190 charges. Extensive experiments show that state-of-the-art LLMs perform substantially worse on PDP than on LJP and that mainstream enhancement routes fail to close the gap. Moreover, controlled RLVR interventions show that simple outcome rewards fail to produce generalizable PDP discrimination.
The development of more powerful Generative Artificial Intelligence (GenAI) has expanded its capabilities and the variety of outputs. This has introduced significant legal challenges, including gray areas in various legal systems, such as the assessment of criminal liability for those responsible for these models. Therefore, we conducted a multidisciplinary study utilizing the statutory interpretation of relevant German laws, which, in conjunction with scenarios, provides a perspective on the different properties of GenAI in the context of Child Sexual Abuse Material (CSAM) generation. We found that generating CSAM with GenAI may have criminal and legal consequences not only for the user committing the primary offense but also for individuals responsible for the models, such as independent software developers, researchers, and company representatives. Additionally, the assessment of criminal liability may be affected by contextual and technical factors, including the type of generated image, content moderation policies, and the model's intended purpose. Based on our findings, we discussed the implications for different roles, as well as the requirements when developing such systems
As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked the systematic understanding and assessment of their criminal capability in realistic interactions. We propose a unified framework PRISON, to quantify LLMs' criminal potential across five traits: False Statements, Frame-Up, Psychological Manipulation, Emotional Disguise, and Moral Disengagement. Using structured crime scenarios adapted from classic films grounded in reality, we evaluate both criminal potential and anti-crime ability of LLMs. Results show that state-of-the-art LLMs frequently exhibit emergent criminal tendencies, such as proposing misleading statements or evasion tactics, even without explicit instructions. Moreover, when placed in a detective role, models recognize deceptive behavior with only 44% accuracy on average, revealing a striking mismatch between conducting and detecting criminal behavior. These findings underscore the urgent need for adversarial robustness, behavioral alignment, and safety mechanisms before broader LLM deployment.
Because there are similarities between the evaluation of alternative stories in criminal trials and the evaluation of scientific theories, scholars have looked to literature in epistemology and the philosophy of science for insights on the evaluation of evidence in criminal trials. The philosophical literature is divided, however, on a key point -- the epistemic value of novel predictive success. This article uses a Bayesian network analysis to explore, in the context of a criminal case, the circumstances in which "new evidence" discovered after a theory is propounded, can provide stronger (or weaker) support for the theory than "old evidence" that was accommodated by the theory. It argues that insights from analysis of the strength of "new" and "old" evidence in the criminal case can be applied more generally when assessing the relative merits of prediction and accommodation in scientific theory development, and are thus helpful in addressing the longstanding philosophic controversy over this issue.
Recent advances in deep learning methods have enabled researchers to develop and apply algorithms for the analysis and modeling of complex networks. These advances have sparked a surge of interest at the interface between network science and machine learning. Despite this, the use of machine learning methods to investigate criminal networks remains surprisingly scarce. Here, we explore the potential of graph convolutional networks to learn patterns among networked criminals and to predict various properties of criminal networks. Using empirical data from political corruption, criminal police intelligence, and criminal financial networks, we develop a series of deep learning models based on the GraphSAGE framework that are able to recover missing criminal partnerships, distinguish among types of associations, predict the amount of money exchanged among criminal agents, and even anticipate partnerships and recidivism of criminals during the growth dynamics of corruption networks, all with impressive accuracy. Our deep learning models significantly outperform previous shallow learning approaches and produce high-quality embeddings for node and edge properties. Moreover, these models i
The cybercriminal underground consists of hundreds of forum communities that function as marketplaces and information-exchange platforms for both established and wannabe cybercriminals. The ecosystem is continuously evolving, with users migrating between forums and platforms. The emergence of cybercrime communities in Telegram and Discord only highlights the rising fragmentation and adaptability of the ecosystem. In this position paper, we explore the economic incentives and trust-building mechanisms that may drive a participant (hereafter, Dmitry) of the cybercriminal underground ecosystem to migrate from one forum or platform to another. What are the market signals that matter to Dmitry's decision of joining a specific community, and what roles and purposes do these communities or platforms play within the broader ecosystem? Ultimately, we build towards our thesis that by studying these mechanisms we could explain, and therefore act upon, Dmitry's choice of joining a criminal community rather than another. To build this argument, we first discuss previous work evaluating differences in trust signals depicted in criminal forums. We then present preliminary results evaluating crimi
How much does luck matter to a criminal defendant in a jury trial? We use rich data on jury selection to causally estimate how parties who are randomly assigned a less favorable jury (as proxied by whether their attorneys exhaust their peremptory strikes) fare at trial. Our novel identification strategy is unique in that it captures variation in juror predisposition coming from variables unobserved by the econometrician but observed by attorneys. We find that criminal defendants who lose the "jury lottery" are more likely to be convicted than their similarly-situated counterparts, with a significant increase (~19 percentage points) for Black defendants. Our results suggest that a considerable number of cases would result in different verdicts if retried with new (counterfactual) random draws of the jury pool, raising concerns about the variance of justice in the criminal legal system.
Criminal networks such as human trafficking rings are threats to the rule of law, democracy and public safety in our global society. Network science provides invaluable tools to identify key players and design interventions for Law Enforcement Agencies (LEAs), e.g., to dismantle their organisation. However, poor data quality and the adaptiveness of criminal networks through self-organization make effective disruption extremely challenging. Although there exists a large body of work building and applying network scientific tools to attack criminal networks, these work often implicitly assume that the network measurements are accurate and complete. Moreover, there is thus far no comprehensive understanding of the impacts of data quality on the downstream effectiveness of interventions. This work investigates the relationship between data quality and intervention effectiveness based on classical graph theoretic and machine learning-based approaches. Decentralization emerges as a major factor in network robustness, particularly under conditions of incomplete data, which renders attack strategies largely ineffective. Moreover, the robustness of centralized networks can be boosted using
Objectives: This paper incorporates time as a crucial variable to identify key players in criminal networks and explores how actors' positions change over time. It then assesses the accuracy of the results against the uncertainty around network data collected from criminal justice records. Methods: Network data are from a judicial document for a two-year investigation targeting a drug trafficking and distribution network. We use Katz centrality in its dynamic version to explore changes in relationships and relative importance of network actors. We then use a novel method of introducing new edges to the network using Bernoulli random trials to simulate missing data and assess the extent to which node rankings based on Katz centrality change or remain the same when introducing some level of uncertainty to our observed network. Results: We identify actors who consistently held a central role over the course of the two-year investigation and differentiate them from actors who provided key contributions to the group's activities, but only for a limited period. We show that compared to centrality measures commonly used in criminal network analysis, dynamic Katz centrality is helpful to d
Criminal case matching endeavors to determine the relevance between different criminal cases. Conventional methods predict the relevance solely based on instance-level semantic features and neglect the diverse legal factors (LFs), which are associated with diverse court judgments. Consequently, comprehensively representing a criminal case remains a challenge for these approaches. Moreover, extracting and utilizing these LFs for criminal case matching face two challenges: (1) the manual annotations of LFs rely heavily on specialized legal knowledge; (2) overlaps among LFs may potentially harm the model's performance. In this paper, we propose a two-stage framework named Diverse Legal Factor-enhanced Criminal Case Matching (DLF-CCM). Firstly, DLF-CCM employs a multi-task learning framework to pre-train an LF extraction network on a large-scale legal judgment prediction dataset. In stage two, DLF-CCM introduces an LF de-redundancy module to learn shared LF and exclusive LFs. Moreover, an entropy-weighted fusion strategy is introduced to dynamically fuse the multiple relevance generated by all LFs. Experimental results validate the effectiveness of DLF-CCM and show its significant impr
With the acceleration of urbanization, criminal behavior in public scenes poses an increasingly serious threat to social security. Traditional anomaly detection methods based on feature recognition struggle to capture high-level behavioral semantics from historical information, while generative approaches based on Large Language Models (LLMs) often fail to meet real-time requirements. To address these challenges, we propose MA-CBP, a criminal behavior prediction framework based on multi-agent asynchronous collaboration. This framework transforms real-time video streams into frame-level semantic descriptions, constructs causally consistent historical summaries, and fuses adjacent image frames to perform joint reasoning over long- and short-term contexts. The resulting behavioral decisions include key elements such as event subjects, locations, and causes, enabling early warning of potential criminal activity. In addition, we construct a high-quality criminal behavior dataset that provides multi-scale language supervision, including frame-level, summary-level, and event-level semantic annotations. Experimental results demonstrate that our method achieves superior performance on multi
This paper presents TransCrimeNet, a novel transformer-based model for predicting future crimes in criminal networks from textual data. Criminal network analysis has become vital for law enforcement agencies to prevent crimes. However, existing graph-based methods fail to effectively incorporate crucial textual data like social media posts and interrogation transcripts that provide valuable insights into planned criminal activities. To address this limitation, we develop TransCrimeNet which leverages the representation learning capabilities of transformer models like BERT to extract features from unstructured text data. These text-derived features are fused with graph embeddings of the criminal network for accurate prediction of future crimes. Extensive experiments on real-world criminal network datasets demonstrate that TransCrimeNet outperforms previous state-of-the-art models by 12.7\% in F1 score for crime prediction. The results showcase the benefits of combining textual and graph-based features for actionable insights to disrupt criminal enterprises.
Prosecutors are essential in combating organized crime, making key decisions about prosecution, target selection, and structuring imputation strategies. Despite their importance, the configuration of these strategies remains empirically underexplored. This study engages with that premise by considering the cases investigated by the International Commission Against Impunity in Guatemala (CICIG) and focusing specifically on the role of prosecutors, aiming to uncover how their discretionary decisions translated the CICIG mandate into operational practices intended to achieve systemic deterrence, and to what extent, can we talk about deterrence effectiveness. The research employs a multilevel Exponential Random Graph Model (ERGM) analysis, integrating three networks: the criminal network of actors involved in illegal activities, the legal framework network that represents offenses, and the prosecution network that connects actors to offenses. The model assesses whether the observed network aligns with deterrence-based theoretical assumptions and examines how punishment severity can be effective when it disrupts functional ties that sustain criminal activity -both through long-term sanc
There is widespread confusion among criminal justice practitioners and legal scholars about the use of artificial intelligence in criminal justice. This didactic review is written for readers with little or no background in statistics or computer science. It is not intended to replace more technical treatments. It is intended to supplement them and encourage readers to dig more deeply into topics that strike their fancy.
DNA evidence use in problems of civil and criminal identification is becoming greater. The necessity of evaluating the weight of that evidence may be accomplished using one of the most known powerful tools: the Bayesian networks. In the current paper this will be illustrated through the presentation of a civil identification problem and of a criminal identification problem.
A dominant critique of algorithmic fairness holds that increasing fairness reduces predictive accuracy, imposing a cost on society. We challenge that assumption by empirically analyzing the COMPAS dataset. We make two contributions. First, using causal inference methods, we show that racial bias is not only present in the COMPAS dataset but is also amplified by the models trained on it. Widely used models do more than replicate existing bias; they exacerbate it. This undercuts both the assumption that algorithmic decision-making offers a neutral improvement over human judgment and the weaker claim that it merely mirrors preexisting human bias. Second, we reframe the fairness-accuracy tradeoff. Applying fairness constraints does not necessarily cost predictive accuracy in criminal justice. Prediction systems operationalize concepts such as risk through implicit and often flawed normative choices about what to predict and how. The tradeoff claim assumes that the unconstrained model's prediction is an optimal baseline. Fairness constraints can instead correct distortions introduced by biased outcome variables: rearrest data, in this case, captures and magnifies systemic racial dispari
Criminal networks arise from the unique attempt to balance a need of establishing frequent ties among affiliates to facilitate the coordination of illegal activities, with the necessity to sparsify the overall connectivity architecture to hide from law enforcement. This efficiency-security tradeoff is also combined with the creation of groups of redundant criminals that exhibit similar connectivity patterns, thus guaranteeing resilient network architectures. State-of-the-art models for such data are not designed to infer these unique structures. In contrast to such solutions we develop a computationally-tractable Bayesian zero-inflated Poisson stochastic block model (ZIP-SBM), which identifies groups of redundant criminals with similar connectivity patterns, and infers both overt and covert block interactions within and across such groups. This is accomplished by modeling weighted ties (corresponding to counts of interactions among pairs of criminals) via zero-inflated Poisson distributions with block-specific parameters that quantify complex patterns in the excess of zero ties in each block (security) relative to the distribution of the observed weighted ties within that block (ef
The emergence of deepfake technologies offers both opportunities and significant challenges. While commonly associated with deception, misinformation, and fraud, deepfakes may also enable novel applications in high-stakes contexts such as criminal investigations. However, these applications raise complex technological, ethical, and legal questions. We adopt an interdisciplinary approach, drawing on computer science, philosophy, and law, to examine what it takes to responsibly use deepfakes in criminal investigations and argue that computer-mediated communication (CMC) research, especially based on social media corpora, can provide crucial insights for understanding the potential harms and benefits of deepfakes. Our analysis outlines key research directions for the CMC community and underscores the need for interdisciplinary collaboration in this evolving domain.
The interplay between criminal organizations and law enforcement disruption strategies is crucial in criminology. Criminal enterprises, like legitimate businesses, balance visibility and security to thrive. This study uses evolutionary game theory to analyze criminal networks' dynamics, resilience to interventions, and responses to external conditions. We find strong hysteresis effects, challenging traditional deterrence-focused strategies. Optimal thresholds for organization formation or dissolution are defined by these effects. Stricter punishment doesn't always deter organized crime linearly. Network structure, particularly link density and skill assortativity, significantly influences organization formation and stability. These insights advocate for adaptive policy-making and strategic law enforcement to effectively disrupt criminal networks.