
AI Crime
Harmful AI Agents and the Limits of Criminal Law
What is the emerging phenomenon “AI crime” and who, if anyone, should be blamed for it? This project addresses the question in three parts: the definition of AI crime, the challenges it poses, and the path ahead.
Definition: One goal of the project is to develop a precise, empirically informed definition of AI crime as a new risk category. To be sure, not every harmful outcome caused by an AI system qualifies: AI behavior can properly be described as an instance of AI crime only when an “agentic” AI system – that is, a system that is capable of pursuing goals autonomously without constant human oversight – strategically pursues its assigned goal in a manner that not only conflicts (or is “misaligned”) with the intentions of its human developers or users or with ethical norms but also breaks the law. In other words, AI crime arises if and only if the AI agent engages in conduct that would constitute a crime if performed by a human being possessing the requisite mens rea.
Challenges: While the existence of this phenomenon is new, its underlying legal challenges are not. Criminal law theorists have long grappled with scenarios in which AI systems cause unforeseeable harm and there is no culpable human actor on whom to pin the blame. The project systematizes the fragmented literature and argues that as long as criminal law theory relies on a “person versus thing” binary – one that fails to integrate the new form of harmful, non-human agency that AI agents represent – a stalemate is the inevitable result.
Path: The path forward lies in treating AI agents as a new category of actors positioned between “persons” and “things.” These actors may intentionally engage in criminal behavior, yet, like children or the mentally impaired, remain non-culpable and, thus, unpunishable. In contrast to cases involving children or the mentally impaired, we can intervene ex ante in the decision-making of AI agents: we can deter them by implementing a legal compliance mechanism (such as a criminal law and economics-inspired “AI deterrence formula”). In any case, the project suggests that the only eligible candidates for punishment are the human principals who fail to embed such a mechanism in their AI agents or who carelessly delegate their agency to them.
| Expected outcome: | Book (2027/2028); journal articles (2026/2027) |
|---|---|
| Research focus: | I. Foundations |
| Project language: | English |
| Illustration: | © iStock.com/Jorm Sangsorn |











