In the Italian judiciary system, Public Prosecutors’ Offices still rely on heterogeneous and partially paper-based document workflows, where crime reports and related attachments are often printed, manually signed and annotated, scanned, and finally stored as image-based files. As a consequence, a significant portion of prosecutorial documentation remains only partially machine-readable, limiting the effectiveness of digital case-management systems. In this context, Optical Character Recognition (OCR) and Named Entity Recognition (NER) are enabling technologies that transform unstructured, non-searchable judicial documents into computationally usable legal information. This paper analyzes the technical, organizational, and legal challenges associated with OCR-based processing of crime reports, in the Italian and foreign jurisdictions, and identifies the main methodological requirements for governance-aware NER models used in judicial environments. These include layout-aware document analysis, legal-domain adaptation, human-in-the-loop validation, and privacy-aware processing mechanisms that support pseudonymization and controlled access to sensitive data. Finally, the paper discusses the broader topic of computer-aided judicial digitalization, highlighting the need for reliable, privacy-aware pipelines capable of processing documents at scale in contemporary criminal justice systems.
Bridging digitalization gaps: OCR and NER for crime report automation in the Italian criminal justice system / Cerini, S.Y., Deidda, N., Lilli, S., Zucca, M.V., Fumera, G., Giacinto, G., Prinetto, P.. - 16900:(2027), pp. 541-558. (ARES - The International Conference on Availability, Reliability and Security Linköping, Sweden 24–27/08/2026) [10.1007/978-3-032-35576-8_29].
Bridging digitalization gaps: OCR and NER for crime report automation in the Italian criminal justice system
Yves Cerini Samuele
;Deidda Nicola
;Lilli Sara
;Zucca Maria Vittoria
;
2027
Abstract
In the Italian judiciary system, Public Prosecutors’ Offices still rely on heterogeneous and partially paper-based document workflows, where crime reports and related attachments are often printed, manually signed and annotated, scanned, and finally stored as image-based files. As a consequence, a significant portion of prosecutorial documentation remains only partially machine-readable, limiting the effectiveness of digital case-management systems. In this context, Optical Character Recognition (OCR) and Named Entity Recognition (NER) are enabling technologies that transform unstructured, non-searchable judicial documents into computationally usable legal information. This paper analyzes the technical, organizational, and legal challenges associated with OCR-based processing of crime reports, in the Italian and foreign jurisdictions, and identifies the main methodological requirements for governance-aware NER models used in judicial environments. These include layout-aware document analysis, legal-domain adaptation, human-in-the-loop validation, and privacy-aware processing mechanisms that support pseudonymization and controlled access to sensitive data. Finally, the paper discusses the broader topic of computer-aided judicial digitalization, highlighting the need for reliable, privacy-aware pipelines capable of processing documents at scale in contemporary criminal justice systems.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


