The widespread adoption of large language models (LLMs) requires safeguarding them against malicious prompts. Existing neural guardrails suffer from a key limitation: their decisions emerge from opaque, high-dimensional transformations that resist human inspection, making them unsuitable for regulated or safety-critical deployments requiring auditability and human oversight. This article faces this gap by introducing PromptSentinel, a neurosymbolic framework that: 1) transforms input prompts into a structured semantic representation (i.e., intent, actions, target, and constraints) through an LLM-based semantic annotation process; 2) projects these primitives into a supervised, polarized embedding space; and 3) discretizes them into symbolic labels through clustering. A rule-based engine then operates on these symbols producing safety decisions whose decision logic is fully auditable. Evaluated under out-of-distribution, nonadaptive adversarial conditions, PromptSentinel matches the performance of the strongest specialized neural guardrails in both detection and usability. On the most adversarial benchmark, HarmBench, it significantly outperforms ShieldGemma, reducing the attack success rate (ASR) from 8% to 2% (p = 0.012). Crucially, PromptSentinel achieves this level of safety performance while grounding every verdict in an inspectable, human-readable rule set, providing a degree of auditability that black-box neural guardrails cannot offer by construction. Taken together, these results challenge the presumed trade-off between interpretability and safety, showing that transparent, rule-based guardrailing can achieve competitive performance with state-of-the-art specialized neural approaches while remaining inspectable and governable at the decision-rule level.

PromptSentinel: a neurosymbolic system for malicious prompt detection / Di Gisi, M., Fenza, G., Gallo, M., Loia, V., Pedrycz, W.. - In: IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS. - ISSN 2329-924X. - (2026). [10.1109/tcss.2026.3738120]

PromptSentinel: a neurosymbolic system for malicious prompt detection

Di Gisi Maria;
2026

Abstract

The widespread adoption of large language models (LLMs) requires safeguarding them against malicious prompts. Existing neural guardrails suffer from a key limitation: their decisions emerge from opaque, high-dimensional transformations that resist human inspection, making them unsuitable for regulated or safety-critical deployments requiring auditability and human oversight. This article faces this gap by introducing PromptSentinel, a neurosymbolic framework that: 1) transforms input prompts into a structured semantic representation (i.e., intent, actions, target, and constraints) through an LLM-based semantic annotation process; 2) projects these primitives into a supervised, polarized embedding space; and 3) discretizes them into symbolic labels through clustering. A rule-based engine then operates on these symbols producing safety decisions whose decision logic is fully auditable. Evaluated under out-of-distribution, nonadaptive adversarial conditions, PromptSentinel matches the performance of the strongest specialized neural guardrails in both detection and usability. On the most adversarial benchmark, HarmBench, it significantly outperforms ShieldGemma, reducing the attack success rate (ASR) from 8% to 2% (p = 0.012). Crucially, PromptSentinel achieves this level of safety performance while grounding every verdict in an inspectable, human-readable rule set, providing a degree of auditability that black-box neural guardrails cannot offer by construction. Taken together, these results challenge the presumed trade-off between interpretability and safety, showing that transparent, rule-based guardrailing can achieve competitive performance with state-of-the-art specialized neural approaches while remaining inspectable and governable at the decision-rule level.
2026
Explainable AI (XAI), Human-AI interaction, Jailbreak detection, LLM guardrails, LLM safety, Malicious prompt detection, Neurosymbolic AI, Trustworthy AI
File in questo prodotto:
File Dimensione Formato  
PromptSentinel_A_Neurosymbolic_System_for_Malicious_Prompt_Detection.pdf

Accesso aperto

Descrizione: PromptSentinel: A Neurosymbolic System for Malicious Prompt Detection
Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 1.17 MB
Formato Adobe PDF
1.17 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11771/44619
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • OpenAlex 0
social impact