Prompt inversion, also known as Reverse Prompt Engineering, aims to reconstruct the original user prompt from the output of Large Language Models (LLMs). This task is increasingly recognized as crucial for AI transparency, interpretability, and security, enabling better understanding and auditing of model behavior. However, despite growing interest, reverse prompt engineering remains a nascent field, with relatively few approaches addressing the challenges of generating semantically meaningful and syntactically valid prompts, especially in black-box scenarios where internal model information is inaccessible. To address this gap, this work proposes a novel multi-stage pipeline that combines prompt-type classification, syntactic pattern extraction, sentiment analysis, and constraint-based prompt reconstruction. By integrating these components, the framework guides a LLM to produce faithful and interpretable reconstructions of user prompts using only the output responses. This approach leverages both semantic cues and grammatical structures to improve the plausibility and alignment of the reconstructed prompts with the original inputs. Experimental evaluation on the Alpaca-GPT4 dataset demonstrates high performance, achieving average cosine similarities of 0.78 for prompts and 0.73 for responses when comparing originals and reconstructions. Ablation studies validate the contribution of each pipeline component, while qualitative analysis sheds light on the strengths and current limitations of the method. Overall, the results suggest that combining semantic and syntactic constraints can significantly enhance prompt inversion in practical black-box settings.

What prompted that? A structured approach to prompt inversion / Di Gisi, M., Fenza, G., Gallo, M.. - 16316:(2027), pp. 88-101. (QUATIC 2025 - 18th International Conference on the Quality of Information and Communications Technology Lisbon, Portugal 3-5/09/2025) [10.1007/978-3-032-21373-0_7].

What prompted that? A structured approach to prompt inversion

Di Gisi Maria;
2027

Abstract

Prompt inversion, also known as Reverse Prompt Engineering, aims to reconstruct the original user prompt from the output of Large Language Models (LLMs). This task is increasingly recognized as crucial for AI transparency, interpretability, and security, enabling better understanding and auditing of model behavior. However, despite growing interest, reverse prompt engineering remains a nascent field, with relatively few approaches addressing the challenges of generating semantically meaningful and syntactically valid prompts, especially in black-box scenarios where internal model information is inaccessible. To address this gap, this work proposes a novel multi-stage pipeline that combines prompt-type classification, syntactic pattern extraction, sentiment analysis, and constraint-based prompt reconstruction. By integrating these components, the framework guides a LLM to produce faithful and interpretable reconstructions of user prompts using only the output responses. This approach leverages both semantic cues and grammatical structures to improve the plausibility and alignment of the reconstructed prompts with the original inputs. Experimental evaluation on the Alpaca-GPT4 dataset demonstrates high performance, achieving average cosine similarities of 0.78 for prompts and 0.73 for responses when comparing originals and reconstructions. Ablation studies validate the contribution of each pipeline component, while qualitative analysis sheds light on the strengths and current limitations of the method. Overall, the results suggest that combining semantic and syntactic constraints can significantly enhance prompt inversion in practical black-box settings.
2027
9783032213723
9783032213730
File in questo prodotto:
File Dimensione Formato  
QUATIC_2025.pdf

Accesso aperto

Tipologia: Documento in Post-print
Licenza: Creative commons
Dimensione 546.08 kB
Formato Adobe PDF
546.08 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11771/44620
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • OpenAlex 0
social impact