Logo image
A Hybrid Neural Semantic Framework for Prompt Injection Detection in LLMs
Conference proceeding

A Hybrid Neural Semantic Framework for Prompt Injection Detection in LLMs

Hania Siddiqui, Rizwan Taj, Qadeer Yasin, Muhammad Sarfraz Shahid, Saba Mehfooz and Assma Bibi
International Conference on Computing, Mathematics and Engineering Technologies (Online), pp.1-5
5th International Conference on Computing, Mathematics and Engineering Technologies (iCoMET2026) (Sukkur, Pakistan, 22/05/2026–23/05/2026)
05/2026

Abstract

Accuracy Adversarial Attacks Conferences Equations Hybrid Neural-Semantic Framework Labeling Large language models Large Language Models (LLMs) Modeling Printing Prompt Injection Detection Security for AI models Signal detection Tagging Training
Prompt injection attacks imply that they lead to a very severe risk to the dependability of Large Language Model (LLM) systems by allowing a user to control the behavior of the model using constructed inputs. This paper proposes a Hybrid-PI neural-semantic architecture to improve the detection of these attacks. Hybrid-PI with the DeepSet prompt-injection dataset uses a neural prompt encoder with a Semantic Consistency Oracle to jointly learn surface and higher-level semantic inconsistencies via a trade-off between the two. The experimental findings indicate that Hybrid-PI performs much better than the classical and transformer-based baselines, achieving 97.4% accuracy and 97.5% F1-score, with significantly reduced false positives. The robustness tests also indicate the ability to withstand paraphrase and obfuscation-based adversarial variations. The results suggest that Hybrid-PI is a viable and scalable approach to strengthening safety controls around LLMs against emerging threats such as prompt injection.

Details

Metrics

1 Record Views
Logo image