Abstract
The rapid integration of Large Language Models (LLMs) into critical personal and professional environments has exacerbated security risks, particularly adversarial attacks such as prompt injection and jailbreaking, which aim to bypass safety alignment. This study evaluates the efficacy of NVIDIA’s Llama-3.1-nemoguard-8b-content-safety model acting as a semantic firewall to mitigate these threats. To ensure a robust assessment, we utilized the ‘Do Not Answer’ dataset, augmented with 939 synthetically generated benign prompts to create a balanced corpus of 1878 samples. The evaluation methodology encompasses a risk-category analysis, standard binary classification metrics, and a novel metric, the Compensation Rate, which measures the firewall’s ability to block responses when the underlying LLM fails. Results indicate a high Precision (94.57%) but a moderate Sensitivity (51.97%), uncovering a critical performance trade-off: the model exhibits a conservative bias, prioritizing high precision to minimize false positives at the expense of recall for nuanced adversarial prompts, particularly in categories involving sensitive data leakage and misinformation. Furthermore, the proposed Compensation Rate achieved 34.8%, suggesting that the semantic firewall successfully mitigated 34.8% of instances where the foundational LLM’s internal safety alignment failed. These findings indicate that while the system effectively blocks explicit threats, its efficacy as a secondary defense diminishes against context-dependent vulnerabilities, notably data exfiltration and misinformation.
Author supplied keywords
Cite
CITATION STYLE
Azambuja, A. J., Guilherme, M., Castro, J. V. F. de, Lima, J. P. de O., Oliveira, L. B., & Soares, A. da S. (2026). Evaluation of NeMo Guardrails as a Firewall for User–LLM Interaction. Future Internet, 18(5). https://doi.org/10.3390/fi18050252
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.