Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of natural language processing tasks. Yet, their vulnerability to prompt-based adversarial attacks, particularly jailbreaks and prompt injections, poses significant safety and ethical concerns. In this study, we present DualGuard, a novel two-stage defense framework designed to identify and suppress unsafe outputs during prompt-based evaluations of LLMs. DualGuard combines semantic-level self-evaluation, wherein a model reflects on its own response for safety violations, with policy-level verification conducted by a high-accuracy, policy-aligned external judge (e.g., GPT-4o). We benchmark DualGuard against both unprotected LLMs and those employing Semantic Smoothing, a defense technique based on semantically preserving input perturbations. Through systematic experimentation on five prominent LLMs and ten adversarial prompts crafted to bypass standard safeguards, we demonstrate that DualGuard significantly enhances resilience against jailbreak attempts while maintaining interpretability and model usability. Furthermore, we analyze multiple semantic judge variants and a customized hybrid judge, which combines two judge variants. Our results show that DualGuard outperforms existing defense mechanisms across all tested models, blocking up to 100% of harmful responses while maintaining a low false positive rate (6%) on benign prompts, demonstrating strong usability preservation for real-world deployment. We also evaluated DualGuard on a mixed benchmark derived from the HarmBench dataset, consisting of 150 prompts (120 unsafe and 30 safe) covering chemical and biological threats, cybercrime, misinformation, harassment, and copyright violations. On this larger and more diverse benchmark, DualGuard, applied with Llama Guard as external judge, reduces attack success to 0.0-7.5% across four open-weight models while maintaining a high safe answer rate, demonstrating its robustness and model-agnostic design. This work contributes to ongoing LLM safety research by introducing a practical and interpretable defense approach that can be applied across different model families and real-world applications.
Author supplied keywords
Cite
CITATION STYLE
Rasheed, A. S. A., & Masud, M. M. (2026). Effective Defense Strategies Against Jailbreaking in Large Language Models. IEEE Access, 14, 52591–52609. https://doi.org/10.1109/ACCESS.2026.3679168
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.