Abstract
Forensic science encompasses multiple specialized domains that reconstruct past events from traces, including digital forensics, which focuses on devices and networks. Large Language Models (LLMs) create a new kind of evidence, prompt-response traces, that traditional digital forensics does not explicitly address. To fill this gap, we introduce Prompt Forensics, the systematic collection, reconstruction, and analysis of LLM interactions to assess safety and support incident response. Drawing inspiration from classical forensic procedures, we propose Prompt Scene Investigation (PSI), a framework comprising three key components: (i) Prompt Processing, which captures interaction traces as tamper-resistant Prompt IDs forming a digital chain of custody; (ii) Replay, which re-enacts model behaviour under controlled conditions; and (iii) Investigation, which detects adversarial prompting and safety violations, and produces diagnostic reports. Prompt Forensics supports auditing and regulatory compliance for AI systems by making LLM safety incidents observable, replayable, and accountable. Preliminary experiments comparing different small language models provide an initial validation of the PSI framework, making this work a proof of concept for Prompt Forensics as a practical auditing approach for LLM systems.
Author supplied keywords
Cite
CITATION STYLE
Ledjaki, L. F., Chang, Y. C., Loutis, M., & Aïmeur, E. (2026). Prompt Scene Investigation: Uncovering the Evidence in LLMs. In WWW Companion 2026 - Companion Proceedings of the ACM Web Conference 2026 (pp. 284–290). Association for Computing Machinery, Inc. https://doi.org/10.1145/3774905.3795469
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.