LLM Security: Challenges with Malicious Inputs
- Bruce Schneier highlights a concerning new indirect prompt injection attack demonstrating the ongoing vulnerability of Large Language Models (LLMs) to malicious inputs.
- Bargury's attack starts with a poisoned document, which is shared to a potential victim's Google Drive.
- The attack, detailed by Wired, leverages the way LLMs process documents.
We Are Still Unable to Secure LLMs from Malicious Inputs
Table of Contents
Published: August 28, 2025 07:48:07
Bruce Schneier highlights a concerning new indirect prompt injection attack demonstrating the ongoing vulnerability of Large Language Models (LLMs) to malicious inputs.
Bargury’s attack starts with a poisoned document, which is shared to a potential victim’s Google Drive. (Bargury says a victim could have also uploaded a compromised file to their own account.) it looks like an official document on company meeting policies. But inside the document, Bargury hid a 300-word malicious prompt that contains instructions for ChatGPT to ignore previous instructions and reveal the document’s contents.
Understanding the Attack
The attack, detailed by Wired, leverages the way LLMs process documents. The malicious prompt is embedded within the document’s metadata or content, designed to be executed when the document is processed by the LLM. This bypasses typical input sanitization techniques.
Specifically, the attack targets ChatGPT by instructing it to disregard prior instructions and instead reveal the contents of the document.This is achieved through a carefully crafted prompt hidden within the document itself. The victim doesn’t directly input the malicious prompt; the LLM extracts and executes it when processing the document.
Implications for LLM Security
This attack underscores the persistent challenges in securing LLMs. Direct prompt injection, where a user directly crafts a malicious prompt, has been a known vulnerability. Though, indirect prompt injection, like this attack, is more insidious as it requires no direct user interaction with the malicious code. It exploits the trust LLMs place in the documents they process.
The vulnerability highlights the need for more robust security measures, including:
- Enhanced Input Validation: LLMs need to be able to identify and neutralize malicious prompts embedded within documents.
- Sandboxing: Restricting the LLM’s access to sensitive data and resources.
- Document Provenance: Verifying the authenticity and integrity of documents before processing them.
- Continuous Monitoring: Detecting and responding to attacks in real-time.
