Image Prompts: New Hack Steals Data from LLMs
Here’s a breakdown of teh key details from the provided text, focusing on the security vulnerability in AI systems:
the Vulnerability: Hidden Prompts in Images
How it works: Researchers at Trail of Bits discovered a way too hide malicious instructions within images. These instructions are invisible to the human eye in the original image.
The Trigger: When AI systems (like those from google – Gemini, Vertex AI, Assistant) downscale (resize) these images for processing, the downscaling process reveals the hidden text.
The Mechanism: This revelation happens due to image resampling techniques, specifically bicubic interpolation. This method can expose the hidden black text.
Impact & Demonstration
Real-World Threat: This isn’t just theoretical. Researchers successfully exploited the vulnerability to siphon Google Calendar data to an external email address without user permission.
Affected Systems: The attack worked on multiple Google AI platforms:
Gemini CLI
Vertex AI studio
Google Assistant (Android)
Gemini’s web interface
Building on Previous Research: The technique is based on a 2020 paper that identified image scaling as a potential attack surface for machine learning.
Key Takeaways
AI systems are vulnerable to attacks that exploit their image processing steps.
Hidden instructions can bypass initial security checks and be executed during image downscaling.
* This vulnerability has the potential for serious data breaches and unauthorized actions.
