LLM Vulnerabilities: Run-on Sentences & Image Scaling Risks
- Recent research from multiple labs demonstrates that large language models (LLMs), despite achieving high scores on benchmarks and ongoing claims of approaching artificial general intelligence (AGI), remain surprisingly...
- LLMs can be tricked into divulging sensitive data through carefully crafted prompts.
- The vulnerabilities extend beyond text-based interactions.llms are also vulnerable to manipulation through images containing embedded, imperceptible messages to humans.
AI Vulnerabilities Highlight Naiveté Despite Progress
Table of Contents
Published August 27, 2024
the Limits of Current AI Governance
Recent research from multiple labs demonstrates that large language models (LLMs), despite achieving high scores on benchmarks and ongoing claims of approaching artificial general intelligence (AGI), remain surprisingly susceptible to manipulation. These findings suggest a gap between performance in controlled settings and real-world submission, where common sense and critical thinking are essential.
Exploiting Linguistic Ambiguity
LLMs can be tricked into divulging sensitive data through carefully crafted prompts. One technique involves using excessively long, run-on sentences devoid of punctuation – specifically, avoiding periods or full stops. This approach appears to overwhelm the AI’s safety protocols,as governance systems “lose their way” when confronted with such unconventional input. This vulnerability underscores the models’ reliance on structured language and their difficulty processing ambiguity.
The vulnerabilities extend beyond text-based interactions.llms are also vulnerable to manipulation through images containing embedded, imperceptible messages to humans. This demonstrates a basic weakness in how these models interpret and process visual data, raising concerns about their reliability in security-sensitive applications.
