Skip to content

For authorized AI red-teaming and defensive research only. Test systems you own or are permitted to test. Read the boundaries.

Posts tagged: #prompt-injection

What authorized testing means for LLM apps: written scope, why prompt-injection research is legitimate for defenders, responsible disclosure, and what never to do.
The classic markdown-image beacon: how indirect prompt injection turns an auto-loaded image into a data-leak channel, and how to shut it down.
A layered blue-team defense for LLM apps: structured messages, instruction/data separation, I/O filtering, least-privilege tools, egress control, and monitoring.
How attacker-planted instructions in retrieved documents hijack a RAG assistant, why retrieval amplifies the risk, and the defenses that contain it.
A quick-reference of common LLM jailbreak techniques with one-line how-it-works and one-line defense for each, for authorized red-team use.
Quick reference for encoding and obfuscation bypasses used in prompt injection, with tiny examples and defenses like normalization and Unicode filtering.