Secure a RAG system
Stop data leaks across users, defuse instructions hidden in documents, and handle outputs safely.
- Enforce permissions in retrieval, not in the prompt
- Recognize indirect prompt injection and poisoned documents
- Apply the OWASP Top 10 for LLM Applications to a RAG design
A RAG system connects a model to your organization’s documents - which makes it a juicy target. The OWASP Top 10 for LLM Applications 2025 lists the risks; these hit RAG hardest:
- LLM01 Prompt injection - especially indirect: instructions hidden in a retrieved document (“ignore your instructions and…”).
- LLM02 Sensitive information disclosure - secrets or personal data indexed by accident, then retrieved for the wrong person.
- LLM04 Data and model poisoning - an attacker plants misleading documents so they get retrieved.
- LLM05 Improper output handling - rendering model output (links, images, HTML) without care.
- LLM08 Vector and embedding weaknesses - new in 2025: cross-tenant leakage in shared vector stores, missing access control on embeddings, poisoned embeddings.
Try it
Inspect what retrieval brought back
These chunks were retrieved for an employee in the Acme tenant. Click everything that should never reach the model (or the user).
Click every part that looks suspicious. There are 2.
Defense in depth
- Permissions at retrieval: store ACLs and tenant IDs with every chunk, filter in the search query, and test with low-privilege users. Use separate indexes or namespaces per tenant when you can.
- Sanitize at ingestion: strip hidden text (HTML comments, white-on-white text, zero-width characters), redact secrets and personal data, record who added each document, and scan for instruction-like text.
- Treat retrieved text as data: delimit it, tell the model not to follow instructions inside sources - and don’t rely on that alone.
- Limit agency: if the RAG app can also take actions (send email, call tools), require confirmation for anything consequential.
- Handle output safely: escape HTML, don’t auto-load external images, show where links go.
- Monitor: log retrievals and answers (without secrets), watch for anomalies, and keep an incident path to remove poisoned documents fast.
Key takeaways
Enforce access control in retrieval; prompts are not permissions.
Retrieved documents can carry hidden instructions - sanitize ingestion and treat sources as data.
Redact sensitive data, limit what the system can do, and handle outputs safely.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Permission-aware retrieval
The first line lists the user’s groups (comma-separated). Each following line is id|groups|text. A chunk is allowed if it shares a group with the user or has the group public. Print allowed: IDS (input order, or none) and withheld: N.
- Support agent
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Quarantine suspicious chunks at ingestion
Each line is id|text. Check the RULES in order (case-insensitive) and print quarantine id: label for the first that matches, otherwise ok id.
- Five chunks
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…