Hallucinations, safety and guardrails
Spot fabricated details, reduce them with grounding and abstention, and keep humans in control of risky actions.
- Recognize common hallucination patterns
- Use grounding, abstention and self-consistency to reduce them
- Apply guardrails: least privilege, confirmation and data protection
A hallucination is fluent, confident output that isn’t supported by facts or sources: invented citations, wrong numbers, non-existent APIs, made-up quotes. It follows from how models work - they generate likely text, and a plausible fake can be likelier than “I don’t know”.
Ways to reduce it:
- Ground answers in retrieved sources or tool results, and require citations you can check.
- Allow abstention: explicitly invite “I don’t know” and reward it in evals.
- Self-consistency: sample several answers; low agreement signals uncertainty.
- Verify important claims in code or with people before acting on them.
Try it
Spot the hallucinations
A model answered a developer’s question. Click every claim that looks fabricated or unverifiable.
Click every part that looks suspicious. There are 3.
Guardrails for real systems
- Least privilege: give an assistant only the tools and data it needs.
- Confirmation for consequential actions - sending, paying, deleting, publishing.
- Untrusted input: documents, web pages and tool results can contain prompt injection; they’re data, not instructions.
- Data protection: don’t send secrets or unnecessary personal data in prompts; check provider data-retention terms.
- Monitoring and feedback: log (without secrets), review failures, and update evals.
The OWASP Top 10 for LLM Applications lists the main risks; the Cybersecurity and RAG tracks cover them in depth.
Key takeaways
Hallucinations are fluent, unsupported claims - invented APIs, citations and numbers are classic signs.
Ground, allow abstention, check consistency, verify what matters.
Least privilege, human confirmation and treating inputs as untrusted keep failures contained.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Self-consistency voting
Each input line is one sampled answer to the same question. Normalize answers (strip spaces, lowercase). If the most common answer’s share is at least 0.6, print answer: X (agreement A); otherwise print abstain: low agreement (A), with A to 2 decimals. Ties go to the alphabetically first answer.
- Agreement
- Disagreement
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Guard an assistant’s actions
Each line is an action the model wants to take. Print allow ACTION for read-only actions (search, read_file, get_weather), confirm ACTION for consequential ones (send_email, delete_file, make_payment), and block ACTION for anything else.
- Mixed actions
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…