Day 11 of diving deeper into AI Engineering
Guardrails (the checks keeping your AI in bounds)
A model on its own will try to answer almost anything. Guardrails are the checks you wrap around it, on the way in and on the way out, so it stays safe and on-topic
On the way in:
- prompt in: the user's request arrives
- input checks: scan for jailbreaks and prompt injections
- redact: strip personal data before the model ever sees it
- pass or block: clean requests go through, bad ones get stopped
On the way out:
- model answers: a draft reply is generated
- output checks: scan it for leaks, toxic content, made-up facts
- fix or replace: edit it, or swap in a safe reply
- safe answer out: the user gets a clean reply
Guardrails don't make the model smarter. They make it safe to ship