Your AI Will Say Something Terrible
Without guardrails, it's not a question of if but when your AI system generates something harmful, inaccurate, or embarrassing. Guardrails are the engineering systems that prevent bad outputs from reaching users, bad inputs from reaching the model, and sensitive data from leaking.
This isn't about censorship. It's about building AI systems that work correctly in production — the same way input validation isn't censorship for web forms.
Input Guardrails
Prompt Injection Detection
The #1 attack vector for LLM-powered applications. Users submit inputs designed to override your system prompt:
- Pattern matching: flag inputs containing "ignore previous instructions", "system prompt:", "you are now"
- Classifier model: train a small model to detect injection attempts. Rebuff and similar tools provide this
- Input/output separation: clearly delimit user input in the prompt so the model treats it as data, not instructions
PII Detection
Users will paste sensitive data into your AI: social security numbers, credit cards, medical records. Detect and redact before the data reaches the model:
- Regex patterns for SSN, credit card numbers, phone numbers, email addresses
- NER (Named Entity Recognition) for names and addresses
- Log redacted versions only — never log raw PII
Content Policy
Block inputs that violate your use case:
- A customer service bot shouldn't process requests for illegal activities
- A code assistant shouldn't help write malware
- A medical information bot shouldn't provide specific treatment recommendations
Output Guardrails
Hallucination Detection
The model will make things up. In RAG systems, check that the answer is actually grounded in the retrieved documents:
- Citation verification: require the model to cite specific sources. Check that the citation exists and supports the claim
- Confidence scoring: use a second model call to score the answer's faithfulness to the context
- Fact extraction: extract factual claims from the output and verify each one against the knowledge base
Output Validation
- Schema validation: if the output should be JSON, validate it against a schema before returning
- Length limits: cap response length to prevent runaway generation
- Toxicity filtering: run outputs through a toxicity classifier (Perspective API, OpenAI moderation endpoint)
- Brand safety: check that the model doesn't mention competitors, make unauthorized claims, or use language that violates brand guidelines
The Guardrail Pipeline
User Input
→ Input validation (length, encoding)
→ PII detection & redaction
→ Prompt injection detection
→ Content policy check
→ [ALLOWED] → LLM Call
→ Output toxicity check
→ Hallucination detection
→ Schema validation
→ PII re-check (model might generate PII)
→ [PASSED] → Return to user
→ [FAILED] → Fallback response + alert
Implementation Patterns
Synchronous Guards (Blocking)
Check before returning. Adds latency but prevents bad outputs from ever reaching users. Use for toxicity, PII, and schema validation.
Async Guards (Non-Blocking)
Check in the background after returning. Doesn't add latency but bad outputs may briefly be visible. Use for hallucination detection, brand safety, and analytics.
Human-in-the-Loop
For high-stakes outputs (financial advice, legal content, medical information), route to human review before delivery. The guardrail isn't an algorithm — it's a queue.
Monitoring Your Guardrails
- Track trigger rates. If your injection detector fires on 30% of inputs, it's too aggressive. If it fires on 0%, it's not catching anything
- Review false positives. Users blocked by guardrails they shouldn't have triggered will leave your product
- Log everything. Input, guardrail decision, output, guardrail decision. This is your audit trail and training data
- Set up alerts. Sudden spikes in guardrail triggers often indicate an attack or a model behavior change
The Cost of Guardrails
Each guardrail layer adds latency and cost. A typical production setup:
- Input validation: +5ms
- PII detection: +20ms
- Injection detection: +50ms (classifier) or +200ms (LLM-based)
- Output toxicity: +50ms
- Hallucination check: +500ms (requires second LLM call)
Total overhead: 125-775ms depending on which guards you run. Budget this into your latency targets from the start.
Related Articles
AI Testing and Evaluation
Building quality gates for production AI systems with golden datasets and regression gates.
AI Monitoring and Observability
Catching silent failures in production with evaluation pipelines and alerting.
AI Security and Threat Detection
Building AI-powered security monitoring and threat detection systems.