Your AI Will Say Something Terrible

Without guardrails, it's not a question of if but when your AI system generates something harmful, inaccurate, or embarrassing. Guardrails are the engineering systems that prevent bad outputs from reaching users, bad inputs from reaching the model, and sensitive data from leaking.

This isn't about censorship. It's about building AI systems that work correctly in production — the same way input validation isn't censorship for web forms.

Input Guardrails

Prompt Injection Detection

The #1 attack vector for LLM-powered applications. Users submit inputs designed to override your system prompt:

  • Pattern matching: flag inputs containing "ignore previous instructions", "system prompt:", "you are now"
  • Classifier model: train a small model to detect injection attempts. Rebuff and similar tools provide this
  • Input/output separation: clearly delimit user input in the prompt so the model treats it as data, not instructions

PII Detection

Users will paste sensitive data into your AI: social security numbers, credit cards, medical records. Detect and redact before the data reaches the model:

  • Regex patterns for SSN, credit card numbers, phone numbers, email addresses
  • NER (Named Entity Recognition) for names and addresses
  • Log redacted versions only — never log raw PII

Content Policy

Block inputs that violate your use case:

  • A customer service bot shouldn't process requests for illegal activities
  • A code assistant shouldn't help write malware
  • A medical information bot shouldn't provide specific treatment recommendations

Output Guardrails

Hallucination Detection

The model will make things up. In RAG systems, check that the answer is actually grounded in the retrieved documents:

  • Citation verification: require the model to cite specific sources. Check that the citation exists and supports the claim
  • Confidence scoring: use a second model call to score the answer's faithfulness to the context
  • Fact extraction: extract factual claims from the output and verify each one against the knowledge base

Output Validation

  • Schema validation: if the output should be JSON, validate it against a schema before returning
  • Length limits: cap response length to prevent runaway generation
  • Toxicity filtering: run outputs through a toxicity classifier (Perspective API, OpenAI moderation endpoint)
  • Brand safety: check that the model doesn't mention competitors, make unauthorized claims, or use language that violates brand guidelines

The Guardrail Pipeline

User Input
  → Input validation (length, encoding)
  → PII detection & redaction
  → Prompt injection detection
  → Content policy check
  → [ALLOWED] → LLM Call
    → Output toxicity check
    → Hallucination detection
    → Schema validation
    → PII re-check (model might generate PII)
    → [PASSED] → Return to user
    → [FAILED] → Fallback response + alert

Implementation Patterns

Synchronous Guards (Blocking)

Check before returning. Adds latency but prevents bad outputs from ever reaching users. Use for toxicity, PII, and schema validation.

Async Guards (Non-Blocking)

Check in the background after returning. Doesn't add latency but bad outputs may briefly be visible. Use for hallucination detection, brand safety, and analytics.

Human-in-the-Loop

For high-stakes outputs (financial advice, legal content, medical information), route to human review before delivery. The guardrail isn't an algorithm — it's a queue.

Monitoring Your Guardrails

  • Track trigger rates. If your injection detector fires on 30% of inputs, it's too aggressive. If it fires on 0%, it's not catching anything
  • Review false positives. Users blocked by guardrails they shouldn't have triggered will leave your product
  • Log everything. Input, guardrail decision, output, guardrail decision. This is your audit trail and training data
  • Set up alerts. Sudden spikes in guardrail triggers often indicate an attack or a model behavior change

The Cost of Guardrails

Each guardrail layer adds latency and cost. A typical production setup:

  • Input validation: +5ms
  • PII detection: +20ms
  • Injection detection: +50ms (classifier) or +200ms (LLM-based)
  • Output toxicity: +50ms
  • Hallucination check: +500ms (requires second LLM call)

Total overhead: 125-775ms depending on which guards you run. Budget this into your latency targets from the start.

Related Articles

TestingEvaluation

AI Testing and Evaluation

Building quality gates for production AI systems with golden datasets and regression gates.

MonitoringObservability

AI Monitoring and Observability

Catching silent failures in production with evaluation pipelines and alerting.

SecurityAI

AI Security and Threat Detection

Building AI-powered security monitoring and threat detection systems.