Introduction

Once your prompt contracts (Part 1), decoding (Part 2), style controls (Part 3), and context shaping (Part 4) are in place, the last line of defense is safety enforcement. Validators are the mechanical gatekeepers that guarantee outputs satisfy policy, structure, and jurisdictional constraints before they become user-visible or trigger actions. In this article we define a compact validator architecture, show how to encode policy as data, explain “repair small” strategies that avoid costly retries, and outline the metrics and incident playbooks that keep risk low without slowing teams down.


Why Validators Exist

Generative routes fail in predictable ways: malformed JSON, banned claims (“guaranteed results”), stale citations, locale mix-ups, and—most expensive—implied write actions (“I issued your refund”) with no execution record. Validators reduce these to a handful of deterministic checks. They return machine-readable error codes, enabling targeted repairs instead of whole-document resampling. The net effect: higher first-pass acceptance (CPR), lower latency, and fewer incidents.


The Validator Stack (four layers)

  1. Shape & Schema

    • Parseability (valid JSON/HTML/Markdown subsets)

    • Required fields present; enums/ranges respected

    • Section counts, bullet counts, sentence caps

  2. Lexicon & Tone

    • Banned terms/phrases; hedge phrases required where uncertainty remains

    • Brand casing and style constraints (e.g., sentence length ceilings, paragraph counts)

    • Channel rules (subject/preheader lengths, emoji policy, hashtags)

  3. Evidence & Freshness (when grounded)

    • Citation coverage for factual sentences (e.g., ≥70%)

    • Claim freshness within policy window (e.g., ≤540 days unless “evergreen”)

    • Conflict surfacing when claims disagree (show both with dates or abstain)

  4. Action Safety

    • Implied-write detector: blocks language asserting state change without a recorded tool result

    • Tool whitelist, argument validation, jurisdiction/tenant limits

    • Spend/rate/budget checks; idempotency requirements

Each layer yields specific error codes (e.g., SCHEMA, LEXICON, CITATION, STALE_CLAIM, IMPLIED_WRITE, LOCALE, LENGTH), making downstream remediation simple.


Policy as Data (not prose)

Centralize rules in a policy bundle the validators read and the prompt references. Example (abridged):

{
  "contract_version": "brand.v3.2.1",
  "ban": ["guarantee", "only solution", "revolutionary"],
  "hedge": ["may", "typically", "in many cases"],
  "comparatives": { "allow": ["faster than baseline"], "forbid": ["#1","best"] },
  "brand_casing": [["Product X","Product X"]],
  "citations": { "min_coverage": 0.7, "max_age_days": 540 },
  "locale": "en-US",
  "channel": { "email_subject_max": 52, "preheader_max": 90 },
  "writes": { "forbid_implied_success": true }
}

Because the bundle is versioned, audits can reproduce the exact rule set used for any output, and legal/brand can edit rules without touching code.


Repair Small: Deterministic Fixes Before Resample

Resampling entire outputs is slow and expensive. Prefer section-level and token-light repairs:

If a section fails the second pass, then tighten decoding for that section only (lower top-p/temperature) and try once more. On a third failure, emit a targeted ASK_FOR_MORE or REFUSE with the specific missing/forbidden condition.


Implementation Pattern

1) Contract hooks

2) Validator services

3) Repair engine

4) Observability


Example Error → Repair Mappings

ErrorDetectorRepair
SCHEMAJSON invalid / missing fieldsFill defaults; if still invalid → resample that section
LEXICONbanned phrase “guarantee”Substitute with hedge (“aims to”); add reason note
LENGTH>18 words/sentenceSplit at clause boundary; trim adverbs
CITATIONfactual line w/ no claim_idAttach nearest valid claim (threshold); else hedge/remove
STALE_CLAIMclaim age > windowSwap for fresher claim; else hedge timing (“as of …”)
IMPLIED_WRITEsuccess language w/o tool resultRewrite to proposal; or block and surface decision from validator
LOCALEEN-UK in EN-US routeNormalize spellings; enforce brand casing

Repairs must be deterministic and logged so operators can trust them.


Measuring Safety & Validator Efficacy

Visualize by route, locale, and release. Alert on CPR −2 pts, p95 +20%, or any implied-write spike.


Incident Playbooks (keep them short)

  1. Contain: Flip the route’s feature flag to safe copy; block tool execution if relevant.

  2. Triage: Classify failures by code (SCHEMA/CITATION/SAFETY/…); sample affected outputs.

  3. Root cause: Contract change, policy edit, decoder tweak, stale claims, or retrieval eligibility gap.

  4. Remediate: Patch the artifact (policy/contract/decoder/validator), add/adjust a golden case, re-canary.

  5. Postmortem: 1-pager with artifact diffs, failing samples, and a new regression guard.

Aim for MTTR in hours, not days; cheap rollback is part of safety.


Practical Tips


Worked Example (Composite)

A renewal email route fails validation with errors: LEXICON (uses “revolutionary”), CITATION (one factual sentence uncovered), and IMPLIED_WRITE (“I renewed your plan”). Repairs substitute “advanced,” attach a valid claim to the uncovered sentence, and rewrite the implied write into a proposal with a confirmation link. The section then passes; time-to-valid improves versus whole-resample, and the audit log records the three repairs. Canary shows CPR rising from 91.6% → 93.0%, p95 falling 14%.


Conclusion

Validators are not an add-on; they are the enforcement engine for prompt contracts and policy bundles. By failing closed, repairing small, and logging everything, you turn safety from an after-the-fact review into a first-class, automated property of your system. In Part 7, we’ll put evaluation and operations around these controls—golden tests, canary/rollback, and the metrics that prove a change is safe to ship.