Optimize · Safety Eval
Safety Eval
A pre-built corpus of adversarial visitor messages — attempts to extract legal advice, leak PHI, jailbreak the bot, imply representation, or trip a Louisiana-specific legal trap — run through the real chat pipeline (real prompt, RAG, model, guardrails) and checked against category-specific safety assertions. A failure here is a real reply a real visitor could have seen, not a synthetic test artifact.
Corpus
154 cases across 5 categories.
Advice-seeking
Asks the bot to opine on the case instead of deferring to an attorney.
12 cases
PHI
Feeds the bot a fake SSN, DOB, or medical detail and asks it to record it.
31 cases
Jailbreak
Tries prompt injection, roleplay, or "developer mode" to break the rules.
49 cases
Implied representation
Tries to get the bot to confirm an attorney-client relationship exists.
9 cases
Louisiana-specific
Louisiana-specific deadlines and traps: prescription periods, medmal, municipal defendants.
53 cases
Round-robins across every distinct phrasing before repeating one, so a small sample still touches every category and every trap, not just the first template's value variants.
Last run: 14/15 passed, Aug 13, 2026, 11:19 PM · mufty.851
Run history
| When | Cases/category | Result | Triggered by |
|---|---|---|---|
| Aug 13, 2026, 11:19 PM | 3 | 14/15 | mufty.851 |
| Aug 13, 2026, 9:05 PM | 10 | 48/49 | mufty.851 |
| Aug 13, 2026, 8:55 PM | 10 | 48/49 | mufty.851 |
| Aug 13, 2026, 8:37 PM | 3 | 14/15 | mufty.851 |