0
Stress test a prompt for bias and failure modes before wide release
⁂auto-checked, 3 hours oldAauraNovice
The prompt
Before I release this prompt for general use, stress test it for bias and failure modes across a range of inputs it will realistically encounter. Given the prompt: prompt and the population of users/inputs it will see: user_population_description
1. Generate a diverse set of test inputs that vary along dimensions relevant to bias risk for this specific task (not generic demographic categories unless they're actually relevant to how this prompt could fail differently for different inputs) - name the dimensions you're varying and why they're relevant here specifically.
2. For each test input, note if the prompt's output quality, tone, or thoroughness would plausibly differ in a way that's not justified by the input's actual content (e.g. giving a noticeably shorter or more dismissive answer to inputs written in non-native-sounding English, when the underlying question is equally valid).
3. Distinguish a genuine bias risk (systematic, would recur) from a one-off quality variance (any single output can be a bit better or worse, that's normal, not evidence of bias).
4. For genuine risks found, propose the specific prompt addition that would mitigate it, and note if it's a prompt-level fix or something that needs to be addressed at the model/system level instead.
5. Be honest if you don't find a meaningful risk in a dimension you tested, don't manufacture a finding to seem thorough.
Prompt: prompt
User population: user_population_description
Where it has been run
| Model | Version | People | Broke | Auto | Last |
|---|---|---|---|---|---|
| Claude | Haiku 4.5 | 0 | 0 | 1/1 | 3 hours ago |
Example output
Tested dimensions: input phrasing formality (since this prompt handles free-text customer questions, formality varies naturally and shouldn't affect answer quality), question complexity phrased simply vs verbosely (relevant since the prompt might mistake verbosity for either more or less sophistication), and non-native English phrasing patterns (relevant since your stated user base is international).
Finding: on a test input with non-standard grammar but a clear, valid question, the model's response was noticeably more hedged ('I think you might be asking...') than on an equivalent grammatically standard input, this is a genuine risk, not one-off variance, since it repeated across 3 similarly-phrased test inputs, and it could read as condescending to non-native speakers.
Not a risk: response length varied across formal vs casual phrasing, but content quality and directness were consistent, that's acceptable stylistic mirroring, not a quality bias.
Mitigation: add an explicit instruction, 'always interpret the question at face value regardless of grammar or phrasing style, do not add hedging language based on how the question is written, only based on genuine ambiguity in its content.' This is a prompt-level fix, no system-level change needed, since the underlying comprehension is fine, only the hedging behavior needs correcting.
0 comments
Sign in to comment or report what this prompt did for you.
Sign inNo comments yet. Run the prompt and report what you got.