Prompt injection detection guardrail
Binary classification of incoming prompts to detect prompt injection attempts and block malicious inputs before LLM processing
STop tier. Meets A, and the payoff is major with high confidence.5.94
Key facts
- Vertical
- Software & tech
- Function
- Agents & dev
- Status
- Seen in the wild
- Volume
- high
- Value
- major
- Risk
- high
- Evidence
- described plan
- Flags
- human-review
Source: https://docs.mozilla.ai/any-guardrail/api-reference/index/prompt-injection/injec-guard
Build this with a classifier
Define a typed decision with a bounded answer, then evaluate it on examples.
{
"decision_type": "yes_no",
"question": "Does this input match the decision in “Prompt injection detection guardrail”?",
"input": "<input to classify>",
"output": "yes | no"
}Related use cases
Cite this
Copy a link in your preferred format.