# LLM output safety classification

Detect unsafe categories in LLM responses \(S1-S5 safety taxonomy\) for content filtering.

Grade: S — Top tier. Meets A, and the payoff is major with high confidence.. Score: 5.78.

Vertical: Software & tech.

Function: Agents & dev.

Status: Seen in the wild.

Volume: high.

Value: major.

Risk: high.

Evidence: described plan.

Flags: human-review.

- [Source](https://arxiv.org/pdf/2408.15488)
- [Vertical](/verticals/saas_tech/)
- [Function](/functions/agents_dev/)
- [Prompt injection detection guardrail](/use-cases/prompt-injection-detection-guardrail/)
- [LLM cost-quality routing with RouteLLM](/use-cases/llm-cost-quality-routing-with-routellm/)
- [Claude Code hooks pre-tool-use validation](/use-cases/claude-code-hooks-pre-tool-use-validation/)
- [Log sensitive-data leak](/use-cases/log-sensitive-data-leak/)
- [Prompt injection scanner](/use-cases/prompt-injection-scanner/)
- [Claude Code agent hooks for guardrails](/use-cases/claude-code-agent-hooks-for-guardrails/)
