Skip to main content

Uniphore Help Center Portal

Banned Topic Guardrails

It's essential to have a safe and respectful conversation when interacting with people. Without proper safeguards in place, the LLM might receive and process harmful, offensive, malicious or illegal inputs, resulting in inappropriate responses that reduce customer trust or potentially damage brand reputation. Companies using AI-based assistants must follow strict ethical and legal guidelines for implementing AI features, including how content is moderated.

Uniphore uses a 'guardrail' mechanism to moderate all user input to an AI-based conversation, implemented via a list of topics that are specifically banned from reaching the LLM.

Whenever a banned topic is detected during a conversation, the guardrail will be triggered. The LLM will be instructed to return a polite, generic, pre-defined response. Finally, the full guardrail event will be logged in detail for later analysis and compliance purposes.

Currently, Uniphore guardrails will prevent any LLM processing for the following topics:

  • Controlled and/or Regulated Substances

  • Guns and Illegal Weapons

  • Harassment

  • Hate and Identity Hate

  • Personally Identifiable Information (PII/Privacy)

  • Profanity

  • Response Unsafe

  • Sexual

  • Sexual (Minor)

  • Suicide and Self-Harm

  • Threat

  • User Unsafe

  • Violence

Allowing Usage of a Specific Banned Topic

Although the above-referenced guardrail topics are banned globally for each Uniphore account, it is possible to select a specific banned topic that will be allowed for use during an AI-based conversation. This whitelisting mechanism can be performed in the relevant Task of an AI Agent:

xf-TaskBannedTopic-Example_240226.png

For more details on using this setting, click here.

Tip

We recommend that you consult with your Account Admin and other key stakeholders regarding which banned topics should be whitelisted for a specific Task in an associated AI Agent.